← Back to live feed1 story
1
Inception Launches Mercury 2.5 LLM at 1,100 Tokens Per Second, Topping OpenRouter Speed Leaderboard
Models23d agoThe new Mercury 2.5 model from Inception costs $0.04 per million input tokens and $0.15 per million output tokens at general availability. Available on OpenRou…
Sign in to suggest edits
Key sources
- SOURCE@ollama“DeepSeek-V4.1-Flash is now being rolled out on Ollama's cloud, starting with Max and Team accounts.”x.com
- SUPPORT@askalphaxiv“Together, these reduce global KV cache to 890 bytes/token, ~4x smaller than V4-Flash, and persistent KV by ~8x, while supporting 1M token contexts with much stronger overall performance.”x.com
- SUPPORT@teortaxestex“DeepSeek V4.1 Flash takes first place on AutomationBench-AA with 69%, equal to GPT-6 Astra (69%) and slightly above Grok 4.6 (67%).”x.com
- SUPPORT@ekzhang1“DeepSeek V4.1 has mathematically different outputs depending on whether you prefix cache hit”x.com
- SUPPORT@ollama“DeepSeek-V4.1-Flash is now rolling out for Pro plan subscribers.”x.com
- SOURCE@askvenice“fully private.”x.com
- SOURCE@justinsuntron“Supporting a 1M context window and up to 384K output”x.com
- SUPPORT@therundownai“It beats GPT-5.6 Sol and Claude Opus 5 on several agentic, coding, and cyber benchmarks.”x.com