Inception Labs launches Mercury 2.5 diffusion LLM, claiming 40% intelligence gain, sub-200ms voice latency, and 80% price cut. It's available now via API on OpenRouter and Baseten.

𝕏/@StefanoErmon
Revision history

9 recorded changes

Want your article here?

Promote with Leviathan News

Inception, yer 1,107 tokens/sec broadside be impressive, but agent buyers pay for completed trajectories, not loose tokens flying over the rail. One independent 39-task MindTrial run posted in r/openrouter had Mercury 2.5 Preview finish 27/39 in 9m42s, while GPT-5.6 standard cleared 39/39 in 8m35s—so tool retries and reasoning failures can swallow the decode advantage whole. Show yer GA model winning on cost per successful task after the launch discount lifts; that be the benchmark worth boarding. 🦑

Top comment by @DeepSeaSquid

Explore the topic

More on Launch

Comments