Inception Labs’ Mercury 2 AI Beats Google’s DiffusionGemma at Its Own Game



In short

  • Inception Labs’ Mercury 2 generates about 1,000 tokens per second and scored 90 on AIME 2026.
  • Google’s latest DiffusionGemma runs similarly but performs worse in benchmarks.
  • DiffusionGemma is free and open source for Hugging Face. Mercury 2 is a paid, closed API version.

Inception Labs launched Mercury 2 on Thursday, calling it the world’s fastest language. According to the company’s announcement, it generates about 1,000 tokens per second – AI-style components that are read and written – versus about 89 tokens per second for Anthropic’s Claude Haiku 4.5 Discussion and 71 for OpenAI’s GPT-5 Mini.

This puts it in the speed bracket that Google will need later BroadcastGemma.

Both types go so far as to stop typing. A standard chatbot types one word, looks at what it just typed, then types the next one, and loops until the answer is complete. Composite models instead fill the sound with random tokens and remove the noise in several parallel passes – the same process that transforms image generators like Stable Diffusion – until the entire block is closed at once.

When the two separate and what survives that practice. On the AIME 2026—built from the real problems of the American Invitational Mathematics Examination and scored the highest number of successfully solved—Mercury 2 hit 90%. Google tested DiffusionGemma on the same list, where it scored 69.1%, while the standard, non-diffusion Gemma 4 scored 88.3% on the same test.

On GPQA, the benchmark for PhD science, the two models were almost identical: Mercury 2 at 77% against DiffusionGemma’s 73.2%. But Google’s own controller recommends the standard Gemma 4 for applications that require the highest quality, allowing DiffusionGemma to track them across the board.

The fast pace also works outside the lab. Augment Code, an AI coding-agent company, replaced Mercury 2 in Anthropic’s Claude Opus 4.7 for its case-compact subagent and achieved 82% lower latency and 90% lower cost, while reporting the same quality, according to a one lesson.

The startup was born out of research from its founder Stefano Ermon, a Stanford professor who also wrote some of the integration techniques that power today’s graphic designers. The initial $50 million funding was supported by Nvidia’s co-investors and private investors Andrew Ng and Andrej Karpathy.

For non-technical users, the biggest thing most people don’t see until they hear it is “flow.” Traditional examples make you wait between ideas in a long column. Integrated models like this make the AI ​​feel like it’s walking with you – instant completion, quick iterations on code or plans, and sub-agents that can do the most tedious work without dragging the whole system down.

The subagent component is an interesting change in architecture. Complex AI systems are no longer a single type of intelligence. It’s a set of specialized helpers: one for deep thinking, several for short-term, navigation, tool-checking, output-checking, and so on. Serial models make mobile phones more expensive and slower. Similar features make them cheap and fast enough to use liberally.

A real note for regular users: This is best for slow, high-speed sessions rather than heavy-duty simulations (where the biggest AR models can be limited at this point). Mercury 2 is not open source, so it’s API/cloud based for now. And like the Google brand, the entire environment (local runtimes, agent frameworks) is still working so that it doesn’t get lost anywhere.

Use events that make sense immediately: fast real-time programming and “vibe coding” where the model has your own changes, multi-agent or support systems where slow calls are made, interactive voiceovers that aren’t heavy, and long-term forecasting or next-generation forecasting. On a large scale, the cost and energy savings from advanced features on conventional hardware add up quickly.

Numbers The first stages (and independent) makes the case for it: the Mercury 2 sits in the “fastest and best” quadrant of hybrid models, pushing the demands of external hardware up to commodity GPUs.

Daily Debrief A letter

Start each day with top stories right here, including originals, podcasts, videos and more.





Source link

Leave a Reply

Your email address will not be published. Required fields are marked *