Mercury 2.5 Preview
Unknown · 1 configuration
Tariff: NanoGPT · Source: models-dev · USD per million tokens
What does more thinking buy you?
Compare this model’s measured thinking configurations.
No pair of measurements available
Change the filters or pick another benchmark. Missing values are not estimated.
Mercury 2.5 Preview is Inception's latest and most intelligent diffusion language model. Instead of generating tokens strictly one at a time, it produces and refines multiple tokens in parallel, reaching up to 1,107 tokens per second on standard GPUs. It delivers a 10+ point intelligence gain over Mercury 2, with tunable reasoning, parallel tool calls, schema-aligned JSON output, and a 260K context window. It is built for latency-sensitive production work such as search agents, voice pipelines, customer support, rapid coding iteration, and coding subagents.
This model is in the catalog but has no imported benchmarks. Capabilities are not estimated.