Phi-4
Microsoft · 1 configuration
Intelligence5.9
Input / 1M$0.13
Output / 1M$0.50
Context16K
Tariff: Microsoft · Source: artificial-analysis · USD per million tokens
Thinking configuration
What does more thinking buy you?
Compare this model’s measured thinking configurations.
Scroll the chart ↔ · tap a point for details
Efficient frontier: best measured score at each cost1 measured configurations · 0 without both values
Selected: thinking default · Click a point to switch configuration
Microsoft Research Phi-4 is designed to perform well in complex reasoning tasks and can operate efficiently in situations with limited memory or where quick responses are needed. At 14 billion...
Intelligence IndexArtificial Analysis composite index, not a percentage. Different versions of the index are not directly comparable.
5.9 pointsGPQA DiamondGraduate-level scientific questions. Measures specialist reasoning rather than web-search quality.
57.5 %Humanity’s Last ExamDifficult questions across many disciplines, compared under the same evaluation setup.
3.8 %τ²-BenchTool use and interaction in assistance scenarios. Does not replace an evaluation of your integration.
0.0 %IFBenchVerifiable instruction following. A proxy for precision rather than a benchmark of creative style.
23.5 %Omniscience accuracyFactual accuracy in Artificial Analysis's Omniscience evaluation.
14.1 %Specifications from the catalog
Inputtext
Outputtext
Tool callingNo
Structured outputYes
Speed40.2 tok/s
First token2.59 s
Maximum output14.7K
Catalog listing date2025-01-10