The leaderboard placed Claude Fable 5 with an Opus 4.8 fallback next, at 62. A high-effort Claude Opus 5 configuration and GPT-5.6 Sol at maximum settings each scored 61. The ordering shows that inference settings and model configurations are treated as distinct entries, so the table should not be read simply as a ranking of provider brand names.
Artificial Analysis constructs the index from a set of evaluations rather than a single test. The listed components include GDPval-AA v2, a banking assessment, Terminal-Bench v2.1, SciCode, Humanity’s Last Exam, GPQA Diamond, CritPt, an omniscience measure and a long-context reasoning assessment. Different applications may place different weight on those abilities, and a composite score does not guarantee leadership on every task.
The service also tracks characteristics beyond its intelligence measure. These include output speed, latency, context capacity, model availability and token pricing. Its cost calculations account for input, cache use, reasoning and answer-token charges, while speed figures use first-party application programming interfaces when available or a median across providers where no first-party interface exists. Time-to-first-token and total time for a 500-token response are reported separately.
On the page’s speed comparison, Celeris-1 led at 1,574.3 generated tokens per second, ahead of Mercury 2 at 977.2. Llama 3.1 Instruct 8B and Granite 4.2 3B were listed among the least expensive options at a blended $0.02 per million tokens. Those categories underline a practical limitation of an intelligence-only ranking: the highest composite score may not be the best choice when responsiveness, deployment rights or cost dominate a workload.
The leaderboard identified Kimi K3 at maximum settings as the highest-ranked open-weights entry, scoring 60. It counted 97 open-weights models among the 177 systems evaluated. Qwen3.8 2.4T A95B followed at 58, with GLM-5.3-Flash at 57. The site separately labels models whose licenses restrict commercial use or prohibit it.
Leaderboard positions can change as models, configurations and methodology are updated. The result therefore records a particular version of Artificial Analysis’s framework, not a permanent or universal verdict on model quality. For buyers and developers, its most informative use is comparative: intelligence, speed, price, latency and licensing can be viewed together against the needs of a specific system. Direct benchmark comparisons also inherit the assumptions of their test harnesses, scoring rules and provider settings, making methodology as important as a model’s headline position.



