DeepSeek V4 Flash 0731 placed among the stronger open-weight reasoning models in an evaluation by Artificial Analysis, combining a high composite intelligence score with quick output but generating substantially more tokens than the median system in its comparison group.
The model, released on July 31, scored 52 on version 4.1.1 of the Artificial Analysis Intelligence Index. The median for comparable models was 28, according to the evaluator. The index combines results from a collection of tests covering practical analytical work, scientific coding, difficult general questions, knowledge reliability and long-context reasoning.
Artificial Analysis reported that the model produced output at 119 tokens per second, which it characterized as notably fast. It supports text input and text output and offers a context window of one million tokens. A context window of that size can support workflows that place long documents or retrieved material into a single request, though the maximum input capacity alone does not establish how reliably a model uses every part of that material.
The evaluation also identified verbosity as a tradeoff. DeepSeek V4 Flash 0731 generated about 210 million tokens while completing the Intelligence Index, compared with a 110 million-token median for the class. Longer reasoning traces or answers can affect latency and the total cost of tasks even when a provider's per-token rate is competitive.
Listed prices were 44 cents per million input tokens and $1.32 per million output tokens. Artificial Analysis described each rate as somewhat above its comparison medians of 30 cents and $1.17 respectively. The complete Intelligence Index run cost $323.26, reflecting both those rates and the model's high output volume.
The reported score should be read within the evaluator's methodology rather than as a universal ranking. Version 4.1.1 combines GDPval-AA v2, τ³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience and AA-LCR. Different production uses may give those components different importance. A coding assistant, document-analysis tool and customer-facing chatbot can each value accuracy, speed, refusal behavior, context length and brevity differently.
Artificial Analysis also notes that its speed figures use a first-party interface when one exists, or the median across providers for models without one. Real-world throughput can therefore vary with hosting, load, prompt size and deployment configuration.
On the supplied evidence, DeepSeek V4 Flash 0731's main proposition is a strong intelligence result without sacrificing generation speed. Its unusually high token consumption complicates that picture: teams comparing models will need to consider cost per completed task and answer efficiency, not only list prices or benchmark position. The reasoning configuration evaluated here may also differ from lighter-effort settings used in latency-sensitive applications.



