aiexpert
Home / News / Brief
Research · Aug 10, 2026, 02:03 PM · 4 sources

DeepSeek V4-Flash official beats V4-Pro on agent benchmarks; Terminal Bench 2.1 hits 82.7

DeepSeek released DeepSeek-V4-Flash-0731 to public beta on July 31, 2026, as an official release of the model that beat V4-Pro-Preview on all nine published agent benchmarks, with Terminal Bench 2.1 scoring 82.7 at the same $0.14 per 1M input and $0.28 per 1M output pricing.

The V4-Flash-0731 uses the exact same architecture and size as the April preview version (284B total params, ~13B active in Mixture-of-Experts)—only the post-training was redone. Agent benchmarks included Cybergym 76.7, NL2Repo 54.2, DeepSWE 54.4, Toolathlon verified 70.3, Agent Last Exam 25.2, Automation Bench 25.1, DSBench-FullStack 68.7 and DSBench-Hard 59.6.

The official V4-Flash natively supports the Responses API format and has been specifically adapted for Codex, allowing direct integration without translation shim. At $0.14 per million input tokens, V4-Flash is the most cost-effective agent model on the market, delivering 82.7 on Terminal Bench 2.1 versus Opus-4.8's 85.0, a gap of just 2.3 points, at a fraction of the price. For builders: a cheaper model beating the expensive tier on agentic work through post-training signals a narrowing quality gap on standard agent tasks. Verify against your own workload before swapping.

Sources

Everything this brief rests on
  1. 01 Primary source deepseek.ai
  2. 02 flowtivity.ai flowtivity.ai “DeepSeek V4-Flash just scored 82.7 on Terminal Bench 2.1, beating V4-Pro-Preview by 14.7%”
  3. 03 morphllm.com morphllm.com “DeepSeek V4-Flash: 284B total, 13B active, $0.14/M input, $0.28/M output”
  4. 04 deepseekv4guide.org deepseekv4guide.org “Terminal-Bench 2.1 hits 82.7; DeepSWE 54.4; Agent Last Exam 25.2”