Signals
DeepSeek Gave V4 Flash a Major Agent Upgrade Without Raising Its Price

DeepSeek re-post-trained V4 Flash and replaced the model behind the existing deepseek-v4-flash API slug with a much stronger agent-focused version. Its vendor-reported agent results jumped, native Responses API and Codex support arrived, and the price stayed at $0.14 per million cache-miss input tokens and $0.28 per million output tokens.
What Changed
- On July 31, DeepSeek replaced the preview model behind the existing deepseek-v4-flash slug with V4-Flash-0731. The architecture and size stayed the same; the improvement came from agent-focused re-post-training.
- DeepSeek reports Terminal-Bench 2.1 at 82.7, NL2Repo at 54.2, DeepSWE at 54.4, Toolathlon Verified at 70.3, and Automation Bench Public at 25.1.
- It also adds native Responses API and Codex support, a 1M-token context window, and up to 384K output.
- The already-low list price did not rise: $0.0028 per million cache-hit input tokens, $0.14 per million cache-miss input tokens, and $0.28 per million output tokens.
- These are vendor-reported results run at maximum effort with DeepSeek's forthcoming minimal harness; two disclosed test sets are internal.
Who Should Care
- Teams running high-volume coding agents, background workers, subagents, or research loops where frontier-model pricing makes experimentation uneconomic.
- Existing deepseek-v4-flash users, because the upgraded model arrived behind the same slug without a configuration change.
- Builders using OpenAI-compatible tooling or Codex who can now test a far more capable agent model at the same unusually low API price.
Who Should Not Care
- Teams that cannot send the workload through their chosen hosting route for privacy, residency, compliance, or policy reasons.
- Products that require a pinned model revision, independently reproduced quality evidence, or mature production guarantees before adoption.
- Workloads where reliability and support matter more than token economics; the release remains a public beta and needs careful validation.
The Verdict
The real Signal is a major increase in usable agent capability without a price increase. DeepSeek is making long-running agent loops and background work economically viable at volumes that would be difficult to justify on frontier-priced models. That does not prove equal quality on every workload, but it absolutely changes which models deserve an evaluation. Benchmark the 0731 model on your own tasks, route it through the provider boundary you are comfortable with, and monitor quality because the stable public slug can hide future backing-model changes.
Take This To Your Agent
Copy a ready-to-paste investigation handoff that asks your coding agent to check this update against your product and recommend what to do next.