DeepSeek V4.1 Flash vs Opus 5: Where It Wins and Loses
Todd Brooks, Founder6 min read

Jump to a section11
TL;DR
- Where it wins: on DeepSeek's own table, V4.1 Flash edges Opus 5 on DeepSWE, Terminal-Bench 2.1, AutomationBench, Agent's Last Exam and HLE with tools.
- Where it loses: GPQA Diamond, Terminal-Bench 3.0 and 4.0, whole-repository coding, ExploitGym, hard reasoning without tools and three image-reasoning tests, several by double digits.
- The pattern: it has caught up on older, well-worn tests. The frontier still leads on the newest, hardest ones.
- What to do: route by task, test on your own work, and score the model like any vendor with the AI Vendor Reliability & Spend-Justification Scorecard.
Is DeepSeek V4.1 Flash better than Claude Opus 5?
Not across the board. On DeepSeek's published table, V4.1 Flash leads Opus 5 on DeepSWE (74.2 to 74.0), Terminal-Bench 2.1, AutomationBench and Agent's Last Exam, and edges it on HLE with tools. It trails Opus 5 on GPQA Diamond, Terminal-Bench 3.0 and 4.0, NL2Repo-Bench, ProgramBench, ExploitGym, HLE without tools and three image-reasoning tests. All of these are the vendor's own figures, not independent results.
Best for: teams deciding which work to try on V4.1 Flash first, and which to leave where it is.
Every model launch comes with a chart where the new model wins. DeepSeek's chart is more honest than most, because it includes the rows where V4.1 Flash loses. It's tempting to quote only the rows where it wins.
This is the whole table, with what I'd take from it. It's part of DeepSeek V4.1 Flash changes the economics of AI agents.
All 19 benchmarks
| Benchmark | V4.1 Flash | Opus 5 | GPT-5.6 Sol |
|---|---|---|---|
| GPQA Diamond | 90.9 | 93.4 | 94.1 |
| HLE, no tools | 36.8 | 56.3 | 44.5 |
| Codeforces rating | 3471 | — | — |
| MathArena Apex | 65.6 | — | — |
| Terminal-Bench 2.1 | 90.6 | 89.1 | 88.8 |
| Terminal-Bench 3.0 | 30.0 | 43.3 | 34.4 |
| Terminal-Bench 4.0 | 31.2 | 51.8 | 39.9 |
| DeepSWE v1.1 | 74.2 | 74.0 | 73.0 |
| ProgramBench | 20.3 | 37.0 | 23.0 |
| NL2Repo-Bench | 64.0 | 75.3 | 56.8 |
| CyberGym | 88.1 | — | 84.5 |
| SEC-Bench Pro | 62.8 | — | 74.3 |
| ExploitGym | 15.3 | 22.1 | 33.7 |
| HLE with tools | 63.9 | 63.6 | — |
| AutomationBench | 54.8 | 50.3 | 45.8 |
| Agent's Last Exam | 31.8 | 28.6 | 26.7 |
| Chartography with tools | 78.9 | 84.0 | 79.9 |
| BabyVision with tools | 89.6 | 94.1 | 88.9 |
| ZeroBench with tools | 49.0 | 52.0 | 53.0 |
From the DeepSeek V4.1 Flash model card, September 10, 2026. A dash means the table shows no score for that model. DeepSeek's table also compares other models not shown here.
A misquote to watch for
If you've seen CyberGym reported as 88.1 for V4.1 Flash against 84.5 for Opus 5, the 84.5 belongs to GPT-5.6 Sol. DeepSeek's table shows no Opus 5 score for CyberGym at all. The same slip is easy to make on AutomationBench: 45.8 is Sol's score. Opus 5 scored 50.3.
It's an easy mistake when three columns sit side by side. It also flips the story from "V4.1 Flash beats Opus 5 everywhere" to something more useful.
The pattern: old tests versus new tests
Terminal-Bench is the clearest case, because it comes in three versions. On version 2.1, V4.1 Flash leads with 90.6. On 3.0 it trails Opus 5 by 13 points. On 4.0 it trails by 21. The newer the test, the wider the gap.
The coding picture looks the same. V4.1 Flash wins DeepSWE by two-tenths of a point. It trails Opus 5 by about 11 points on NL2Repo-Bench, where the task is building a whole repository, and by nearly 17 points on ProgramBench.
V4.1 Flash score minus Opus 5 score, in points, from DeepSeek's table. Bar length is roughly scaled to the largest gap.
My read is that V4.1 Flash has caught the frontier on the kinds of tasks the benchmarks have mostly solved, and still trails where the tests are newest and hardest. For routine agent work, that may be all you need. For long, novel engineering jobs, it probably isn't yet.
Three things the table doesn't tell you
- How hard it was trying. Every V4.1 Flash score was run at maximum reasoning effort, 100 out of 100, which is also the setting that uses the most tokens. Cheaper settings may score lower, by an amount I haven't seen published.
- How the other columns were run. DeepSeek ran its coding agents in its own harness with a 1M-token context window. The model card I read doesn't say how the Opus 5 and GPT-5.6 Sol numbers were produced.
- How it does on your work. A benchmark is someone else's tasks. The only score that decides your routing is the share of your own tasks each model gets right.
How I'd route work
| Kind of work | Where I'd start testing | Why |
|---|---|---|
| High-volume triage, tagging, extraction with an automatic check | V4.1 Flash | A miss is cheap to catch and retry, so price dominates |
| Routine coding and terminal tasks, with tests gating every merge | V4.1 Flash | Leads DeepSWE and Terminal-Bench 2.1 |
| Long, unfamiliar engineering or whole-repository work | Opus 5 | 51.8 to 31.2 on Terminal-Bench 4.0, 75.3 to 64.0 on NL2Repo-Bench |
| Hard reasoning with no tools | Opus 5 or GPT-5.6 Sol | HLE: 56.3 and 44.5 against 36.8 |
| Reading charts and images | Opus 5 | Leads Chartography and BabyVision |
None of this is a final answer. It's where I'd spend the first week of testing, using the per-task math in DeepSeek V4.1 Flash pricing: what a task really costs.
A model is a vendor
Once V4.1 Flash is doing real work, stop thinking of it as a benchmark entry and start treating it like any other supplier. Does it stay up? Does it hold quality from week to week? Does its behavior change without notice?
That last one is live right now. DeepSeek's changelog says that after 12:00 Beijing time on September 14, 2026, every request to deepseek-v4-pro will be served by V4.1 Flash and billed at V4.1 Flash prices, until V4.1 Pro ships. If you're on V4 Pro, your model changes that day and your code doesn't. More on that in is DeepSeek V4.1 Flash safe for business data?
Score the model like any other vendor
The AI Vendor Reliability & Spend-Justification Scorecard ($79, one-time) scores each AI vendor on reliability and whether the spend is justified, and returns RENEW, RENEGOTIATE or DO NOT RENEW. Its reliability-floor gate refuses to justify renewing a business-critical vendor you can't rely on, however cheap it is.
Get the Scorecard — $79 →Decision Guide
Start here if: you're choosing which workloads to test on V4.1 Flash.
Skip it if: you've already run your own side-by-side on real tasks. Your numbers beat this table.
Best first step: pick the one workflow in the first row of the routing table and test it this week.
More in this guide
FAQ
Does DeepSeek V4.1 Flash beat Opus 5?
On some benchmarks. DeepSeek's table has it ahead on DeepSWE, Terminal-Bench 2.1, AutomationBench, Agent's Last Exam and HLE with tools, and behind on GPQA Diamond, Terminal-Bench 3.0 and 4.0, NL2Repo-Bench, ProgramBench, ExploitGym, HLE without tools and three image tests.
Is DeepSeek V4.1 Flash better than GPT-5.6 Sol?
It's split. On DeepSeek's table, V4.1 Flash leads Sol on DeepSWE, Terminal-Bench 2.1, NL2Repo-Bench, CyberGym, AutomationBench, Agent's Last Exam and BabyVision. Sol leads on GPQA Diamond, HLE without tools, Terminal-Bench 3.0 and 4.0, ProgramBench, SEC-Bench Pro, ExploitGym, Chartography and ZeroBench.
What did Opus 5 score on CyberGym?
DeepSeek's table shows no Opus 5 score for CyberGym. The 84.5 sometimes quoted against V4.1 Flash's 88.1 is GPT-5.6 Sol's score.
Are these benchmarks independent?
No. DeepSeek published them on its own model card. They're useful for deciding what to test, but they don't replace testing on your own tasks.
What reasoning effort were the benchmarks run at?
Every V4.1 Flash result used the maximum setting, reasoning effort 100, which is also the setting that uses the most output tokens.
Cheap isn't the same as reliable
Score every AI vendor on reliability and spend with a gate that won't be talked out of it. $79, offline.
Get the AI Vendor Reliability & Spend-Justification Scorecard — $79 →

The gate this post refers to, drawn from the tool’s own logic. See the tool.