DeepSeek V4.1 Flash vs Opus 5: Where It Wins and Loses

Todd Brooks, Founder6 min read

A woman crouches by a blank gauge on a long test rail where glowing blue blocks stop short of a red line
Jump to a section11

TL;DR

  • Where it wins: on DeepSeek's own table, V4.1 Flash edges Opus 5 on DeepSWE, Terminal-Bench 2.1, AutomationBench, Agent's Last Exam and HLE with tools.
  • Where it loses: GPQA Diamond, Terminal-Bench 3.0 and 4.0, whole-repository coding, ExploitGym, hard reasoning without tools and three image-reasoning tests, several by double digits.
  • The pattern: it has caught up on older, well-worn tests. The frontier still leads on the newest, hardest ones.
  • What to do: route by task, test on your own work, and score the model like any vendor with the AI Vendor Reliability & Spend-Justification Scorecard.

Is DeepSeek V4.1 Flash better than Claude Opus 5?

Not across the board. On DeepSeek's published table, V4.1 Flash leads Opus 5 on DeepSWE (74.2 to 74.0), Terminal-Bench 2.1, AutomationBench and Agent's Last Exam, and edges it on HLE with tools. It trails Opus 5 on GPQA Diamond, Terminal-Bench 3.0 and 4.0, NL2Repo-Bench, ProgramBench, ExploitGym, HLE without tools and three image-reasoning tests. All of these are the vendor's own figures, not independent results.

Best for: teams deciding which work to try on V4.1 Flash first, and which to leave where it is.


Every model launch comes with a chart where the new model wins. DeepSeek's chart is more honest than most, because it includes the rows where V4.1 Flash loses. It's tempting to quote only the rows where it wins.

This is the whole table, with what I'd take from it. It's part of DeepSeek V4.1 Flash changes the economics of AI agents.

All 19 benchmarks

BenchmarkV4.1 FlashOpus 5GPT-5.6 Sol
GPQA Diamond90.993.494.1
HLE, no tools36.856.344.5
Codeforces rating3471
MathArena Apex65.6
Terminal-Bench 2.190.689.188.8
Terminal-Bench 3.030.043.334.4
Terminal-Bench 4.031.251.839.9
DeepSWE v1.174.274.073.0
ProgramBench20.337.023.0
NL2Repo-Bench64.075.356.8
CyberGym88.184.5
SEC-Bench Pro62.874.3
ExploitGym15.322.133.7
HLE with tools63.963.6
AutomationBench54.850.345.8
Agent's Last Exam31.828.626.7
Chartography with tools78.984.079.9
BabyVision with tools89.694.188.9
ZeroBench with tools49.052.053.0

From the DeepSeek V4.1 Flash model card, September 10, 2026. A dash means the table shows no score for that model. DeepSeek's table also compares other models not shown here.

A misquote to watch for

If you've seen CyberGym reported as 88.1 for V4.1 Flash against 84.5 for Opus 5, the 84.5 belongs to GPT-5.6 Sol. DeepSeek's table shows no Opus 5 score for CyberGym at all. The same slip is easy to make on AutomationBench: 45.8 is Sol's score. Opus 5 scored 50.3.

It's an easy mistake when three columns sit side by side. It also flips the story from "V4.1 Flash beats Opus 5 everywhere" to something more useful.

The pattern: old tests versus new tests

Terminal-Bench is the clearest case, because it comes in three versions. On version 2.1, V4.1 Flash leads with 90.6. On 3.0 it trails Opus 5 by 13 points. On 4.0 it trails by 21. The newer the test, the wider the gap.

The coding picture looks the same. V4.1 Flash wins DeepSWE by two-tenths of a point. It trails Opus 5 by about 11 points on NL2Repo-Bench, where the task is building a whole repository, and by nearly 17 points on ProgramBench.

V4.1 Flash score minus Opus 5 score, in points, from DeepSeek's table. Bar length is roughly scaled to the largest gap.

My read is that V4.1 Flash has caught the frontier on the kinds of tasks the benchmarks have mostly solved, and still trails where the tests are newest and hardest. For routine agent work, that may be all you need. For long, novel engineering jobs, it probably isn't yet.

Three things the table doesn't tell you

  1. How hard it was trying. Every V4.1 Flash score was run at maximum reasoning effort, 100 out of 100, which is also the setting that uses the most tokens. Cheaper settings may score lower, by an amount I haven't seen published.
  2. How the other columns were run. DeepSeek ran its coding agents in its own harness with a 1M-token context window. The model card I read doesn't say how the Opus 5 and GPT-5.6 Sol numbers were produced.
  3. How it does on your work. A benchmark is someone else's tasks. The only score that decides your routing is the share of your own tasks each model gets right.

How I'd route work

Kind of workWhere I'd start testingWhy
High-volume triage, tagging, extraction with an automatic checkV4.1 FlashA miss is cheap to catch and retry, so price dominates
Routine coding and terminal tasks, with tests gating every mergeV4.1 FlashLeads DeepSWE and Terminal-Bench 2.1
Long, unfamiliar engineering or whole-repository workOpus 551.8 to 31.2 on Terminal-Bench 4.0, 75.3 to 64.0 on NL2Repo-Bench
Hard reasoning with no toolsOpus 5 or GPT-5.6 SolHLE: 56.3 and 44.5 against 36.8
Reading charts and imagesOpus 5Leads Chartography and BabyVision

None of this is a final answer. It's where I'd spend the first week of testing, using the per-task math in DeepSeek V4.1 Flash pricing: what a task really costs.

A model is a vendor

Once V4.1 Flash is doing real work, stop thinking of it as a benchmark entry and start treating it like any other supplier. Does it stay up? Does it hold quality from week to week? Does its behavior change without notice?

That last one is live right now. DeepSeek's changelog says that after 12:00 Beijing time on September 14, 2026, every request to deepseek-v4-pro will be served by V4.1 Flash and billed at V4.1 Flash prices, until V4.1 Pro ships. If you're on V4 Pro, your model changes that day and your code doesn't. More on that in is DeepSeek V4.1 Flash safe for business data?

Key insight: a model that costs 33 to 42 times less per token than Opus 5 doesn't earn a place in a business-critical workflow on price alone. If you can't rely on it, the cheap rate is the wrong number to look at.

Score the model like any other vendor

The AI Vendor Reliability & Spend-Justification Scorecard ($79, one-time) scores each AI vendor on reliability and whether the spend is justified, and returns RENEW, RENEGOTIATE or DO NOT RENEW. Its reliability-floor gate refuses to justify renewing a business-critical vendor you can't rely on, however cheap it is.

Get the Scorecard — $79 →

Decision Guide

Start here if: you're choosing which workloads to test on V4.1 Flash.

Skip it if: you've already run your own side-by-side on real tasks. Your numbers beat this table.

Best first step: pick the one workflow in the first row of the routing table and test it this week.

More in this guide

FAQ

Does DeepSeek V4.1 Flash beat Opus 5?

On some benchmarks. DeepSeek's table has it ahead on DeepSWE, Terminal-Bench 2.1, AutomationBench, Agent's Last Exam and HLE with tools, and behind on GPQA Diamond, Terminal-Bench 3.0 and 4.0, NL2Repo-Bench, ProgramBench, ExploitGym, HLE without tools and three image tests.

Is DeepSeek V4.1 Flash better than GPT-5.6 Sol?

It's split. On DeepSeek's table, V4.1 Flash leads Sol on DeepSWE, Terminal-Bench 2.1, NL2Repo-Bench, CyberGym, AutomationBench, Agent's Last Exam and BabyVision. Sol leads on GPQA Diamond, HLE without tools, Terminal-Bench 3.0 and 4.0, ProgramBench, SEC-Bench Pro, ExploitGym, Chartography and ZeroBench.

What did Opus 5 score on CyberGym?

DeepSeek's table shows no Opus 5 score for CyberGym. The 84.5 sometimes quoted against V4.1 Flash's 88.1 is GPT-5.6 Sol's score.

Are these benchmarks independent?

No. DeepSeek published them on its own model card. They're useful for deciding what to test, but they don't replace testing on your own tasks.

What reasoning effort were the benchmarks run at?

Every V4.1 Flash result used the maximum setting, reasoning effort 100, which is also the setting that uses the most output tokens.

Cheap isn't the same as reliable

Score every AI vendor on reliability and spend with a gate that won't be talked out of it. $79, offline.

Get the AI Vendor Reliability & Spend-Justification Scorecard — $79 →
How it decides
Diagram of the AI Vendor Reliability & Spend-Justification Scorecard: six weighted signals scored to 0–100 and a reliability-floor gate demoting a 75-point critical vendor to DO NOT RENEW.

The gate this post refers to, drawn from the tool’s own logic. See the tool.