Latest API update: DeepSeek lists V4 Flash public-beta agent scores including Terminal Bench 2.1 at 82.7, NL2Repo at 54.2, Cybergym at 76.7, DeepSWE at 54.4, Toolathlon verified at 70.3, Agent Last Exam at 25.2, Automation Bench Public at 25.1, and DSBench-FullStack at 68.7.
Open model-card rows: Flash Base is listed at 68.3 on MMLU-Pro, 69.5 on HumanEval, and 44.7 on LongBench-V2. Flash Max is listed at 79.0 on SWE Verified and 56.9 on Terminal Bench 2.0.
Same-family read: Pro Base leads Flash Base on MMLU-Pro, HumanEval, and LongBench-V2. Flash becomes interesting when cost, speed, and repeated execution matter more than the hardest reasoning rows.
Your benchmark should measure completed work: accepted outputs, failed-state handling, citation quality, latency, cost, reviewer time, and tool-state reliability. Official scores choose candidates; your harness decides the route.