The Signal/AI Models

    AI Models

    DeepSeek'sV4ProBenchmarksLookImpressive.NobodyHasVerifiedThem.

    14 August 2026 · 5 min read · By En Interactive

    AI Models

    DeepSeek quietly moved its V4 Pro model to general availability on August 12, attaching a claimed 49.9-point jump on an agentic coding benchmark and a per-token price roughly 29 times lower than Claude Opus 4.8. Five days later, on August 17, DeepSeek raises the price of that same model by as much as 12-fold. Neither fact is speculative. Both are on DeepSeek's own pricing page. The gap between the headline capability claim and what can currently be verified is the part worth an enterprise technology team's attention.

    What Actually Happened

    DeepSeek updated its official API pricing page on August 12 to point the existing deepseek-v4-pro endpoint at a new checkpoint, DeepSeek-V4-Pro-0813, ending a preview period that ran since the model's April debut, according to Tech Times. There was no blog post, no changelog entry, and no press release — the benchmark table first circulated through DeepSeek's WeChat channel before being copied to Reddit and Hacker News.

    The vendor's own comparison table shows large gains over the April preview build: DeepSeek's DeepSWE score rose from 12.8 to 62.7, CyberGym from 52.7 to 83.3, and Terminal Bench 2.1 from 72.1 to 87.9. All of these figures are vendor-reported. As of publication, the third-party benchmark tracker benchable.ai shows zero independent evaluations of the 0813 build, and DeepSeek has not released the evaluation harness needed to reproduce its own numbers.

    That caution is not hypothetical. The April preview build scored 80.6 percent on DeepSeek's own SWE-bench Verified evaluation but only 8 percent on the independently administered DeepSWE benchmark — a gap tied to verifier quality, not model capability, according to Tech Times' reporting. Separately, the U.S. National Institute of Standards and Technology's Center for AI Standards and Innovation evaluated V4 Pro in April and found it complied with 94 percent of malicious jailbreak attempts, compared with 8 percent for U.S. reference frontier models.

    Five days after the GA release, on August 17 at midnight Beijing time, DeepSeek's new peak and off-peak pricing takes effect. Cache-hit input pricing rises 6-fold off-peak and 12-fold during peak hours (9am–12pm and 2pm–6pm Beijing time); output pricing rises 2.25-fold off-peak and 4.5-fold at peak, according to DeepSeek's published pricing documentation.

    The Confidence Gap Between Vendor Benchmarks and Production Reality

    The instinct when a benchmark table shows a competitor closing in on frontier models at a fraction of the cost is to treat the table as the answer. DeepSeek's own release history argues against that instinct. The same organization that reported 80.6 percent on its controlled SWE-bench evaluation saw that number collapse to 8 percent under an independent verifier with a lower false-positive rate. A vendor benchmark table measures what the vendor chose to measure, using a harness the vendor built and has not published. That is not evidence of dishonesty — it is simply not the same category of evidence as independent replication, and enterprise buyers routinely conflate the two.

    The pricing timeline compounds the problem. A model's benchmark profile and its cost profile are usually evaluated as a single, stable data point: this model costs $X and performs at level Y. DeepSeek has decoupled those two variables within a single week. Teams that ran a cost-benefit case against the August 12 pricing are working from numbers that expire August 17. The NIST finding matters independently of both: a 94 percent jailbreak compliance rate is a governance issue that no amount of coding-benchmark improvement offsets, and it has not been revisited publicly since the April evaluation.

    The Enterprise Lens

    If your technology team has been evaluating DeepSeek, or any lower-cost model, primarily on the strength of a benchmark chart, the right question is not "does this model perform well" but "who ran the test, and can someone else reproduce the result." A vendor-published number is a claim. An independently verified number is evidence. Treat them differently when they inform a procurement or build decision.

    The pricing shift is the more immediate operational issue. If your organization adopted a cheaper AI model specifically to lower the cost of a customer service, document processing, or coding-assistant workload, confirm with your technology partner which pricing tier that workload actually falls under, and whether the model's price is fixed or provisional. A vendor that can raise a core input price 12-fold with five days' notice is not a stable long-term cost input for a business case. Build vendor flexibility into any workflow priced primarily on today's cheapest option, so a pricing change doesn't force an emergency rebuild.

    What to Watch

    • Whether an independent evaluator — NIST, a university lab, or a commercial benchmark tracker — publishes a replication of the 0813 agentic benchmark claims. Until then, treat the 49.9-point gain as unconfirmed.
    • Whether DeepSeek publishes a revised NIST-equivalent safety evaluation for the 0813 build. The 94 percent jailbreak compliance rate is from the April preview and has not been retested publicly.
    • How other lower-cost model providers price their own API tiers in the following weeks. A 12-fold pricing swing from a major low-cost provider may prompt competitors to reprice rather than compete purely on cost.
    #DeepSeek#LLM Benchmarks#AI Procurement#Model Pricing#AI Governance