Skip to content

DeepSeek V4 Pro Officially Launched

Aug 13, 00:12
DeepSeek V4 Pro Officially Launched
Original Article Title: "DeepSeek V4 Pro Officially Launched"
Original Article Author: InsightBeating


DeepSeek V4 Pro is here.


On the evening of August 12, the DeepSeek official API documentation has updated the model version corresponding to deepseek-v4-pro to DeepSeek-V4-Pro-0813. Developers can now directly call the new version through the API, with a price of ¥3 per million Token input, ¥6 per million Token output, and a cache hit input price of only ¥0.025 per million Token.


V4 Pro-0813 continues to support a 1 million Token context window, with a maximum output of 384,000 Tokens and supports both deliberative and non-deliberative modes. Abilities such as JSON Output and Tool Calls have also been opened up.


As of the evening of August 12, the DeepSeek official update log has not yet separately released an announcement for the V4 Pro official launch. However, during the previous V4 Flash official release, DeepSeek clearly stated that the V4 Pro official version would be released subsequently; and now the API backend model node has been switched from the previous version to DeepSeek-V4-Pro-0813.


It has been nearly four months since the first appearance of V4.


On April 24 of this year, DeepSeek released the V4 Preview and introduced V4 Pro and V4 Flash. Among them, V4 Pro is the flagship version of the entire series, with a total parameter size of 1.6T, activating about 49B parameters per Token, and supporting a 1 million Token context. This generation of V4 significantly enhanced coding, tool calling, and long-term Agent tasks from the beginning.


On July 31, DeepSeek launched V4-Flash-0731. The official statement mentioned that this version maintains the original model architecture and size, mainly undergoing retraining, significantly strengthening Agent capabilities, and natively integrating Responses API, while also adapting to Codex. At that time, DeepSeek specifically emphasized that this update only involved Flash, with no changes to the Pro and web-based models, and the V4 Pro official version would be released subsequently.


Now, the final piece of the puzzle has been added to Pro.



With this generation of DeepSeek, the focus is increasingly on Agent.


The timing of the V4 Pro official release coincides with a new wave of product emphasis in the overall large model industry.


When model manufacturers unveiled new products this year, the emphasis on simple comparisons of knowledge QA, math, and traditional benchmarks has been diminishing. Metrics such as Coding, Computer Use, Tool Use, long-task completion rates, and the cost of completing a task are becoming the new core indicators.


During the July launch of GPT-5.6, OpenAI devoted a significant amount of space to discussing Agent and Coding capabilities. GPT-5.6 Sol introduced support for a new max reasoning effort, further unveiling an ultra mode where multiple Agents work in parallel to accomplish complex tasks. The Responses API also began allowing models to directly write programs to coordinate tools, process intermediate results, and then determine the next steps. OpenAI even made "how much work can be done per dollar" one of the main selling points of this generation of products.


Anthropic has followed a similar path.


With the late-June release of Claude Sonnet 5, the main enhancements were also focused on Coding, Agent, Computer Use, and long tasks. Anthropic offers different effort levels, allowing developers to choose inference intensity based on the task; by August 10, Anthropic announced that the Sonnet 5's original promotional prices of $2 per million inputs and $10 per million outputs will be permanently retained.


Model companies are now facing more than just who is smarter.


An Agent may run continuously for tens of minutes or even hours, repeatedly reading and writing code, searching for information, invoking browsers and external tools. Once truly in a production environment, the model invocation rate will quickly increase. Capability, speed, Token consumption, and invocation price are now collectively determining whether a product's Agent can run.


The design of DeepSeek V4 is also moving in this direction.


With 1 million Token contexts, the model can read in a larger codebase, corporate knowledge base, and task history in one go; Tool Calls handle interaction with the external environment; compatibility with Anthropic API and the Responses API already supported by Flash are aimed at lowering the barrier for models to integrate into existing Agent toolchains. When DeepSeek released V4 in April, it had simultaneously opened up interfaces for OpenAI ChatCompletions and Anthropic formats.


This is also where V4 Pro's official version is more noteworthy than just another larger DeepSeek model. DeepSeek is further turning the model into infrastructure that Agents can directly invoke.


After the launch of V4 Pro, the hierarchy of this generation of DeepSeek products becomes clearer.


Currently, the V4 Flash input price is 1 yuan per million tokens, while Pro is 3 yuan; Flash output is 2 yuan, and Pro is 6 yuan. The prices happen to differ by exactly three times. Both V4 Pro and Flash support 1 million contexts, but Pro uses a larger set of activation parameters, focusing more on complex reasoning, coding, and more challenging Agent tasks.


Flash handles a large amount of high-frequency, cost-sensitive tasks, leaving the heavier work to Pro.


In the 2025 large model competition, discussions often revolve around "which flagship model scores first." By 2026, models are increasingly resembling true software infrastructure, requiring both intelligence and a stable API, long contexts, tool invocation, sufficiently high concurrency, and a price point that developers can afford.


After the launch of V4 Pro's official version, the product profile of this generation DeepSeek V4 is finally complete.


Original Article Link


Recommended

After Suspending 8 ETFs and Downsizing by 14%, Bitwise Forges Ahead with New Product Launch

Aug 12, 19:06
After Suspending 8 ETFs and Downsizing by 14%, Bitwise Forges Ahead with New Product Launch

Will Tonight's U.S. CPI Data Crush Expectations for a September Rate Hike?

Aug 12, 17:49
Will Tonight's U.S. CPI Data Crush Expectations for a September Rate Hike?

CoreWeave Delivers Record-Breaking Financial Report, Why is Bernstein Still Bearish?

Aug 12, 17:30
CoreWeave Delivers Record-Breaking Financial Report, Why is Bernstein Still Bearish?

Manus and Lin Junyang, Collective Return

Aug 12, 14:38
Manus and Lin Junyang, Collective Return

StonkBroker is undoubtedly the Robinhood of the Metaverse, bro.

Aug 12, 12:18
StonkBroker is undoubtedly the Robinhood of the Metaverse, bro.

Arthur Hayes New Article: Betting on Yen Appreciation, ENA Could Surge 5 to 10 Times in the Coming Months

Aug 12, 11:29
Arthur Hayes New Article: Betting on Yen Appreciation, ENA Could Surge 5 to 10 Times in the Coming Months