Skip to main content
TechnologyApr 25, 2026· 2 min read

DeepSeek Launches V4: Drastically Reduced AI Agent Costs and Top-Tier Performance

The Chinese lab DeepSeek is back in the spotlight with the launch of its new model V4, marking a significant advance in the field of artificial intelligence. After already surprising the industry with the R-1 model, the company now introduces an even more efficient solution, capable of drastically reducing costs associated with AI agents and complex workflows.

Compared to the previous generation, DeepSeek V4 offers significant improvements, especially in terms of computational efficiency. In scenarios with very broad contexts, the model requires about 27% of the computing power and only 10% of the temporary memory compared to V3.2. This translates into a concrete reduction in operational costs: tasks that previously could cost around 10 dollars could now drop to a range between 1.5 and 2.5 dollars.

From a performance perspective, V4 positions itself at the top of the industry, competing with advanced models developed by OpenAI, Google, and Anthropic. In some benchmarks, it is surpassed only by particularly advanced versions like Gemini 3.1 Pro, while in other areas it reaches or exceeds state-of-the-art performance.

Additional Details on DeepSeek's New Model

One of the most relevant elements is the open-weight nature of the model, which theoretically allows it to be run locally on suitable infrastructures, albeit with extremely high hardware requirements. This approach could favor a greater spread of AI, reducing reliance on proprietary services and lowering entry barriers.

Another crucial aspect concerns the hardware used. DeepSeek has adopted a combination of solutions, leveraging both NVIDIA chips and Ascend processors developed by Huawei. The model has been heavily optimized on the latter, marking a significant moment: for the first time, high-level AI performances are competitive even on non-U.S. hardware.

The technical innovations behind V4 include new attention compression techniques, such as hybrid configurations that optimize token usage, and advanced internal information routing systems. These improvements allow the model to handle longer and more complex contexts with greater efficiency. Despite the advancements, the market remains uncertain.