DeepSeek Unveils DeepSeek-V4-Flash-0731: High-Performance Compact Model Matching Claude Sonnet 5 Standards at Ultra-Low API CostAI research firm DeepSeek has officially launched DeepSeek-V4-Flash-0731, an updated iteration of its Flash model series originally introduced in April. Retaining its core architectural design and parameter scale, the refreshed model maintains DeepSeek's signature ultra-low API pricing structure while delivering substantial benchmark improvements.
Despite operating on a compactMoE (Mixture-of-Experts) framework with a 283B total / 13B active parameter profile (283B-A13B), internal test scores released by DeepSeek demonstrate exceptional capabilities in complex software engineering and terminal execution tasks:
Terminal Bench: Achieved a score of 82.7.
DeepSWE: Reached 54.4, outperforming the previously released DeepSeek V4 Pro and placing its performance on par with Claude Sonnet 5 (max).
API pricing for DeepSeek-V4-Flash remains locked at $0.14 per million input tokens and $0.28 per million output tokens. While direct access is currently limited to API endpoints, DeepSeek has fully released the model weights under the permissive MIT License, enabling third-party cloud providers and open-source hosts to deploy and serve the model globally.
The way a low-parameter Mixture-of-Experts (MoE) model can achieve cutting-edge intelligence—by enabling just 13 billion parameters out of 283 billion for any given token—DeepSeek-V4-Flash achieves the inference speed and low cost of a small-scale model while maintaining the deep contextual reasoning capabilities of larger underlying models.
The price comparison of $0.14/0.28 per million tokens to proprietary models underscores why open models are transforming software development workflows. At these prices, engineering teams can directly integrate real-time code completion, automated request pool monitoring, and automated agent loops into ongoing integration/collaboration (CI/CD) pipelines without the high overhead of cloud APIs.
Commercial freedom through the MIT licensing underscores the importance of the MIT license, which, unlike open licenses that restrict commercial revenue or prohibit competitive customization, allows startups, enterprises, and hyperscaler providers (such as AWS, Together AI, or Ollama) to freely host, modify, and monetize DeepSeek-V4-Flash. This will help accelerate deployment across organizations in regional data centers.
Source: @deepseek_ai
DeepSeek Unveils DeepSeek-V4-Flash-0731: High-Performance Compact Model Matching Claude Sonnet 5 Standards at Ultra-Low API CostAI research firm DeepSeek has officially launched DeepSeek-V4-Flash-0731, an updated iteration of its Flash model series originally introduced in April. Retaining its core architectural design and parameter scale, the refreshed model maintains DeepSeek's signature ultra-low API pricing structure while delivering substantial benchmark improvements.
Despite operating on a compactMoE (Mixture-of-Experts) framework with a 283B total / 13B active parameter profile (283B-A13B), internal test scores released by DeepSeek demonstrate exceptional capabilities in complex software engineering and terminal execution tasks:
Terminal Bench: Achieved a score of 82.7.
DeepSWE: Reached 54.4, outperforming the previously released DeepSeek V4 Pro and placing its performance on par with Claude Sonnet 5 (max).
API pricing for DeepSeek-V4-Flash remains locked at $0.14 per million input tokens and $0.28 per million output tokens. While direct access is currently limited to API endpoints, DeepSeek has fully released the model weights under the permissive MIT License, enabling third-party cloud providers and open-source hosts to deploy and serve the model globally.
The way a low-parameter Mixture-of-Experts (MoE) model can achieve cutting-edge intelligence—by enabling just 13 billion parameters out of 283 billion for any given token—DeepSeek-V4-Flash achieves the inference speed and low cost of a small-scale model while maintaining the deep contextual reasoning capabilities of larger underlying models.
The price comparison of $0.14/0.28 per million tokens to proprietary models underscores why open models are transforming software development workflows. At these prices, engineering teams can directly integrate real-time code completion, automated request pool monitoring, and automated agent loops into ongoing integration/collaboration (CI/CD) pipelines without the high overhead of cloud APIs.
Commercial freedom through the MIT licensing underscores the importance of the MIT license, which, unlike open licenses that restrict commercial revenue or prohibit competitive customization, allows startups, enterprises, and hyperscaler providers (such as AWS, Together AI, or Ollama) to freely host, modify, and monetize DeepSeek-V4-Flash. This will help accelerate deployment across organizations in regional data centers.
Source: @deepseek_ai
Comments
Post a Comment