Models & Pricing - DeepSeek
The prices listed below are in units of per 1M tokens. A token, the smallest unit of text that the model recognizes, can be a word, a number, or even a punctuation mark.

- The prices listed below are in units of per 1M tokens.
- The legacy names deepseek-v4-flash and deepseek-v4-flash-vision-exp are still accepted, but the corresponding models have been retired, their requests are served by the DeepSeek-V4.1-Flash model and billed at the Flash price.
- (2) Off-peak rates are half of the peak rates.
The prices listed below are in units of per 1M tokens. A token, the smallest unit of text that the model recognizes, can be a word, a number, or even a punctuation mark. We will bill based on the total number of input and output tokens by the model. MODEL deepseek-flash (1) deepseek-v4-pro BASE URL (OpenAI Format) https://api.deepseek.com BASE URL (Anthropic Format) https://api.deepseek.com/anthropic MODEL VERSION DeepSeek-V4.1-Flash DeepSeek-V4-Pro-0813 THINKING MODE Supports both non-thinking and thinking (default) modes See Thinking Mode for how to switch CONTEXT LENGTH 1M MAX OUTPUT MAXIMUM: 384K FEATURES Json Output ✓ ✓ Tool Calls ✓ ✓ Responses API ✓ ✓ Anthropic API ✓ ✓ Chat Prefix Completion(Beta) ✓ ✓ FIM Completion(Beta) Non-thinking mode only Non-thinking mode only Vision ✓ Not supported PRICING (2) 1M INPUT TOKENS (CACHE HIT) OFF-PEAK $0.003 $0.022 PEAK $0.006 $0.044 1M INPUT TOKENS (CACHE MISS) OFF-PEAK $0.15 $0.66 PEAK $0.3 $1.32 1M OUTPUT TOKENS OFF-PEAK $0.6 $1.98 PEAK $1.2 $3.96 Concurrency Limit (3) 2500 500 (1) Use deepseek-flash as the model name. The legacy names deepseek-v4-flash and deepseek-v4-flash-vision-exp are still accepted, but the corresponding models have been retired, their requests are served by the DeepSeek-V4.1-Flash model and billed at the Flash price. (2) Off-peak rates are half of the peak rates. Peak hours are 01:00 - 04:00 and 06:00 - 10:00 UTC, Monday through Friday, excluding Chinese public holidays. All other hours are off-peak, including weekends and Chinese public holidays in full. (3) For more details on concurrency limits, please refer to Rate Limit & Isolation . The corresponding fees will be directly deducted from your topped-up balance or granted balance, with a preference for using the granted balance first when both balances are available. Product prices may vary and DeepSeek reserves the right to adjust them. We recommend topping up based on your actual usage and regularly checking this page for the most recent pricing information.
Sources
Related stories

The pacing era's first launch day
PLUS: Build, test, and publish an app without leaving Codex Good morning, AI enthusiasts, and welcome to the 20,221 new readers who joined us yesterday. Barely two weeks into the AI industry’s new “pacing” era, the frontier’s two biggest rivals are still launching on each other’s heels.

Apparently, OpenAI isn't trying to build "magic intelligence in the sky" anymore
Apparently, OpenAI isn't trying to build "magic intelligence in the sky" anymore OpenAI CEO Sam Altman is pushing back against religious analogies tied to AI models.

How NVIDIA GPUs Help Accelerate OpenAI’s GPT-6 Astra Ultrafast
GPT-6 Astra Ultrafast, running on NVIDIA Blackwell GPUs, is available now in the OpenAI API and to eligible ChatGPT Work and Codex users.

Last Week in AI #345 - 5 new models, 9 misalignment incidents, some Dots
Top NewsAnthropic and OpenAI race to release smarter and cheaper modelsSources:Anthropic launches Claude Opus 5.5 with stricter safeguards for…