DeepSeek V4.1 Flash pricing & price history
DeepSeek V4.1 Flash by DeepSeek costs $0.14 per 1M input tokens and $0.42 per 1M output tokens, with a 1M token context window. The cheapest provider currently charges $0.05 input / $0.25 output.
Price history
No price changes since tracking began on 2026-09-24. We check hourly.
Providers serving DeepSeek V4.1 Flash (26)
| Provider | Input | Output | Context | Uptime 30m | Uptime 24h |
|---|---|---|---|---|---|
| Relace | $0.05 | $0.25 | 1M | 99.9% | 100.0% |
| Morph | $0.072 | $0.29 | 1M | 100.0% | 99.8% |
| OpenInference fp4 | $0.10 | $0.50 | 1M | 100.0% | 99.2% |
| DekaLLM | $0.10 | $1.00 | 1M | 99.9% | 95.4% |
| Sail Research fp4 | $0.13 | $0.75 | 1M | 99.9% | 99.8% |
| DeepInfra fp8 | $0.14 | $0.42 | 1M | 99.8% | 99.7% |
| DeepSeek | $0.15 | $0.60 | 1M | 100.0% | 100.0% |
| StreamLake fp8 | $0.20 | $0.79 | 1M | 99.7% | 99.8% |
| Wafer | $0.20 | $0.60 | 1M | 100.0% | 100.0% |
| CoreWeave fp8 | $0.20 | $0.65 | 1M | 99.8% | 99.5% |
| Fireworks | $0.22 | $0.66 | 1M | 99.8% | 99.9% |
| Krea fp8 | $0.23 | $0.90 | 1M | 97.2% | 96.3% |
| GMICloud fp8 | $0.23 | $0.90 | 1M | 98.8% | 99.3% |
| Phala | $0.28 | $1.10 | 1M | 100.0% | 99.1% |
| Novita fp8 | $0.28 | $1.14 | 1M | 100.0% | 99.9% |
| AtlasCloud fp8 | $0.30 | $1.20 | 1M | 99.7% | 99.8% |
| BaseTen fp8 | $0.30 | $1.20 | 1M | 100.0% | 99.2% |
| Makora fp8 | $0.30 | $1.20 | 1M | 98.0% | 98.6% |
| DigitalOcean | $0.30 | $1.20 | 1M | 99.8% | 98.6% |
| Alibaba | $0.30 | $1.20 | 1M | 100.0% | 99.9% |
| Together | $0.30 | $1.20 | 1M | 99.9% | 99.9% |
| SiliconFlow fp8 | $0.30 | $1.20 | 1M | 99.9% | 99.2% |
| Modal | $0.30 | $1.20 | 1M | 99.7% | 99.8% |
| BaseTen fp8 | $0.30 | $1.20 | 1M | 99.5% | 99.0% |
| Parasail fp8 | $0.30 | $1.20 | 1M | 93.9% | 95.6% |
| Venice fp8 | $0.38 | $1.50 | 1M | 99.6% | 98.6% |
Provider availability updated 54m ago.
Cost calculator
Cheaper alternatives to DeepSeek V4.1 Flash
| Model | Input | Output | Context | Intel. | |
|---|---|---|---|---|---|
O GPT-6 Luna OpenAI | $0.10 | $0.50 | 1.1M | 37.3 | compare |
About DeepSeek V4.1 Flash
DeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the first built on the company's Causal Encoder-Decoder (CED) architecture. It activates 8B parameters on input and 16B on...
FAQ
How much does DeepSeek V4.1 Flash cost?
DeepSeek V4.1 Flash by DeepSeek costs $0.14 per 1M input tokens and $0.42 per 1M output tokens, with a 1M token context window. The cheapest provider currently charges $0.05 input / $0.25 output. A typical request with 2,000 input and 500 output tokens costs about $0.0005.
What is the context window of DeepSeek V4.1 Flash?
DeepSeek V4.1 Flash supports up to 1,048,576 tokens of context and up to 131,072 output tokens per request.
Has DeepSeek V4.1 Flash’s price changed?
No price changes recorded since ModelBank started tracking it on 2026-09-24.
Is DeepSeek V4.1 Flash being deprecated?
No retirement date has been announced for DeepSeek V4.1 Flash as of 2026-09-24.
What are cheaper alternatives to DeepSeek V4.1 Flash?
Cheaper options include GPT-6 Luna ($0.10/$0.50).