GLM 5.3 Flash pricing & price history
GLM 5.3 Flash by Z.ai (Zhipu) costs $0.15 per 1M input tokens and $0.50 per 1M output tokens, with a 1.3M token context window. The cheapest provider currently charges $0.045 input / $0.14 output.
Price history
No price changes since tracking began on 2026-09-24. We check hourly.
Providers serving GLM 5.3 Flash (30)
| Provider | Input | Output | Context | Uptime 30m | Uptime 24h |
|---|---|---|---|---|---|
| InferenceNet fp4 | $0.045 | $0.14 | 1M | 99.0% | 98.9% |
| Relace | $0.07 | $0.28 | 1M | 99.8% | 99.9% |
| DeepInfra fp4 | $0.075 | $0.25 | 1M | 99.7% | 99.4% |
| Morph | $0.088 | $0.31 | 1M | 78.3% | 79.5% |
| Wafer | $0.089 | $0.35 | 1M | 99.8% | 99.9% |
| Inceptron fp8 | $0.09 | $0.28 | 1M | 99.8% | 98.8% |
| GMICloud fp8 | $0.09 | $0.30 | 1M | 99.6% | 99.1% |
| OpenInference fp4 | $0.10 | $0.50 | 1M | 99.9% | 99.9% |
| Io Net fp8 | $0.13 | $0.42 | 262K | 99.4% | 99.7% |
| Phala fp8 | $0.13 | $0.42 | 1M | 99.2% | 99.0% |
| Novita fp8 | $0.13 | $0.44 | 1M | 99.4% | 99.5% |
| StreamLake fp8 | $0.14 | $0.47 | 1M | 99.5% | 99.4% |
| Sail Research fp8 | $0.14 | $0.47 | 1M | 100.0% | 100.0% |
| BaseTen fp8 | $0.15 | $0.50 | 1M | 99.1% | 98.2% |
| Near AI fp8 | $0.15 | $0.50 | 1M | 99.9% | 99.3% |
| Crusoe fp4 | $0.15 | $0.50 | 1M | 96.3% | 92.7% |
| CoreWeave nvfp4 | $0.15 | $0.50 | 1M | 100.0% | 99.9% |
| AtlasCloud fp8 | $0.15 | $0.50 | 1M | 99.7% | 99.5% |
| Fireworks | $0.15 | $0.50 | 1M | 94.1% | 99.5% |
| Friendli | $0.15 | $0.50 | 1M | 99.5% | 99.0% |
| SiliconFlow fp8 | $0.15 | $0.50 | 1M | 99.4% | 99.6% |
| DigitalOcean | $0.15 | $0.50 | 1M | 99.2% | 97.0% |
| Together | $0.15 | $0.50 | 1M | 99.8% | 99.9% |
| Reka fp8 | $0.15 | $0.50 | 262K | 99.9% | 97.9% |
| Parasail fp8 | $0.15 | $0.50 | 1M | 99.2% | 99.1% |
| BaseTen fp8 | $0.15 | $0.50 | 1M | 99.0% | 98.6% |
| Venice | $0.15 | $0.50 | 1M | 99.5% | 98.8% |
| Z.AI fp8 | $0.15 | $0.50 | 1M | 99.5% | 99.7% |
| NextBit fp8 | $0.17 | $0.55 | 1M | 99.7% | 99.4% |
| Cloudflare | $0.30 | $1.00 | 1.3M | 100.0% | 93.1% |
Provider availability updated 54m ago.
Cost calculator
Cheaper alternatives to GLM 5.3 Flash
| Model | Input | Output | Context | Intel. | |
|---|---|---|---|---|---|
O GPT-6 Luna OpenAI | $0.10 | $0.50 | 1.1M | 37.3 | compare |
D DeepSeek V4.1 Flash DeepSeek | $0.14 | $0.42 | 1M | 39.5 | compare |
About GLM 5.3 Flash
GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while...
FAQ
How much does GLM 5.3 Flash cost?
GLM 5.3 Flash by Z.ai (Zhipu) costs $0.15 per 1M input tokens and $0.50 per 1M output tokens, with a 1.3M token context window. The cheapest provider currently charges $0.045 input / $0.14 output. A typical request with 2,000 input and 500 output tokens costs about $0.0006.
What is the context window of GLM 5.3 Flash?
GLM 5.3 Flash supports up to 1,310,720 tokens of context and up to 943,718 output tokens per request.
Has GLM 5.3 Flash’s price changed?
No price changes recorded since ModelBank started tracking it on 2026-09-24.
Is GLM 5.3 Flash being deprecated?
No retirement date has been announced for GLM 5.3 Flash as of 2026-09-24.
What are cheaper alternatives to GLM 5.3 Flash?
Cheaper options include GPT-6 Luna ($0.10/$0.50), DeepSeek V4.1 Flash ($0.14/$0.42).