Gemini 3.8 Flash vs GLM-5.3-FlashX pricing
- For the Chatbot workload, GLM-5.3-FlashX costs 2.6× less than Gemini 3.8 Flash.
- For the RAG workload, GLM-5.3-FlashX costs 2.2× less than Gemini 3.8 Flash.
- For the Coding agent workload, GLM-5.3-FlashX costs 2.0× less than Gemini 3.8 Flash.
Prices
| USD per 1M tokens | Gemini 3.8 Flash | GLM-5.3-FlashX |
|---|---|---|
| Input | $0.75 | $0.37 |
| Output | $3.75 | $1.25 |
| Cache read | $0.075 | $0.075 |
| Cache write | Not published | Free |
| Batch input | $0.375 | Not published |
| Batch output | $1.875 | Not published |
Verified against the provider's official pricing page on Oct 9, 2026. Official pricing page Long-context and separately billed token prices, where shown, are list prices compiled by models.dev (checked Oct 10, 2026, 09:00 UTC).
Z.AI's list price for its own API, as compiled by models.dev and checked Oct 10, 2026, 09:00 UTC. Prices on cloud platforms and resellers can differ. Provider documentation
Monthly cost by workload
| Workload | Gemini 3.8 Flash | GLM-5.3-FlashX |
|---|---|---|
| Chatbot1,000 in / 500 out tokens, 10,000 requests a day | $787.50/mo | $298.50/mo |
| RAG8,000 in / 500 out tokens, 5,000 requests a day | $1,181/mo | $537.75/mo |
| Coding agent50,000 in / 2,000 out tokens, 1,000 requests a day, 80% cached input | $540.00/mo | $276.00/mo |
Cost = input tokens × input price + output tokens × output price, 30 days a month. Cached input uses the cache read price. Cache write fees are not included.
Context and capabilities
| Gemini 3.8 Flash | GLM-5.3-FlashX | |
|---|---|---|
| Context | 1,048,576 tokens | 1,000,000 tokens |
| Max output | 65,536 tokens | 131,072 tokens |
| Input | text, image, video, audio, pdf | text, image, video, pdf |
| Tool calling | Yes | Yes |
| Reasoning | Yes | Yes |
| Structured outputs | Yes | Yes |
Questions
- What is the price difference between Gemini 3.8 Flash and GLM-5.3-FlashX?
- Gemini 3.8 Flash: $0.75 per 1M input tokens and $3.75 per 1M output tokens. GLM-5.3-FlashX: $0.37 per 1M input tokens and $1.25 per 1M output tokens. For the Chatbot workload (1,000 in / 500 out tokens, 10,000 requests a day), that is about $787.50 per month for Gemini 3.8 Flash and $298.50 for GLM-5.3-FlashX.
- Which has the longer context window, Gemini 3.8 Flash or GLM-5.3-FlashX?
- Gemini 3.8 Flash has the longer context window: up to 1,048,576 tokens, compared with 1,000,000 for GLM-5.3-FlashX.
- Do Gemini 3.8 Flash and GLM-5.3-FlashX have prompt caching and batch prices?
- Gemini 3.8 Flash: cached input $0.075 per 1M tokens; batch $0.375 per 1M input tokens and $1.875 per 1M output tokens. GLM-5.3-FlashX: cached input $0.075 per 1M tokens; no batch price published.