DeepSeek V4.1 Flash vs Gemini 3.8 Flash API Pricing
Current prices as of September 10, 2026. This comparison separates DeepSeek's time- and cache-dependent rates from Gemini's flat headline rates.
Quick answer: DeepSeek V4.1 Flash has the lower token prices in every published peak/off-peak and cache scenario. Gemini 3.8 Flash is simpler to budget because its headline input and output rates are flat. Price alone does not establish which model completes a coding or agent task more reliably; test both on your own acceptance criteria.
| Model / billing state | Input / 1M | Output / 1M | Operational note |
| DeepSeek V4.1 Flash — off-peak, cache hit | $0.003 | $0.60 | Lowest input rate; requires a reusable cached prefix |
| DeepSeek V4.1 Flash — off-peak, cache miss | $0.15 | $0.60 | Off-peak windows: Mon–Fri 01:00–04:00 and 06:00–10:00 UTC |
| DeepSeek V4.1 Flash — peak, cache hit | $0.006 | $1.20 | Cache discount still applies |
| DeepSeek V4.1 Flash — peak, cache miss | $0.30 | $1.20 | Highest published DeepSeek state |
| Gemini 3.8 Flash — introductory | $0.75 | $3.75 | Flat headline pricing through December 31, 2026; output includes thinking tokens |
Example workload: 10K input and 2K output tokens
These are arithmetic examples for one user-specified workload, not claims about average tokens needed to finish a task.
Cost per request
DeepSeek off-peak, cache hit$0.00123
DeepSeek off-peak, cache miss$0.00270
DeepSeek peak, cache hit$0.00246
DeepSeek peak, cache miss$0.00540
Gemini 3.8 Flash$0.01500
Which pricing model fits which workload?
- Choose DeepSeek for cost-sensitive scheduled or cache-friendly work: batch code analysis, repeated repository context and background agents can exploit the lower off-peak or cache-hit rates.
- Choose Gemini when predictable billing and Google integration matter: the flat rate is easier to forecast, while actual task cost still depends on generated and thinking tokens.
- Do not infer quality from token price: vendor benchmarks are not a substitute for testing completion rate, tool-call accuracy, latency and retries on your own workflow.
DeepSeek V4 Pro migration
According to owner-provided DeepSeek release communication, V4 Pro remains in service until September 14, 2026 at 12:00 Beijing time. After that cutoff, Pro requests route to V4.1 Flash at Flash pricing. Teams should test behavior before the cutoff rather than treating the lower price as proof of identical outputs.
Use the DeepSeek calculator for peak/cache scenarios, the Gemini calculator for Gemini workloads, or compare broader model choices in the low-cost coding API guide.
Archived May 2026 comparison
The following table and scenarios are retained as a historical snapshot and must not be used as current pricing guidance.
| Model |
Input/1M |
Output/1M |
Context |
Notes |
| Gemini 2.5 Flash-Lite |
$0.10 |
$0.40 |
1M |
Cheapest input — but deprecated, shuts down June 1 |
| DeepSeek V4 Flash |
$0.22 |
$0.66 |
1M |
Cheapest output, best value overall |
| DeepSeek V4 Pro |
$0.66 |
$1.98 |
1M |
75% discount through May 31 |
| Gemini 2.5 Pro |
$1.25 |
$10.00 |
1M |
Strong reasoning, thinking model |
| Gemini 3.1 Pro |
$2.00 |
$12.00 |
1M |
Google's latest flagship |
Budget Tier: DeepSeek V4 Flash vs Gemini 2.5 Flash-Lite
On the surface, Gemini 2.5 Flash-Lite wins on input price ($0.10 vs $0.22). But there are important differences:
- Output price: DeepSeek V4 Flash is cheaper ($0.66 vs $0.40) — and chatbots are output-heavy
- Availability: Gemini 2.0 Flash was deprecated and shut down June 1, 2026. DeepSeek V4 Flash is the current model.
- Context window: Both have 1M token context windows
- Max output: DeepSeek V4 Flash supports 384K output tokens vs Gemini's 8K
Verdict: DeepSeek V4 Flash is the better choice for output-heavy workloads. DeepSeek wins on output price ($0.66 vs $0.40), which matters most for chatbots. Gemini 2.5 Flash-Lite is a solid budget option for input-heavy tasks like classification and extraction.
Mid-Tier: DeepSeek V4 Pro vs Gemini 2.5 Pro
For tasks that need stronger reasoning:
- Price: DeepSeek V4 Pro is 2.8x cheaper on input and 11.5x cheaper on output (with the 75% discount)
- Quality: Gemini 2.5 Pro is generally considered stronger on complex reasoning tasks
- Thinking: Gemini 2.5 Pro is a "thinking" model that shows its reasoning process. DeepSeek V4 Pro also supports thinking mode.
- Discount risk: DeepSeek V4 Pro's 75% discount expires May 31. After that, it reverts to $1.74/$3.48 — still cheaper than Gemini 2.5 Pro.
Verdict: DeepSeek V4 Pro for cost-sensitive workloads. Gemini 2.5 Pro when quality matters more than cost.
Real Cost Comparison: 3 Scenarios
Chatbot (10K conversations/month, 500 input + 300 output tokens each)
Monthly cost comparison
DeepSeek V4 Flash$1.05
Gemini 2.5 Flash-Lite$1.20
DeepSeek V4 Pro$4.73
Gemini 2.5 Pro$12.38
Gemini 3.1 Pro$18.60
RAG Pipeline (100K queries/month, 2K input + 500 output tokens each)
Monthly cost comparison
DeepSeek V4 Flash$4.20
Gemini 2.5 Flash-Lite$4.00
DeepSeek V4 Pro$13.15
Gemini 2.5 Pro$37.50
Gemini 3.1 Pro$60.00
Code Assistant (50 devs, 20 requests/day each, 3K input + 1K output tokens)
Monthly cost comparison
DeepSeek V4 Flash$7.35
Gemini 2.5 Flash-Lite$7.50
DeepSeek V4 Pro$25.99
Gemini 2.5 Pro$75.00
Gemini 3.1 Pro$108.00
When to Choose DeepSeek
- Cost is the primary concern — DeepSeek is consistently 3-15x cheaper
- High-volume workloads — the savings compound at scale
- Code generation — DeepSeek V4 Pro excels at coding tasks
- Long outputs — 384K max output vs Gemini's 8K-65K
When to Choose Gemini
- Complex reasoning — Gemini 2.5 Pro's thinking model is stronger on multi-step problems
- Multimodal inputs — Gemini natively handles images, audio, and video
- Google ecosystem — better integration with Google Cloud, Vertex AI
- Reliability — Google's infrastructure has better uptime guarantees
The Deprecation Problem
Google is aggressively deprecating older Gemini models:
- Gemini 3 Pro: Shut down March 9, 2026 (replaced by 3.1 Pro)
- Gemini 2.5 Flash-Lite: Shutting down June 1, 2026
DeepSeek has been more stable — V3 is still available alongside V4. If you're building production systems, model stability matters.
— See if you're overpaying for AI APIs
🎯 API Cost Score
Rate your API setup — get a letter grade in 30 seconds
💸 Looking for DeepSeek Alternatives?
We compared 5 alternatives with better quality, reliability, or compliance at comparable prices.
See 5 DeepSeek Alternatives →
Related Reading
5 Cheaper Gemini Alternatives → Save 17-97%
🎯 Rate Your API Setup in 30 Seconds
Get an A+ to F grade on your AI API costs. See how you compare and find cheaper alternatives instantly.
Get Your Cost Score →
📊 Generate Your Personalized API Cost Report
Select your model, enter your monthly spend, and get a custom savings report with cheaper alternatives — free, in 60 seconds.
Want to optimize your AI API costs?
APIpulse includes free cost comparisons, exports, and recommendations that can save you up to 40%.
Free Cost Audit →
This was a snapshot. What about next month?
Prices change. New models launch. Our tools catch what a one-time calculation can't — and saves you money every month.