gpt-5.6-luna vs Gemini 3.6 Flash
The usual choice for high-volume work, where a small difference per million tokens becomes a large difference per month.
Prices in USD per 1M tokens, standard tier, updated 2026-08-11.
gpt-5.6-luna
- Input / 1M
- $0.200
- Cached input / 1M
- $0.020
- Output / 1M
- $1.20
- Context window
- 1.1M
- Source
- First-party
Gemini 3.6 Flash
- Input / 1M
- $1.50
- Cached input / 1M
- $0.150
- Output / 1M
- $7.50
- Context window
- 1M
- Source
- First-party
Which is cheaper depends on the workload
Estimated monthly cost for three common request shapes. This is the part a price list cannot answer.
| Workload | gpt-5.6-luna | Gemini 3.6 Flash | Cheaper |
|---|---|---|---|
| Chat assistant1,500 in / 600 out × 100,000/mo | $102.00 | $675.00 | gpt-5.6-luna · 6.6× cheaper |
| RAG / document Q&A20,000 in / 500 out × 30,000/mo | $138.00 | $1.0K | gpt-5.6-luna · 7.3× cheaper |
| Coding agent30,000 in / 4,000 out × 20,000/mo | $216.00 | $1.5K | gpt-5.6-luna · 6.9× cheaper |
Excludes caching discounts, so real bills with repeated prompts are lower. Price your own workload →
Common questions
- Is gpt-5.6-luna or Gemini 3.6 Flash cheaper?
- It depends on the shape of your requests. gpt-5.6-luna has the lower input price ($0.200 per 1M). gpt-5.6-luna has the lower output price ($1.20 per 1M). Because output usually costs several times more than input, the cheaper choice flips depending on how much text you generate.
- What is the context window on gpt-5.6-luna and Gemini 3.6 Flash?
- gpt-5.6-luna accepts 1,050,000 tokens and Gemini 3.6 Flash accepts 1,048,576 tokens. gpt-5.6-luna holds more.
- Do these prices include caching discounts?
- The table shows standard rates. gpt-5.6-luna bills cached input at $0.020 per 1M and Gemini 3.6 Flash at $0.150. If you resend the same system prompt or document, that discount usually matters more than the headline difference.