Gemini models and capabilities

The Gemini family in 2026 spans three generations on the Developer API plus the Vertex-only enterprise mirror: the current Gemini 3.x generation (3.5 Flash GA + 3.1 Pro preview + 3 Flash preview + 3.1 Flash-Lite), the still-recommended Gemini 2.5 generation (Pro / Flash / Flash-Lite), the 2.0 line (Flash + Flash-Lite), and the 1.5 legacy line (Pro / Flash) which is on a deprecation track. The Gemini Ultra brand was retired after the 1.x era; “frontier” capability shifted to Pro at the 2.5 tier and to Flash at the 3.x tier (per ai.google.dev 2026-05, Gemini 3.5 Flash is positioned as “most intelligent model for sustained frontier performance on agentic and coding tasks”). Native modality coverage — text, image, audio, video, PDF, code — is the line’s signature differentiator: every current Gemini model accepts all six.

See also

1. Current generation (Gemini 3.x)

Per ai.google.dev/gemini-api/docs/models and ai.google.dev/gemini-api/docs/pricing as of 2026-05-25.

FeatureGemini 3.5 FlashGemini 3.1 ProGemini 3 FlashGemini 3.1 Flash-Lite
StatusStable / GAPreviewPreviewStable + Preview tracks
API IDmodels/gemini-3.5-flashmodels/gemini-3.1-promodels/gemini-3-flashmodels/gemini-3.1-flash-lite
Vertex AI IDgemini-3.5-flashgemini-3.1-progemini-3-flashgemini-3.1-flash-lite
Input price (USD / MTok)$1.50(preview pricing varies)(preview pricing varies)$0.25
Output price (USD / MTok)$9.00(preview pricing varies)(preview pricing varies)$1.50
Context window1,000,000 tokens1,000,000+ tokens1,000,000 tokens1,000,000 tokens
Max output tokens≥ 65,536≥ 65,536≥ 65,536≥ 65,536
Modalities (input)text + image + audio + video + PDFsame + extended reasoningsamesame
Modalities (output)text + optional structured JSONsamesamesame
Thinking controlthinkingLevel (minimal / low / medium / high)thinkingLevel defaults to highthinkingLevelthinkingLevel
Cachingimplicit + explicit; min 1,024 tokensimplicit + explicitimplicit + explicitimplicit + explicit
Function callingunique id per call required (Gemini 3 change — per ai.google.dev/gemini-api/docs/function-calling)samesamesame
Knowledge cutoff(observed late-2025)(observed late-2025)(observed late-2025)(observed late-2025)

Selection heuristic (per the docs’ positioning):

  • 3.5 Flash — current general-purpose default. Best price/performance on agentic coding, tool use, structured tasks. Replace Gemini 2.5 Flash for most workloads.
  • 3.1 Pro (preview) — for the hardest reasoning tasks where thinkingLevel: high is needed and 3.5 Flash isn’t enough.
  • 3.1 Flash-Lite — high-throughput, low-cost endpoint for classification, extraction, simple Q&A. Replaces Gemini 2.5 Flash-Lite for most workloads.

Per ai.google.dev/gemini-api/docs/models and …/pricing. All stable / GA as of 2026-05-25; many production deployments are still on 2.5 because the API surface is identical and the 3.x generation hadn’t fully GA’d.

FeatureGemini 2.5 ProGemini 2.5 FlashGemini 2.5 Flash-Lite
API IDmodels/gemini-2.5-promodels/gemini-2.5-flashmodels/gemini-2.5-flash-lite
Vertex AI IDgemini-2.5-progemini-2.5-flashgemini-2.5-flash-lite
Input price (USD / MTok)$1.25 (≤200k tokens) / $2.50 (>200k)$0.30$0.10
Output price (USD / MTok)$10.00 (≤200k) / $15.00 (>200k)$2.50$0.40
Context window1,048,576 tokens (~1M)1,048,576 tokens1,048,576 tokens
Max output tokens65,53665,53665,536
Thinking controlthinkingBudget 128 - 24,576 (cannot disable)thinkingBudget 0 - 24,576 (set 0 to disable)not supported
Caching minimum4,096 tokens1,024 tokens1,024 tokens
Modalities (input)text + image + audio + video + PDFsametext + image + PDF (no audio / video)
Free-tier rate5 RPM (limited)higherhigher
Knowledge cutoff (observed)Jan 2025Jan 2025Jan 2025

Important pricing notch: Gemini 2.5 Pro has a >200k-token surcharge — request input over the 200,000-token watermark doubles to $2.50 / MTok input and $15 / MTok output. The watermark is on the input size, not the context cap. Workloads near 200k should consider chunking + caching to stay under, or accept the surcharge as an explicit budget item.

3. Gemini 2.0 generation (deprecating)

FeatureGemini 2.0 FlashGemini 2.0 Flash-Lite
API IDmodels/gemini-2.0-flashmodels/gemini-2.0-flash-lite
StatusStable, deprecating in favor of 2.5 FlashStable, deprecating in favor of 2.5 Flash-Lite
Input / Output(legacy pricing, see archived docs)(legacy)
Context1,048,5761,048,576
Modalitiestext + image + audio + video + PDFtext + image + PDF

Migrate to 2.5 Flash / Flash-Lite (same surface, better behavior). Per ai.google.dev 2026-05, 2.0 Flash receives only critical-fix updates.

4. Gemini 1.5 legacy (retired soon)

FeatureGemini 1.5 ProGemini 1.5 FlashGemini 1.5 Flash-8B
API IDmodels/gemini-1.5-promodels/gemini-1.5-flashmodels/gemini-1.5-flash-8b
Contextup to 2,000,000 tokens (Pro is the only 2M model)1,000,0001,000,000
Modalitiestext + image + audio + video + PDFsamesame
StatusRetiring — migrate to 2.5 ProRetiring — migrate to 2.5 FlashRetiring

2M context survives only on 1.5 Pro; the 2.5 and 3.x generations cap at ~1M. Workloads that genuinely need 2M (e.g. analyzing 20+ hours of video transcripts in a single call, ingesting 1500-page legal corpora) must either stay on 1.5 Pro for now or chunk + use context caching on 2.5/3.x.

5. Multimodal sister models

Distinct model series accessible from the Gemini API surface but not part of the Gemini-family chat completion path:

FamilyModelsPurpose
Nano BananaNano Banana 2, Nano Banana ProNative image generation + editing (replaces / supplements Imagen)
ImagenImagen 4 (and predecessor lines)High-quality image generation; $0.02-$0.06 per image
VeoVeo 3.1, Veo 3.1 LiteVideo generation; $0.05-$0.60 per second
LyriaLyria 3 Pro, Lyria 3 Clip, Lyria RealTimeMusic generation
Live APIGemini 3.1 / 2.5 Flash LiveReal-time bidirectional voice + video
TTSGemini 3.1 / 2.5 Flash TTSText-to-speech variants
AgentsDeep Research, Antigravity Agent, Computer UseHosted agent endpoints

Per ai.google.dev/gemini-api/docs/models 2026-05.

6. Modality matrix

Every Gemini chat-tier model accepts the same six input modalities. The cell content is the model’s per-unit token consumption.

ModalityInput cost (tokens)Notes
Text1 token ≈ 4 charsStandard SentencePiece-like tokenizer
Image (≤ 384 × 384 px)258 tokens flatWhole-image fits a single tile
Image (larger)258 × tile_countImage is split into 768×768 tiles
PDF (per page)~258 tokens / pagePer ai.google.dev/gemini-api/docs/vision
Video (per second)(model-dependent; ~263 tokens / frame for 1-fps sampling)Sampled at native rate by the Files API
Audio (per second)32 tokens / secondPer Gemini docs; ~115k tokens / hour

Practical: a 1,000-page PDF ≈ 258k tokens (fits in 1M context with 700k+ headroom). A 1-hour video ≈ 950k tokens (close to the 1M limit; chunk if longer). 9.5 hours of audio ≈ 1.1M tokens (just over 1M).

7. Pricing details beyond per-token

Free tier

Per ai.google.dev/gemini-api/docs/pricing: limited access to certain models with free input + output tokens, Google AI Studio web access, and standard rate limits. Content from free-tier requests may be used to improve Google products — explicit opt-out requires paid tier. Practical use: prototyping only.

Batch mode

50 % discount on input + output via the Batch API. JSONL input format, 24-hour SLA, max 2 GB per input file. See gemini-api-and-sdks section on batch.

Context caching

Three components:

  1. Implicit cache (automatic on 2.5+) — best-effort, no cost-saving guarantee. Just works.
  2. Explicit cacheclient.caches.create(...), guaranteed cost reduction on hits, with hourly storage fee (~$1.00 per 1M tokens / hour).
  3. Hit pricing — cache reads charged at a reduced per-MTok rate vs base input.

See context-caching-gemini for full mechanics.

Google Search grounding

On Gemini 3.x: per-query billing (each search the model decides to execute is one billable unit). Free tier: 5,000 grounded prompts / month on Gemini 3 models. Paid tier above free: $14 per 1,000 queries.

On Gemini 2.5 and older: per-prompt billing (any request that triggered grounding counted once, regardless of how many sub-searches happened). The 3.x change makes grounding-heavy agentic workflows materially more expensive — budget accordingly.

Code execution

No additional charge for enabling code_execution. The model is billed on:

  • Original prompt → input tokens
  • Generated code + execution result → intermediate tokens (billed as output)
  • Summary / final answer → output tokens

Flex inference (where supported)

50 % cost reduction with relaxed SLA. Useful for non-interactive workloads (data processing, periodic summarization) where 10-15 minute response is acceptable.

8. Benchmarks

Per Google DeepMind technical reports + Gemini 3 launch materials (2026-Q1). Treat as directional; vendor self-reports.

BenchmarkGemini 3.5 FlashGemini 2.5 ProGemini 2.5 FlashNotes
MMLUhigh 80smid 90slow 90sKnowledge breadth
GPQA Diamondhigh 70shigh 80smid 70sGraduate STEM
SWE-bench Verifiedstrong (positioned for agentic coding)mid rangelowerSoftware engineering
MMMU (multimodal)high 70shigh 70smid 70sMultimodal reasoning
LiveCodeBenchleaderstrongmidLive competitive programming

Per ai.google.dev model cards 2026-05; the DeepMind Gemini technical report carries audited numbers.

9. Endpoint types

EndpointHostname patternAuthBest for
Gemini Developer API (AI Studio)generativelanguage.googleapis.comAPI key (x-goog-api-key)Prototyping, indie, sub-enterprise
Vertex AI (enterprise)aiplatform.googleapis.com (regional)GCP IAM service-account / ADCProduction, compliance, private VPC, audit logs
Vertex AI Express ModeregionalAPI key (simpler, no IAM)On-ramp from Developer API to Vertex without full IAM setup

The same model IDs work on both, but Vertex IDs typically drop the models/ prefix (e.g. gemini-2.5-pro vs models/gemini-2.5-pro). See vertex-ai-enterprise-gemini for the full delta.

10. Programmatic model discovery

from google import genai
client = genai.Client(api_key="...")
for m in client.models.list():
    print(m.name, m.display_name, m.input_token_limit, m.output_token_limit,
          m.supported_actions)

REST:

curl "https://generativelanguage.googleapis.com/v1beta/models?key=$GEMINI_API_KEY"

The supported_actions field tells you which methods the model accepts (generateContent, streamGenerateContent, embedContent, countTokens, batchEmbedContents, cachedContents, etc.). Use this for capability-aware routing rather than hardcoding lists.

11. Migration guide

Quick reference for choosing the migration target:

FromRecommended targetWhy
Gemini 1.5 ProGemini 2.5 ProSame context (1M is enough for most cases); much better quality + price; same API surface. Stay on 1.5 Pro only if you genuinely need 2M context.
Gemini 1.5 FlashGemini 2.5 FlashStrict upgrade
Gemini 2.0 FlashGemini 2.5 FlashStrict upgrade; same modality matrix
Gemini 2.5 FlashGemini 3.5 FlashBetter agentic + coding; 3.5 has thinkingLevel semantics — adapt any code that sets thinkingBudget directly
Gemini 2.5 ProGemini 3.1 Pro (when preview lifts)Wait for GA unless preview is fine for your use case

12. Further reading