Google

Gemini 2.5 Flash Lite

1.05M ctxTextImage Tools

Released Jul 22, 2025 · 129B tokens this week · 2 providers

Overview

The cheapest Gemini tier, optimised for very high request volumes. Best for routing, tagging and short answers where cost per call dominates the design. Keeps the long context window of its bigger siblings.

translationmarketingagents

Providers

Reference data — not yet measured from live traffic

PROVIDERMAX OUTOUTPUT /MQUANT
Google AI StudioCheapest
1.05M66K$0.10$0.40230ms34099.94%Full
Google Vertex AI
1.05M66K$0.11$0.42270ms31099.86%Full

Latency p50 reflects a provider’s API endpoint responsiveness — the round-trip to its API, not per-token inference time. These figures, with throughput and uptime, are reference data until the gateway aggregates its own traffic.

Price comparison

Blended $ per 1M tokens

Gemini 2.5 Flash Lite$0.33
GPT-5.2$10.94

Blended rate per 1M tokens, weighted one part prompt to three parts completion — roughly the shape of a chat workload. Your mix will move the number.

Call it

OpenAI-compatible — swap the base URL and go

curl https://model.cards/api/v1/chat/completions \
  -H "Authorization: Bearer $MODELCARDS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "google/gemini-2.5-flash-lite",
    "messages": [
      { "role": "user", "content": "Summarize the tradeoffs of speculative decoding." }
    ]
  }'