OpenAI
GPT-4o Mini
Released Jul 18, 2024 · 214B tokens this week · 2 providers
Overview
The small sibling of GPT-4o, with a very low price and quick responses. Popular for chat widgets, summarisation and bulk data cleanup where frontier quality is not required. Handles images, which is unusual at this price point.
Providers
Reference data — not yet measured from live traffic
| PROVIDER | MAX OUT | OUTPUT /M | QUANT | |||||
|---|---|---|---|---|---|---|---|---|
OpenAICheapest | 128K | 16K | $0.15 | $0.60 | 208ms | 180 | 99.94% | Full |
Azure AI Foundry | 128K | 16K | $0.16 | $0.63 | 380ms | 150 | 99.82% | Full |
Latency p50 reflects a provider’s API endpoint responsiveness — the round-trip to its API, not per-token inference time. These figures, with throughput and uptime, are reference data until the gateway aggregates its own traffic.
Price comparison
Blended $ per 1M tokens
Blended rate per 1M tokens, weighted one part prompt to three parts completion — roughly the shape of a chat workload. Your mix will move the number.
Call it
OpenAI-compatible — swap the base URL and go
curl https://model.cards/api/v1/chat/completions \
-H "Authorization: Bearer $MODELCARDS_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "openai/gpt-4o-mini",
"messages": [
{ "role": "user", "content": "Summarize the tradeoffs of speculative decoding." }
]
}'