The balanced Claude: near-frontier coding and agent quality at a mid-tier price. It is the default choice for production assistants that run all day and need consistent, well-formed tool calls. Long context and steady instruction-following make it a safe general upgrade.
Models
Every model reachable through one endpoint, on your own provider keys. Compare context windows, per-token pricing, and the providers serving each model before you commit a single line of code.
58 models
The high-throughput member of the Gemini 3 family: most of Pro quality for a small fraction of the cost. The huge context window and low latency make it a favourite for retrieval and document pipelines at scale. Multimodal input is included at the same price.
OpenAI's flagship general model, tuned for long multi-step work and tool-heavy agents. It holds instructions across very long sessions and can switch between quick answers and extended deliberation. The default pick when you want maximum reliability rather than the lowest price.
A sparse-attention update to the V3 line that cut long-context serving costs dramatically. It folds chat and reasoning modes into one checkpoint and is among the best value-for-quality models available. A frequent choice for cost-sensitive production traffic.
Google's flagship multimodal model, reading text, images and audio inside a million-token window. It is strong at long-document analysis, chart and video-frame understanding, and agentic planning. The obvious pick when the input is large and messy.
A cost-optimised Grok 4 with a two-million-token window and unusually low pricing for its class. Designed for search-heavy and high-volume agent workloads where you cannot afford flagship rates. The value leader in the xAI line by a wide margin.
Anthropic's most capable model, aimed at hard engineering, research and long-horizon agent work. It sustains focus across very long tool-using sessions and is unusually literal about following detailed specifications. The one to reach for when a wrong answer is expensive.
A fast, inexpensive Claude tuned for high-volume interactive work. It keeps the family discipline around instructions and tool schemas while answering in a fraction of the time. Good for inline product features, triage and sub-agents inside a larger system.
A cost-reduced GPT-5.2 that keeps most of the reasoning quality for a fraction of the price. Well suited to high-volume classification, extraction and everyday chat. Exposes the same tool-calling interface as the flagship, so swapping between them is a one-line change.
A fast, low-cost Gemini with an optional thinking budget you control per request. Heavily used for summarisation, extraction and chat over large corpora. A dependable workhorse that is hard to beat on cost per useful token.
The previous flagship of the GPT-5 line, still strong at analysis, coding and structured output. Kept in the catalog for workloads pinned to its exact behaviour. Reads images alongside text and supports adaptive reasoning effort.
The small sibling of GPT-4o, with a very low price and quick responses. Popular for chat widgets, summarisation and bulk data cleanup where frontier quality is not required. Handles images, which is unusual at this price point.
Qwen's flagship coding mixture-of-experts, trained heavily on repository-level tasks and tool use. It handles long files and multi-step edits well and is the most common open-weight substitute for proprietary coding models. Pairs nicely with agentic CLI harnesses.
A GPT-5.1 variant specialised for software engineering: repository-scale edits, test-driven iteration and long agent loops in a terminal. It is trained to keep working through multi-file refactors instead of answering in one shot. Less chatty than the general model and noticeably better at patch formats.
An open-weight mixture-of-experts model from OpenAI with configurable reasoning effort. Because the weights are public it is served by many accelerator vendors at a fraction of API-model pricing. A strong default when you want GPT-family behaviour on commodity inference.
The previous Gemini flagship, with built-in thinking and a million-token window. Still very capable on long-context reasoning, code and multimodal analysis. Widely deployed and well understood.
A dense 70B model that matched the far larger 3.1-405B on most instruction-following benchmarks. The pragmatic open-weight default for general assistants and fine-tuning. Supported by essentially every inference vendor, so pricing is competitive.
A reasoning variant of Kimi K2 that interleaves thinking with hundreds of sequential tool calls. Aimed at deep research and long autonomous sessions rather than snappy chat. Holds up remarkably well across very long trajectories.
The refreshed R1 reasoning model, which emits a full chain of thought before answering. Very strong on mathematics and algorithmic problems relative to its price. Expect high completion-token counts, since the thinking is billed as output.
OpenAI's 2024 omni-modal workhorse, taking text and images with fast conversational output. Still a dependable general assistant and one of the most common baselines in published evaluations. Widely supported by third-party tooling.
Z.ai's flagship open-weight model, with a large context window and strong coding performance. It has become a popular cheap drop-in inside agentic coding harnesses. Good tool-calling discipline for the price.
xAI's flagship, built around long reasoning traces and live tool use. Competitive at the top of mathematics and science evaluations and comfortable with images. Its search integration makes it useful for questions about very recent events.
The cheapest Gemini tier, optimised for very high request volumes. Best for routing, tagging and short answers where cost per call dominates the design. Keeps the long context window of its bigger siblings.
Meta's mixture-of-experts flagship, with native image understanding and a very large context window. Competitive on general chat and coding while staying inexpensive on open-weight hosts. A practical frontier substitute when you need weights you can move between providers.
A mid-2025 Sonnet with strong coding and reasoning across a large context window. A reasonable step down from 4.5 when you want an older, well-characterised behaviour. Still perfectly capable for most assistant workloads.
A large mixture-of-experts with only about 22B active parameters per token, so it serves far more cheaply than its size suggests. Supports switchable thinking and non-thinking modes in one checkpoint. A strong open-weight generalist.
Moonshot's open-weight mixture-of-experts, built for tool calling and agentic workflows. It is notably good at following structured API schemas and long instruction lists. A capable open alternative for assistants that live inside other software.
The smallest and cheapest member of the GPT-5 family, built for latency-sensitive glue work. Good at routing, tagging, short rewrites and pre-filtering before a larger model runs. Fast enough to sit in the request path of an interactive UI.
Alibaba's largest Qwen, a trillion-parameter-class mixture-of-experts served through hosted APIs. Excellent multilingual coverage, particularly across Asian languages, and competitive on agentic benchmarks. The top of the Qwen range for general quality.
A compact reasoning model that spends extra tokens thinking before it answers. It punches far above its size on mathematics, competitive programming and structured logic. A good value pick when correctness matters more than response speed.
The smaller Llama 4 mixture-of-experts, built for long-context retrieval on modest hardware. Cheap, fast and multimodal, with quality close to the previous generation of dense 70B models. A sensible default for document-heavy pipelines on a budget.
A compact mixture-of-experts aimed squarely at coding agents and multi-step tool use. Fast and cheap enough to run long agent loops without watching the meter. Punches above its weight on end-to-end task completion.
A small, extremely cheap open model that still handles chat, summarising and simple extraction well. Common as a draft model for speculative decoding and as a cheap classifier. Runs comfortably on a single consumer GPU if you self-host.
A mid-2025 refresh of DeepSeek V3 with better coding and front-end generation. A solid, inexpensive general model with fully open weights. Still widely served because so many providers have it warm.
The smaller open-weight OpenAI release, light enough to run on a single accelerator. Handles chat, tool calls and light reasoning at close to zero cost per token. Often used as the draft model in speculative decoding setups.
The lightweight member of the GLM 4.5 family, tuned for fast agentic use on modest hardware. Keeps tool calling and light reasoning while cutting cost substantially. A sensible sub-agent model inside a larger pipeline.
The first Claude with a visible extended-thinking mode you can budget by tokens. Reliable for code review, drafting and document analysis. A common choice for teams that want thinking traces they can inspect.
A dense 32B Qwen that fits comfortably on two GPUs. Balanced quality across chat, translation and code at a very low hosted price. A good fine-tuning base when you need predictable dense-model behaviour.
A mid-sized multimodal model positioned as frontier-class quality at commodity pricing. A good all-rounder for enterprise chat, retrieval and coding. Deployable on-premises for customers that need it.
Mistral's code-specialised model, trained for fill-in-the-middle completion across dozens of languages. Tuned for editor autocomplete, where first-token latency matters more than depth. Handles long files thanks to its extended context window.
A small, cheap Claude from 2024 for quick classification and short-form generation. Fast enough for inline UI features and cheap enough for overnight bulk jobs. Text-only, with solid instruction-following for its size.
The previous Opus generation, still excellent at deep analysis and large refactors. Priced at the premium tier and mostly kept around for teams that validated their evaluations against it. Superseded by Opus 4.5 on both quality and cost.
A compact multimodal model that runs comfortably on a single GPU. Improved instruction-following and function calling over 3.1 at the same low price. A good open-weight base for fine-tuning narrow assistants.
Google's open-weight model with vision input and a permissive licence. A good self-hostable option for multilingual chat and light multimodal work. Small enough to fine-tune on a single node.
Mistral's dense flagship, with strong multilingual coverage and reliable function calling. Popular with European deployments that need EU data residency. A solid general model for enterprise assistants and document work.
The previous xAI flagship, strong on general knowledge and conversation. Still used where its style and instruction-following were validated in production. Superseded by Grok 4 on reasoning-heavy work.
A 32B model trained specifically for long, deliberate reasoning. It reaches well beyond its parameter count on mathematics and logic, at the cost of many thinking tokens. Best used with a generous output budget and a patient client.
Meta's largest dense open model, useful when you want frontier-adjacent quality with weights you control. Heavier and slower to serve than the mixture-of-experts generation that followed it. Still a strong teacher model for distillation.
A small reasoning-capable Grok that thinks before answering at a fraction of flagship cost. Good for quantitative tasks on a tight budget. Exposes its reasoning effort as a request parameter.
A very low-cost multimodal Nova for high-volume document and image workloads. Trades depth for throughput and price. Useful for first-pass extraction before a stronger model reviews the hard cases.
An agentic coding model built to operate inside a repository: reading files, running tests and applying patches. Designed for scaffolded software agents rather than single-turn question answering. Strong on issue-resolution benchmarks for its size.
Perplexity's search-grounded model, which retrieves live web results and answers with citations. Best for questions where freshness matters more than raw reasoning depth. Search fees are folded into the token price here.
Cohere's enterprise flagship, tuned for retrieval-augmented generation with inline citations. Covers 23 languages and is designed to run on relatively little hardware for its class. A natural fit for regulated document workflows.
Amazon's mid-tier multimodal model, reading text, images and video frames. Priced aggressively inside Bedrock and wired into the rest of the AWS tooling. A convenient default for teams already billing through AWS.
NVIDIA's alignment-tuned Llama 3.1 70B, reworked with reward-model feedback for more helpful answers. Effectively a drop-in upgrade for teams already serving Llama 70B. Same architecture, so existing serving stacks need no changes.
A reasoning-tuned search model that plans a query strategy, reads sources, then explains how it got there. Suited to research questions that need several hops across the web. Slower than plain Sonar but much better at synthesis.
Our free house model. It returns deterministic synthetic completions so you can exercise the gateway end to end — auth, streaming, usage accounting — before connecting a provider key. Always available, and the only model anonymous visitors can call. It answers in prose and never calls tools, so requests carrying `tools` are refused here rather than answered blind.
The smart default. Every request is scored for complexity — length, code, reasoning demands, conversation depth — and routed to the lightest model that can genuinely handle it, chosen from the providers you have connected keys for. Simple asks land on fast, cheap models; hard ones escalate to frontier reasoning models. The router learns from your own usage log: models that recently failed for you are routed around automatically. You pay only what your own provider bills for the model that served the request.