LLM Model Comparison
Context windows, pricing, and best-fit for the leading models — filter by provider.
| Model | Provider | Context | In / 1M | Out / 1M | Best for |
|---|---|---|---|---|---|
| Claude Opus | Anthropic | 200K+ | $15 | $75 | Hardest reasoning, agents, long docs |
| Claude Sonnet | Anthropic | 200K+ | $3 | $15 | Best all-round; coding & agents |
| Claude Haiku | Anthropic | 200K | $0.80 | $4 | Fast, cheap, high-volume tasks |
| GPT (flagship) | OpenAI | 128K+ | $2.50 | $10 | General-purpose, tooling ecosystem |
| GPT mini | OpenAI | 128K | $0.15 | $0.60 | Cheap classification & extraction |
| Gemini Pro | 1M+ | $1.25 | $5 | Huge context, multimodal | |
| Gemini Flash | 1M | $0.075 | $0.30 | Ultra-cheap, long-context, fast | |
| Llama (open) | Meta / self-host | 128K | self-host | self-host | On-prem, data residency, no per-token fee |
Prices are approximate public list prices (USD per 1M tokens) and change often — verify with each provider. We're model-agnostic and pick the best model per task on quality, latency & cost.
How the model comparison works
A tool that will not show its arithmetic is worth less than one that does.
- 1
Models are listed side by side with context window, published pricing, modality support and stated strengths.
- 2
Figures come from each provider’s own documentation rather than from third-party benchmarks.
- 3
You filter by what you actually need — long context, low cost, vision, tool use — instead of reading a single ranking.
What it assumes
Provider-published specifications. Real-world behaviour on your task can differ substantially from any headline capability.
This is deliberately not a leaderboard. Benchmark rankings move constantly and correlate weakly with performance on a specific application.
Pricing is list price, excluding committed-use or batch discounts.
How to read the result
Use this to shortlist two or three candidates, then evaluate them on your own data with your own scoring — that is the only comparison that predicts anything. The right model for a task is frequently not the highest-ranked one: a smaller model with a longer context window, or one with better structured-output support, often wins on the thing you are actually building.
Questions
- Which model is best?
- Not answerable in the abstract, and anyone who answers it confidently is selling something. Best depends on your task, your latency budget, your cost ceiling and your accuracy floor. Shortlist here, then measure on your own inputs.
- How often does this change?
- Frequently. Treat any comparison of this kind as a starting point and confirm current specifications and prices with the provider before committing.