Your Usage
Advanced: cache write
Results
Frequently asked questions
Where are the full FAQ lists?
All questions — text/token pricing, image and video generation pricing, and plain-language benchmark explainers — live on the consolidated FAQ page.
Cheapest LLM API models right now
Ranked by effective cost at a typical agentic mix (2.5% input, 97% cached input, 0.5% output). Prices are USD per million tokens. Use the calculator above to compute your exact workload cost.
| Org | Provider | Model | Input $/M | Output $/M | Cache $/M | Blended $/M |
|---|---|---|---|---|---|---|
| singularity | singularity | flux-1-schnell | $0 | $0.0020 | — | $0.0000 |
| singularity | singularity | flux-pro-1.1 | $0 | $0.040 | — | $0.0002 |
| nex-agi | nex-agi | Nex AGI: Nex-N2-Mini | $0.025 | $0.100 | $0.0025 | $0.0036 |
| openai | openai | OpenAI: GPT-5 Nano | $0.025 | $0.200 | $0.0025 | $0.0040 |
| inclusionai | novita | Ling-3.0-flash | $0.021 | $0.063 | $0.0042 | $0.0049 |
| openai | singularity | gpt-image-1.5 | $0 | $1.00 | — | $0.0050 |
| openai | singularity | gpt-image-2 | $0 | $1.00 | — | $0.0050 |
| hyper | Gemma 4 26B A4B | $0.120 | $0.420 | $0 | $0.0051 | |
| deepseek | crof | DeepSeek: DeepSeek V4 Flash 0731 | $0.080 | $0.100 | $0.0030 | $0.0054 |
| meta | meta | Meta: Muse Spark 1.2 Contributor | $0.100 | $0.200 | $0.0020 | $0.0054 |
| meta | opencode | Muse Spark 1.2 Contributor | $0.100 | $0.200 | $0.0020 | $0.0054 |
| xiaomi | gmicloud | Xiaomi: MiMo-V2.5 | $0.119 | $0.238 | $0.0026 | $0.0066 |
| deepseek | crof | DeepSeek: DeepSeek V4 Flash | $0.120 | $0.210 | $0.0030 | $0.0070 |
| upstage | upstage | Upstage: Solar Pro 4 | $0.030 | $0.120 | $0.0060 | $0.0072 |
| qwen | alibaba | Qwen: Qwen3.7 Flash | $0.030 | $0.130 | $0.0060 | $0.0072 |
| xiaomi | xiaomi | Xiaomi: MiMo-V2.5 | $0.140 | $0.280 | $0.0028 | $0.0076 |
| xiaomi | opencode | MiMo V2.5 | $0.140 | $0.280 | $0.0028 | $0.0076 |
| openai | hyper | gpt-oss-120b | $0.188 | $0.700 | $0 | $0.0082 |
| nvidia | nebius | NVIDIA: Nemotron 3 Nano 30B A3B | $0.060 | $0.240 | $0.0060 | $0.0085 |
| qwen | hyper | Qwen3 Next 80B A3B Instruct | $0.117 | $1.14 | $0 | $0.0086 |
| ibm | deepinfra | ibm-granite/granite-4.2-3b | $0.030 | $0.120 | $0.0075 | $0.0086 |
| xiaomi | streamlake | Xiaomi: MiMo-V2.5 | $0.168 | $0.336 | $0.0034 | $0.0091 |
| xiaomi | novita | Xiaomi: MiMo-V2.5 | $0.168 | $0.336 | $0.0034 | $0.0092 |
| qwen | crof | Qwen: Qwen3.5 9B | $0.040 | $0.150 | $0.0080 | $0.0095 |
| deepseek | singularity | deepseek-v4-flash | $0.081 | $0.162 | $0.0070 | $0.0096 |
Pricing refreshed 2026-08-31 from public provider APIs. Verify prices on the provider's official pricing page before committing spend.
Explore TokenWatch data
Compare benchmarks by use case · Browse inference providers · Read the pricing methodology · Use the pricing API · Read the FAQ