πŸ’° TokenWatch

WyrdWerk GitHub LinkedIn X

Know what your AI actually costs before the bill arrives

Practical benchmarks β€” agentic coding, reasoning, knowledge work β€” joined to real pay-as-you-go pricing. Scores measure the model; the price is the cheapest provider's blended rate at the agentic mix.

Top benchmarked models and their cheapest blended price

Snapshot of the 25 highest-ranked models (Artificial Analysis Intelligence Index, 258 benchmarked models total). The interactive table above adds use-case tabs, live token-mix pricing, and capability-per-dollar value ranking.

ModelCreatorAA IntelligenceAA AgenticAA CodingLiveBench ReasoningDesign Arena EloFrom $/M
anthropic/claude-sonnet-4-6anthropicβ€”β€”β€”84.8β€”$3.06
anthropic/claude-opus-5anthropic63.159.278β€”1382$0.735
Claude Opus 5 (Fast)anthropic63.159.278β€”1382$1.47
Claude Opus 5 (batch)anthropic63.159.278β€”1382$0.735
anthropic/claude-fable-5anthropic62.156.676.5β€”1383$1.47
Anthropic: Claude Fable 5 (batch)anthropic62.156.676.5β€”1383$1.47
gpt-5.6-solopenai60.957.877.4β€”β€”$0.147
SpaceXAI: Grok 4.6xai60.958.776.8β€”1340$0.565
moonshotai/Kimi-K3moonshot59.754.376.2β€”1441$0.331
Kimi K3 Fastmoonshot59.754.376.2β€”1441$0.441
MoonshotAI: Kimi K3 (batch)moonshot59.754.376.2β€”1441$0.376
zai-org/GLM-5.3z-ai59.559.174.8β€”β€”$0.075
Qwen/Qwen3.8-Maxqwen58.158.471.8β€”1368$0.266
Qwen/Qwen3.8-2.4T-A95Bqwen57.757.171.9β€”β€”$0.274
Qwen: Qwen3.8 2.4T A95B (batch)qwen57.757.171.9β€”β€”$0.274
zai-org/GLM-5.3-Flashz-ai57.558.271.5β€”β€”$0.013
Z.ai: GLM 5.3 Flash (batch)z-ai57.558.271.5β€”β€”$0.016
Anthropic: Claude Opus 4.8 (Fast)anthropic57.349.474.3β€”1310$1.47
Anthropic: Claude Opus 4.8anthropic57.349.474.3β€”1310$0.735
Anthropic: Claude Opus 4.8 (batch)anthropic57.349.474.3β€”1310$0.735
Meta: Muse Spark 1.2β€”56.849.372.2β€”1357$0.198
gpt-5.6-terraopenai56.650.276.7β€”β€”$0.152
OpenAI: GPT-5.5openai56.347.474.989.71336$0.38
google/gemini-3.7-flashgoogle5645.176.1β€”1355$0.055
Google: Gemini 3.7 Flash (batch)google5645.176.1β€”1355$0.055

Scores from Artificial Analysis, LiveBench and Design Arena. "From $/M" is the cheapest provider's blended rate at a cached-heavy workload mix. See the benchmark FAQ for what each score measures.

Frequently asked questions

Where are the full FAQ lists?

All questions β€” text/token pricing, image and video generation pricing, and plain-language benchmark explainers β€” live on the consolidated FAQ page.