Reasoning Tokens Change AI Benchmark Economics More Than List Prices Suggest
Artificial Analysis comparisons show why response speed, output volume and reasoning settings can reverse the apparent value of an AI model.
A current Artificial Analysis model-comparison page provides separate numbers for token prices, output tokens per task, reasoning-token use and time per task. Those categories tell a fuller story than a token-price column alone. In one October comparison of Haiku 5.5 and Gemini 4 Argon, the expensive model is materially higher on several quality measures, while the smaller one is far less expensive to call.
The central measurement trap is mixing unlike workloads. Two models may take different reasoning routes even when given the same user-visible request. A model that spends more time examining a problem can produce a better answer, but it could also breach a service-level target for customer support or a latency-sensitive assistant.
For benchmarking, log the entire cost of a successful outcome: input messages, retrieved documents, tool execution, output tokens, failed retries and human corrections. Compare high and medium reasoning modes separately, and never use a single best-case demonstration as proof of general reliability.
Benchmarks should therefore disclose both a capability measure and a resource budget. Without that pairing, 'best value' is a marketing judgment rather than a repeatable technical conclusion.
Reporting sources & references
These links identify the reporting or public materials on which the article is based; they do not imply our newsroom witnessed the events.