Frontier AI models like GPT-4 and Claude 3.5 carry a higher total cost than suggested by simple token pricing. A new analysis from PlayCode reveals that factors like throughput limits, latency penalties, and required infrastructure (e.g., high-end GPUs) inflate actual usage costs by 30-50% for heavy users. The study compared list prices with real-world deployment scenarios, including batch processing and real-time applications. These findings challenge the assumption that token prices alone determine affordability.


We're still thinking about AI costs like we're buying soda by the can. Token price per million is the headline, but the fine print—throughput caps, cold starts, and GPU rental—is where the real bill lives. Heavy users feel this most. They're not just paying for tokens; they're paying for speed, reliability, and access.

But here's the good news: this forces a necessary maturity. As the market matures, we'll see more transparent pricing and optimized infrastructure. The hidden costs today are signals for innovation. Tomorrow's models will be cheaper to run, not just cheaper to list. That's progress.