Why Your TypeScript Code Costs 73% More on Claude Than GPT
The real prices of frontier models. Tokens * Price, right?

I discovered that comparing AI model prices by tokens per million is misleading because tokenizers differ wildly. The same TypeScript file generates 73% more tokens on Claude than on GPT, inflating costs even with identical rate cards. New tokenizers from Anthropic quietly increase bills by over 30% for the same code, making the sticker price a dangerous illusion for developers.
The price is per token, but a token is not a fixed amount of text: each model's tokenizer cuts the same file into a different number of pieces, and you pay per piece.
- SwellJoe
The fact that OpenAI documents theirs is already a big improvement over Anthropic. But, also, the OpenAI tokenizer got more efficient when they last updated it, rather than less. https://mdstudio.app/o200k-base-tokenizer
- Tiberium
Yeah, Anthropic's current tokenizer in Sonnet 5/Opus 4.8/Fable 5 is much worse than OpenAI's. Also, OpenAI has been using their current o200k_base from the day GPT-4o came out over two years ago. Just a few of my own tests:
- A ~2000-2002 legacy C++ game codebase at about ~90kloc: GPT 1.12M, Claude 2.2M
- A ~30kloc TypeScript codebase: GPT 260K, Claude 437K
In the end, GPT's current tokenizer is ~1.6x-2x better than Claude's current one, depending on your data. And you can check for free for both, for OpenAI just use the open-source libraries, for Anthropic - you have to use their count_tokens endpoint as they don't publish the tokenizer, but the endpoint is free (and allows requests over 1M tokens as well).
- lolinder
This piece focuses on the cost differences from the tokenizer, which do matter, but I wish they emphasized more that even adding the tokenizer to your calculation doesn't provide you with a good way to calculate cost for agentic coding tasks.
Other traits where models differ that have an even greater impact on your total spend:
* How much context do they load in to solve a given task?
* How long do they spend thinking to get equivalent results?
* How many times do they stop and ask you for input, and are you there to respond to them before the cache runs out?
* Etc.
Incorporating the tokenizer just makes a very imprecise measurement of cost a little bit more precise, but in my own experience I have not found that the token cost is a significant driver of task cost whether or not you incorporate the tokenizer. Everything else about the model's behavior has a much larger impact.
- iLoveOncall
Very unpalatable completely LLM-written article, but on top of that a lot of the fundations and conclusions are completely wrong, the main one being this one:
> You will see people claim Claude uses 2x to 4x the tokens of GPT. Our measurements do not support that, and overstating it would undercut the real point.
It's not because a single prompt represents only 1.7x the number of tokens that a model doesn't use 4x as many tokens as another, when running as an agent. This doesn't take at all the number of tokens of the output into account, and the number of tokens of the potential tool calls from this output, which directly feeds back to input tokens.
The article also has a very small test set (16 documents), all of very small length (15K tokens at most, when models go up to 1M in context and agents routinely exceed this and have to summarize).
Complete garbage article.
- jnwatson
The real elephant in the room is pricing for KV cache writes and reads. That makes all the difference for tasks with large context.