GLM 5.2 and the Coming AI Margin Collapse

I argue that the real cost of AI lies in inference, not training, and GLM 5.2 is the first open weights model to genuinely rival Opus and GPT. With costs under 20% of frontier labs and trivial migration via OpenAI-compatible endpoints, this shift threatens to collapse AI margins. While it lacks vision and web search, its price advantage makes it a dangerous drop-in replacement for many agentic workflows.
Training is a fixed, up-front cost, but inference scales with your demand and has genuine marginal costs.
- fny
I'm not convinced raw costs matter:
1. Compute costs collapsed since the advent of Cloud and yet hyperscalers still have fat margins.
2. Many open source office suites exist yet none compete with the ubiquity of gsuite or office. GitHub, Slack are similar examples.
3. Both Windows and macOS dominate the home desktop space despite free alternatives existing for a long time.
4. Many formerly open source infrastructure components like Redis and Elastic Search have Apache equivalents, but they still command healthy margins.
I understand the arguments for a margin collapse, but I don't see any historical analogues. It seems that enterprises will pay top dollar for service guarantees, integration, and someone they can sue.
It's nobody gets fired for buying IBM all over again.
- 01100011
I'll agree but from the other direction. AI continues to absorb my job as a senior systems software engineer (c/c++) and after a couple months I've only spent a few hundred dollars using gpt-5.5/5.6 and codex. I have no idea what people are doing to burn so many tokens but for me this is laughably cheap and every day I discover new capabilities. I don't care if costs go up or down, it's so cheap for what I get that I don't care.
- davedx
Meanwhile:
> China’s Ministry of Commerce has led meetings over the past month with major AI companies, including Alibaba, ByteDance, and http://z.ai/, to discuss measures that would restrict overseas access to cutting-edge AI models, including models that have not yet been released.
> The discussions reportedly include not only closed-source models but also open-weight models.
> Future regulations could take the form of a tiered framework based on technological capability. Basic open-source AI models may be managed through a filing system, high-performance models may be subject to security reviews, and the most sensitive frontier models may be banned from public release or restricted to use within China
https://www.reuters.com/world/beijing-is-looking-curbing-ove...
- KronisLV
They have a vision MCP to make up for the model itself not having the capability natively: https://docs.z.ai/devpack/mcp/vision-mcp-server
I also found their web search to be mostly okay.
Furthermore, in case this is of interest to anyone, if you use their ZCode harness then you get bigger Coding Plan quotas: https://zcode.z.ai/en
Used it for a bit, it sits somewhere between OpenCode Desktop (still new but nice) and Claude Desktop (recent versions are good).
As for GLM 5.2 as a model - with max thinking it’s generally satisfactory, somewhere between Sonnet 5 and Opus 4.8, better than DeepSeek V4 Pro for sure.
Pricing wise, the subscription doesn’t seem as good as expected. I spent like 60% of the weekly limits of the Pro (50 USD) plan in one day, only because each 5 hour limit only gave me 20% to spend, otherwise it’d be 80-100%. Not even doing anything crazy, just parallel long form work on 2 projects with about 96% cache rate and at most 3 parallel code review sub-agents.
Their Max (100 USD) subscription would last me the whole week, but so does Anthropic for the same money and so would OpenAI. Off-peak is more palatable but I can’t just twiddle my thumbs at 9 AM to 1 PM local time.
Proper savings would show up with the Max plan and yearly billing, but that’s more of a tough sell.
- spyckie2
It’s important that none of these entities can collude to price fix. Having China be the competitor ensures that.
Basic microeconomics is still the easiest way to understand token economies. How is it not a competitive market (where profits go to zero?).
Anything A or O does to keep more margin, any competitor can copy or choose to undercut, and undercutting has the benefit of collecting training data. So what is going to stop gross profit of tokens going to zero except for collusion/price fixing?
- typ
Unlike the belief that frontier AI is expensive due to a high margin, and going to be expensive if there is no competition. My understanding is that, under certain circumstances (which is most likely true), the price will be driven down just because of profit seeking.
The frontier LLM labs run on a huge fixed cost and very low marginal cost. They need the economies of scale to make sense of the business (an incentive to expand their user base as large as possible). Imagine that you want to buy a few B300s to run GLM 5.2 and rent the service out to other people. How could this business be viable and sustainable in the first place? You need as many customers as possible. If you charge everyone $1000, you find fewer customers who can afford it. It rots the ROA if the servers are not utilized 100% (you would better buy less compute instead).
Also, the marginal cost for onboarding a new customer is low. And it's getting even lower when you have more customers. You wouldn't leave money on the table (especially for your competitors) if you want to maximize your profit.
By this logic, all frontier AI labs are incentivized to lower the price to maximize their customer base, profit, and ROA.
- pixlmint
Last month, I cancelled my Claude Pro subscription and instead used those 20$ to purchase Openrouter Credits. Most of my knowledge-seeking questions can be answered by Gemma4, for basic code editing, Qwen3.6 27b is enough, and for really difficult tasks, GLM5.2 doesn't leave me hanging. I'm by no means a heavy AI user, so I'm even saving money going the API Credit route and relying on the smallest possible model depending on the task complexity.
- dbalatero
> Of course, this was a hugely poor read of where the costs actually lie in AI. Training - while no doubt capex intensive - is a fixed, up-front cost. You spend hundreds of millions to train a model, then you are "done".
I don't understand this point that people make. If you're consistently needing[0] to train new models and the cost of training relative to the % improvement seems to go higher, isn't this just a constant cost that you continue to bear? The footnote seems to allude to this, but then sort of waves it away anyways. Also are there continuing incremental training costs to keep models relevant? Or do they only have knowledge of events up to the day they were trained?
[0] needing, because you have competitors and people expect more and more.
- budsniffer952
>the least understood upcoming shift in AI economics.
Then proceeds to talk about something in the AI news every day. Hey, did you guys hear? Open source models are cheaper and their quality is increasing!
So, first, by no measure is GLM5.2 as good as Opus.
Second, yes, open source models will put pressure on margins...eventually. Everyone knows that. But do you think today's AI business model is the same as tomorrow's?
- Obertr
Metaphor i like is that it will be as cheap as electricty?
Do you know who is supplying your electricity or which factory it runs on? probably no, bc its a commodity and mostly settled and there is so many energy resources. some are alternative some are coal mines. And they all fight in the supply demand trade for energy which is happening real time ( think open router here)
And eventually the consumer wins bc of the abundance.
I think greatest example of abundance of cheap infinite intelligence will be not glm5.2 but DeepSeek V4 Pro max with $0.435 per 1M input tokens and $0.87 per 1M output tokens