the per-message tax is why your ai app meters you (and how one platform killed it)
Queues, daily caps, paywalled features — every AI companion meter exists because platforms rent their AI per message. Soulkyn just showed what happens when you own the hardware instead.

Let me explain the single most important thing about AI companion pricing that no platform wants to say out loud.
Every time your AI answers you, somebody pays for that answer. If the platform rents its intelligence from an API company — and almost all of them do — each reply costs them real money. Fractions of a cent to whole cents, times every message, times every user, forever. That’s the per-message tax.
Once you know it exists, every annoying limit in this industry stops being mysterious.
Character.AI’s free tier is “unlimited”… with a wait queue. That’s the tax. Chai gives you daily message limits with ads stuffed in the gaps. Tax. Replika locks the emotional and romantic features behind Pro at $19.99 a month. Kindroid, $14. Everybody’s structure is different but the shape is identical: they meter you because someone upstream meters them.
so here’s what soulkyn did instead
This week Soulkyn announced they’re running GLM 5.2 — a 744-billion-parameter frontier-class model, the one that’s been eating benchmark leaderboards since June — on their own private B300 clusters. Their changelog says it flat out: most services pay an outside company for every single AI reply, which is why they meter you. Running the model on your own machines means no per-message tax, and no tax is what makes real unlimited possible.
And then they did the thing the economics finally allowed: unlimited chat on that model for Deluxe and Deluxe+ subscribers. Actually unlimited. On a model in the same size class as the big-name frontier assistants people pay per-message to use for work.
I want to be precise about how unusual this is, because “unlimited” is the most abused word in this industry. Unlimited on a small model is easy — small models are cheap. Unlimited with a queue is a meter wearing a trench coat. Unlimited on a 744B model with no queue is only possible if the marginal message costs the platform close to nothing, and that only happens when the hardware is theirs and it’s sitting there anyway. Which, per the announcement, it now is.
the fun backstory: this model almost died
Here’s the detail that makes me trust the whole thing more, not less. A while back Soulkyn openly told users that GLM was losing them money and might have to go. Platforms don’t usually admit that. The renting economics genuinely didn’t work.
Instead of killing it, they bought the metal, self-hosted, did what they describe as “a lot of optimization work,” and flipped the equation. The model that was bleeding them dry as a rented API is now their flagship unlimited tier. That’s not marketing — that’s an infrastructure decision you can reason about. The math checks.
They’re honest about the rough edges too: their dev is still tuning load balancing, there was a brief mostly-invisible hiccup one night where traffic fell back to outside providers (which stay wired in as exactly that — automatic backup), and they expect a few weird moments while it settles.
what this means if you’re shopping
If you’re paying for an AI companion in 2026, ask one question: does the price buy me a relationship or a ration?
$9.99 at Character.AI buys you a shorter queue. $19.99 at Replika buys features the free tier withholds. Those are rations. The Soulkyn pricing page now lists “beta access to next-gen frontier models — currently GLM 5.2 unlimited” as a Deluxe benefit — that’s a different kind of purchase. You’re not buying permission to send more messages. You’re buying access to hardware they own, and the messages just… don’t get counted, because there’s nothing upstream counting them.
And the tax kill isn’t text-only. The same Deluxe plans include unlimited image generation — seven completely different checkpoints spanning photorealistic through anime (their newest engine goes by ZImage). Images are the OTHER thing this industry meters brutally, because rented image APIs charge per picture the same way rented text APIs charge per message. Voice messages: also unlimited on these plans, same reason. Owning the GPUs kills every per-unit tax at once — text, image, audio — which is why one subscription can be unlimited on all three and your competitor’s can’t be unlimited on any.
Fair disclosure, since they were fair about it themselves: GLM 5.2 as the default is still officially labeled a test. They say they truly hope it stays — and that right now it looks like it can — but they want more time before it’s carved in stone. Given that the fallback for “test fails” is the model they already ran before, the downside seems to be “things go back to being merely good.”
The per-message tax built this industry’s every cage. Watching one platform dynamite it is genuinely satisfying. More of this, please.
