Not in my experience, it tends to pick up subtle orientations given in a question (like "which is better, A and B?" and in the context you add you list a few things for A and B) and will absolutely run with them even if they're not true. Has been an issue with Gemini models since at least 3.0. Maybe that makes them great roleplaying models, but for factual information they just run with the slightest hint in one direction or another and never really push back objectively.
This guy had a terrible broken benchmark that gets hawked every release, and I wish HN would ban accounts that essentially exist to hawk a personally owned site, especially such a bad one.
I get similar results in my own tests. And Gemini 3.1 Pro is consistently on top of my ratings. Not everyone is coding monkey, I prefer staying a programmer.
They're referencing Gemini 3.5 Flash being the top model, you must not be great with detail.
And no (strong) programmer would jump to assuming other people are coding monkeys just because they disagree on what a strong LLM is: that's the kind of thinking reserved for the glorified coding monkeys who wasted their life getting better at writing CRUD apps and are now upset that someone's tooling is dropping the already very low bar there.
Karma systems are never perfect, and most people will not assume this is a pattern.
(ie. won't feel the need to downvote them just for having yet another crappy AI benchmark)
I only recognize it because I build a product that leaves me looking for information on every major release... and every major release a new crop of folks reply confused about the anomalies on top of anomalies that they're seeing, and they slowly learn this person is just way more unserious than the dogged distribution would imply.
As always, note: faster than GLM-5.2 doesn't mean too much, as GLM-5.2 is served by different providers, so the inference speed can vary drastically between providers or over time.
but I don’t mind the party seeing my trade secrets and thoughts compared to an American corporation + the party seeing my trade secrets and thoughts. So thats not a functional difference to me, and the Chinese one won’t reply to subpoenas so thats a value add tbh
I'm biased because I run an inference company, https://synthetic.new. That being said I think we're pretty good at serving at GLM-5.2 — and other models, like Kimi K2.7! — and our privacy policy is quite good: zero data retention for prompts and completions on API requests. Our average streaming TPS for GLM-5.2 (aka, tokens after factoring out time-to-first-token, which varies based on geography) is 97tps over the last 24hrs, although it's slightly lower at peak traffic in the mornings PST where it's 50-70 tps. We're also subscription-based which is nicer for coding than e.g. Fireworks which is per-token billing.
Interesting: I don't see anything in our error logs but we could be missing something (and personally the chat works for me + my unsubscribed test account). If you email us at hi@synthetic.new though we should be able to fix anything you're running into!
Fireworks.ai is solid. And if you care more about speed than cost they have a "fast" variant that I think just throws more hardware at the model for about 2x the cost.
Hi, PM at Fireworks here. We have zero data retention so we do not log any of your API requests. Realize you're talking about website activity which is different and will check and update on that too.
> the Chinese one won’t reply to subpoenas so thats a value add tbh
That's not something that's definite. They are not quite like the Russians. A lot of the governments in Asia are overly pragmatic and will happily strong arm their companies to throw users under the bus for the sake of a trade deal. There's a reason why Snowden ran to the Russians and not China.
Also, if they have any subsidiaries in the US, they may not have a choice in the matter.
Bedrock does not have GLM 5.2 and likely will not for quite some time. It seems like they are doing that on purpose due to pressure from Anthropic. DigitalOcean has it though.
the (imperfect) comparison having used both for planning and execution is that GLM5.2 is too jumpy and eager to do things, often to a fault (e.g. deploying/using git when it shouldn't) while sonnet 5 was much lazier than any Claude model I have used has been, not adding an addendum to a plan that I asked for, then lying that it did when asked. Looking at the analysis[0] I don't think it's worth it for me. Maybe for others. Fable was certainly much better.
Weak spots (categories it fails):
[0]: https://aibenchy.com/compare/anthropic-claude-sonnet-4-6-med...