Settings

Theme

GLM-5.3 (open-weight) beat Anthropic/OpenAI models – for 1/5 the cost

reinvently.co.uk

240 points by ed-is-ai · 132 comments

Reader

35 threads
hellohello2

This whole thing immediately reads as Claude generated, making it hard to take seriously.

Why do these results contradict existing serious attempts at benchmarking LLMs? Namely: https://artificialanalysis.ai/ https://arena.ai/leaderboard/agent

  • seizethecheese

    These results don’t just contradict more serious benchmarks, they are wrong on an entirely different axis. This is a saturated benchmark. Haiku gets 96%. The results here are “not even wrong” and this being #1 on HN right now is a massive smell of either bots or massive ignorance or both.

    • urams

      People REALLY want the open models to be better than the frontier labs'.

      • geek_at

        investors REALLY want the closed models to stay frontier forevery. mmw the future of AI will be local and offline

    • ed-is-aiOP

      Read the benchmark - it's on the basis of being 'good enough' for everyday tasks. Which is what fits most applications right.

      This is not a benchmark for testing them against an Einstein.

  • rdsubhas

    I'm seeing a worrying trend on HN. Nearly for all articles, there is one unquantified, unproven comment at the top saying it's 100% AI — no proof, just baseless emotion of what fits the commenter's writing style. This is the new witch hunt, or virtue trolling.

    Look at the later comments, they have substance oriented discussions.

    HN Mods - can you please consider a policy against such comments, it's overflowing the site and is diluting discourse and value. If people don't like an article, they can simply ignore it. These articles are reaching the top because enough people consider it of value.

    But when an article comes to the top, and the topmost comment and discussion thread is an unqualified witch hunt, it's getting sick.

    • Glyptodon

      I don't know about Claude specifically, but I think people are starting to internalize a sense of when prose reads as having "AI-smell" and I do agree phrases like "Same driver, same track. The LLM is the star." trigger it for me too. That said, that doesn't mean a ton about the whole thing - could be anything from humans starting to echo AI style to someone writing "give my results a headline summary" to an AI to someone saying "here's the data, write an article."

      • rdsubhas

        AI is trained on human data. And high quality human data at it's best.

        Can we assume everything we think as AI - must have had a high-quality human pattern behind it, and there is no way to 100% prove which is which - unless the author shows a screencast of them typing the artice?

        This is not healthy. The right thing to do is – if someone doesn't like an article, they should ignore it – they shouldn't so confidently brand it AI without any proof at all, just because it fits their mood and style.

        • Glyptodon

          I think it's more complicated than that. Training does a lot more than just make models imitate the highest quality training data, and even what high quality training data includes can be subjective depending on your goals and tastes. And a lot of effort does go into making sure the models don't have "bad personality" - I'm sure a lot goes into making sure they trend towards appropriate reading levels and various other things that aren't strictly about being a Mark Twain or JRR Tolkien level writer.

        • adastra22

          At this point AI is trained on AI data, and it is moving AI output into super attractors that have no correlation with human text.

        • avazhi

          > AI is trained on human data. And high quality human data at it's best

          And if there was any doubt that you don’t know how LLMs work, this line sorted it out lol.

        • dgellow

          > And high quality human data at it's best.

          LLMs are trained on Reddit…

        • MallocVoidstar

          AI writing is not high-quality writing. And I don't think it's trained on particularly high-quality writing, I'm pretty sure it's trained on SEO slop. The old outputs of gpt-4o that loved the word "delve" read just like SEO spam blogs.

    • thegeomaster

      For what it's worth, before I hurl such an accusation I always check the post in Pangram (https://pangram.com). It always detects the text at 90+% AI generated.

      Notably, Pangram is very conservative, and it's not difficult to manually get an LLM generated passage of text to turn human-written. So a score of near-100% AI generated means the writer didn't do even very light editing for a large part of the text.

      There are extremely good reasons to be skeptical of fully LLM-written content. Our attention spans and our online platforms were built in a time where a long, data-supported article with references was expensive to produce. The time to write it was vastly longer than the time to read it, which means you could usually rely on some good faith, baseline level of accuracy and thinking on the writer's part.

      With LLM-generated content, it's very difficult to know if 5 minutes, 5 hours or 5 days went into writing of the content. On the surface, it all looks similar, but the 5 minute version usually communicates very little or very shallow ideas, makes factual errors, and is generally lacking a lot of context. It's fast food writing.

      These low effort versions of content take way more to read than they take to write. And combined with the obtuseness of the writing style, it all places undue burden on the reader to figure out the underlying message, because a lot of it has been mangled by the writing process.

      I think LLMs are hugely helpful for writing, but to use their proper potential one needs to use them for feedback and engage with them at a level deeper than simply "write an article about X" or "rewrite this paragraph", and the text then doesn't obviously read AI generated as a bonus - I think nobody really has a problem with this.

    • dtech

      I don't know if it's an intentional joke or something, but your comment smell incredibly like it's written as AI.

      Per the writing, reading AI writing is like having something taste "chemical". Not very specific, but still a very recognizable and bad taste that makes it hard to enjoy and marks the thing as low quality.

    • voidmain

      The first thing I do when I see an interesting article title is click on the comments to see if people have noticed that it's AI generated. If I didn't have this option because of the policy you desire, I would probably just stop reading HN entirely.

    • dgellow

      FWIW this part of your comment does look generated by Claude:

      > I'm seeing a worrying trend on HN. Nearly for all articles, there is one unquantified, unproven comment at the top saying it's 100% AI — no proof, just baseless emotion of what fits the commenter's writing style. This is the new witch hunt, or virtue trolling.

      Was that the case? Genuine question. Here the em dash is an actual em dash symbol, where in the other paragraph you used what looks like a minus symbol. It’s also the type of construct used by Claude.

      • rdsubhas

        This is exactly my point. I've been writing with those dashes (em or en, whatever, just Alt+hyphen on mac instead of two hyphens). And now this is what people and tools use as the judging symbol for AI? That's just one example, AI text doesn't come from nowhere, it's trained on us, people.

        • ilyazub

          And iOS keyboard produces these dashes on double click on the dash (“-“) button: Single press: - Double press: —

    • ed-is-aiOP

      Hmmm, as the author: the feedback is interesting indeed and something to take onboard...

      Ultimately I hope the result provided value by stimulating debate. To me it feels like lots of people rejecting the idea that a Chinese open weights model could be THAT good. Yeap, so the writeup is formatted with AI, but the benchmark and the results has taken meeeeeee weeks of effort to bring it together. The point is I am not going head-to-head against AA with my spare time.

      And the concept is based on a Grand Tour - of if you are Brit Top Gear TV show where we put 1 star in a reasonably priced car (my humble, little test harness)

      Hence 'same drive, same track. The LLM is the Star'. The irony is entirely lost on some of these folk. I am not, in my sparetime trying to be Artificial Analysis. But actually the point is....do you really trust Artificial Analysis: or do you trust a benchmark that is free and open for you all to pull down and run (and adapt to your needs, in your circumstances and your problems).

      If there is anyone who wants to look at things with open eyes instead of the group think, it's all there for deep review

    • HDThoreaun

      I generally agree, but the very first sentence of this post is "Same driver, same track." This prose is so AI coded, that even if its your natural writing style you would change it so as not to confused with AI. If this was in the middle, fine, but as the very first sentence it is quite strange.

      • ljm

        I've been watching old TV shows lately, the original CSI, all those 24 episode a season procedurals and so on.

        Whatever we call 'AI coded' now has an awful lot in common with old TV screenplays where no word of dialogue was wasted. It's all the same style: punchy, plays on words, a bit of smart-ass in there.

        • HDThoreaun

          I think that makes sense. The "ai prose" stuff is a result of post training where the labs force the model to sound a certain way. It makes sense that the general chat models are trained to sound similar to mass market media, revealed preferences show the average person likes that kind of language.

      • lopatin

        To my horror, it kept the metaphor going on and off the whole article. Half way through, there is a section called "Three cars failed the crash test".

    • timmytokyo

      The "About" page basically admits the entire site is written by AI.

      https://reinvently.co.uk/about/

    • itishappy

      > [...] there is one unquantified, unproven comment at the top saying it's 100% [...]

      The parent was clearly not stating anything confidently.

      > Look at the later comments, they have substance oriented discussions.

      There are two sentences in the parent comment. Why is the top response (yours) pointedly ignoring the sentence with substance?

    • lopatin

      The poster's username is "ed-is-ai" so I don't know how unproven it is.

    • lelanthran

      > I'm seeing a worrying trend on HN. Nearly for all articles, there is one unquantified, unproven comment at the top saying it's 100% AI — no proof, just baseless emotion of what fits the commenter's writing style.

      Other than having all the AI tells, what proof can there be? Are you seriously asking that, because there is no way to provide hard proof that we should just stop pointing out obvious AI tells?

    • Farmadupe

      It's a valid shift to move onto actually trying to read the article critically (which I don't mean in an insulting way -- If you assume a writing has something worthwhile to tell you, reading it critically is how you learn the worthwhile thing)

      In this article, _I_ get unstuck right at the very first paragraph:

      > Same driver, same track. The LLM is the star. Seventeen leading models driven round the identical 28-realworld task lap — one harness, same verbatim prompts, deterministic grading — and the results go on the board.

      It jsut doesn't make much sense to me. At best, I think it can be glossed as... "I made an arbitrary benchmark which I'm not going to explain, and I plotted the results."

      ------

      Getting my own opinions out, this is blatant slop. It claims to be "deterministic grading", but then almost the _entire_ webpage is editorialization. Examples:

      * "If you only run one model, run glm-5.3"

      * "opus-5 posts the best rubric on the default panel"

      * "deepseek-v4-pro is nominally cheaper still at $0.0029 [...] treat it as a batch-only option."

      * " It performed well on what it completed"

    • jsnell

      It's not nearly all articles. It's predominantly for the AI-written articles. And no, it's not unsubstantiated allegations or "virtue trolling". The tells are painfully obvious, and can be verified with high quality AI-detectors like Pangram.

      And this really has to be policed. Once some tipping point is reached and too much of the HN homepage is AI slop, the site is dead.

    • seizethecheese

      I’ve previously called for such a policy, but I am slowly changing my mind. This blog post is so obviously slop… and I suspect it’s the product of an upvote ring of some sort (Haiku is 96% on this benchmark.)

    • malshe

      This article is pure AI slop. But if you want an objective metric, Pangram 4.0 says "100 % of this text is AI". In my experience, Pangram has a very low false negative rate and relatively high false positive rate for human writing. That means it errs on the side of humans. So if it says something is 100% AI written, I'm quite convinced it is.

      But beyond that, can't you see how terrible the writing is? This is unadulterated AI slop.

    • hartator

      Because it’s truer and truer.

      I won’t be surprise if 80% of content in social media - including maybe HN - is AI generated as of now.

    • hellohello2

      Respectfully, the linked website reads exactly like the Claude artifacts I read all day, i.e., it is low effort. I do not mind AI generated writing at all, but I do mind bad writing.

      Further, you ignored the actual contents of my comments, to latch onto a superficial aspect. Please tell me: why do these results contradict existing attempts at benchmarking LLMs, which were designed with considerably more effort? Because the website certainly doesn't explain why in a way that is human-readable, which is why I asked.

      EDIT: as explained by another commenter below, its because Fable refused to perform some of the tasks.

    • Jamesbeam

      I am curious why you think the commenter is the problem?

      If authors would be open about their use of AI right in the title, I could decide for myself if I want to read it or not and spend the time, and there would be no need to write a comment to warn other users not to waste theirs.

      So the root of the problem is authors not disclosing properly. Not commenters telling others about the clickscam the author is pulling.

      And the commenter really shouldn’t need to prove it’s AI written. The author needs to prove it’s not. If you want my attention and time, you should have a good reason.

      Right now, authors on HN are just abusing the fact that I need to click their bullshit to find out that it’s bullshit I wouldn’t have read if it was disclosed properly from the start.

      If I want llm output, why would I read your writing if I could just ask one myself and get an equal result?

      Also, it’s short-sighted to think it’s just the writing style that makes people think it’s fully or partially AI-generated. This website has a ton of people with decades of experience and domain knowledge in their fields.

      For people with domain knowledge, it’s quite easy to spot flaws someone without domain knowledge generating words about something they have zero expertise in, or only having it explained by the machine that is making things up, is making.

      I don’t see the need for moderation to police people who think something is AI generated. You tell people they can just ignore the article, but they can’t, they have to click and read to find out, and on the other hand you could also just ignore their comment if you don’t like it, or disagree and tell everyone why you think this is not AI-generated bullshit wasting others people time to generate clicks.

    • johnfn

      This article is so blatantly, obviously, painfully AI generated. The real "worrying trend" is this getting upvoted to the frontpage in the first place. If you want "proof" just chuck it into pangram.

    • angoragoats

      No, I think I will in fact call out AI slop when I see it.

  • Tepix

    Note that in AAs report, Kimi K3 was at #1, then they updated their criteria and published a new report on the same day where it was no longer at the #1 spot. They may be under pressure not to declare a chinese model as #1.

  • SwellJoe

    While I was inclined to push back on the results, with Fable and Sol being so low, I have to admit I've also run into refusals several times since the latest models have arrived, and I've had to use Kimi K3 or DeepSeek to complete the task. Usually security auditing type stuff, but Fable balks at all sorts of ridiculous things, sometimes stupid things. I've even had Fable fall back to Opus and then Opus refused the task as well. So, it actually is becoming hard to use US models for everything because they refuse to work on a pretty broad selection of security and security-adjacent tasks. I guess if you're not at a Fortune 500 or a member of a fascist government, you don't get to use the best models to protect yourself and that's just how it's going to be.

    But, you're right. The prose is miserable Claude-speak, difficult to wade through.

    • hellohello2

      Interesting idea, I had not considered refusals. I have ran into some as well although rarely. I'm not certain the 28 tasks described would trigger it though, if I understand correctly the security tasks are about avoiding prompt injection and not about doing security work.

      EDIT: You were correct, Fable and Opus reject some of the coding tasks, which is why they score lower. Thanks for explaining.

      EDIT2: I believe this benchmark is invalid, my Opus 5 runs the supposedly rejected tasks just fine.

  • EddieLomax

    It's because it was written by Claude Fable 5.

    https://reinvently.co.uk/about/

    > Fable Anthropic * Research Editor * Claude Fable 5

    > Challenges claims, tightens methods and prose, applies British English and removes hype that the evidence cannot support.

  • koe123

    Haha theres even an em-dash in the title

  • ckocagil

    Because it's notoriously hard to benchmark LLMs. Ultimately every benchmark is different and measures different things. This is why companies that make LLM models have private benchmarks - they find the areas where the model is weak and make that their goal.

    • hellohello2

      Yes, I agree, this is why this post is interesting despite being clickbait. You get what you measure but its better than being blind etc.

    • irishcoffee

      The whole concept is kind of silly. We don’t “benchmark” humans. Or do we, via standardized tests? Why don’t we just use those? Or is that what the benchmarks are? I have no idea.

jchw

Almost 10 models are passing the benchmark >95% - isn't that... substantially overly saturated?

I'm also really skeptical of benchmarks that place any Haiku model very high. I've been thoroughly unimpressed with Haiku and my opinion has been that you basically shouldn't use it. Yet here, it ties Kimi K3. HMMMM.

"Rubric Quality" seems a bit more realistic than "Pass Rate", but it is apparently judged by Fable 5...

This entire write-up is also obviously clearly very heavily AI-assisted, which doesn't help matters any.

  • sambusa_123

    Just read through some of the code "benchmarks", and I see why: https://reinvently.co.uk/tools/ed-o-meter/tests/

    Most are extremely trivial tasks. I would be surprised if a model from 2 years ago failed these...

    • nylonstrung

      I think most of them would have actually failed, it's only recently that models were any good at using tool calls and harnesses after they started post-training for that

  • rfgplk

    > Almost 10 models are passing the benchmark >95% - isn't that... substantially overly saturated?

    The only true benchmark for any of these models that I've discovered isn't if they can pass precanned SWE tests, but rather can they create something novel? This isn't even too difficult to test, just give it a seemingly impossible task let it spin and see where it ends up.

  • finaard

    Interesting - when Kimi K2.6 came out I switched over from Anthropic models, with at that time comparable to better results for me. I was using Anthropic via API, heavier months were roughly $400 worth of Anthropic tokens - I can get the same thing done via a $100 ollama subscription.

  • zuzululu

    yeah that haiku really diminishes the claims behind the benchmarks. luna-max is significantly cheap and it is a strong performer but i dont see it on the benchmarks.

    I just get a feeling this site started with an intent to elevate Chinese models above the rest so wouldn't be surprised if the whole prompt sail was set to that tune

gertlabs

One problem is that $30 per run is really noisy for many verifiable tasks. Ours usually run at least an order of magnitude more for a model in GLM 5.3's price class.

We just published GLM 5.3 results on our multi-agent coding evaluations and it's definitely an impressive model, coming in around #6. For the price, it's actually not Pareto optimal, falling slightly behind Grok 4.6 (which is faster and the same price) and Sol 5.6 (which uses far fewer thinking tokens for comparable results). As with most Chinese models, it excels at iterating in a harness while its first answer/base fluid intelligence is below the American frontier.

Data at https://gertlabs.com/rankings

iamcoder18

There's no way GPT 5.5 is better than GPT 5.6 Sol, and gemini-3.6-flash is better than both of those. I wouldn't trust this benchmark at all.

  • zuzululu

    benchmark is saturated but your claim that gemini-3.6-flash is better than 5.6-sol is also not trustworthy or accurate.

    maybe if you mentioned 3.7-flash it might have been slightly more believable (but still false).

    • MostlyStable

      He was remarking on how unbelievable it was. He sentence is a little bit hard to parse but he's saying that both 5.5 being better than 5.6 AND Gemini 3.6 scoring better than both are signs that this benchmark is not very useful.

      The reason for both of those things is, as you point out, the benchmark is very obviously saturated

      • zuzululu

        ah fair enough , 5.5 being better than 5.6 is just plain wrong and yeah makes everyone skeptical

solenoid0937

IDK I use open models every day for personal projects, and closed models for work.

Open models are all decidedly far behind Fable and a good bit behind Opus as well. All of these posts read like motivated/wishful thinking to me.

I get that people badly want the open frontier to be where the closed frontier is, but it is just so obviously not the case if you actually use the models on a real project.

ac29

Not sure I trust a benchmark where Haiku gets a nearly perfect score and Fable is tied for last place

  • svachalek

    Fable got heavily beaten down by its refusals, which is not too surprising; although a couple of problems got refused for reasons I can't even imagine and the page doesn't quote the refusal.

    Some of the other failures like the colicky baby one are also probably soft refusals, it's not clear what the grading criteria are but I'm guessing it got docked for not going anywhere near a possible diagnosis.

    • rfgplk

      They loosened the refusals recently. They aren't anywhere close to the Sol refusals.

  • poincareball

    I've been not just unimpressed by Fable, but actively find it to generate negative value.

    It hallucinates more, and in more destructive ways, than other models I've worked with and generates truly atrocious jargon and bizarre inhuman explanations that end up cluttering things. The code it writes is terrible too. Overly complex with a lot of technical debt.

    • rfgplk

      > I've been not just unimpressed by Fable, but actively find it to generate negative value.

      Fable is good in a few very specific domains (graphics programming) but otherwise it's an overhyped model. Far too expensive too. Opus 5 is outright better in every metric.

      • poincareball

        I honestly find Opus 5 a downgrade too. It's not as bad as fable, but nearly so, and is very argumentative. It'll straight up ignore directives and design decisions.

    • solenoid0937

      If it's overcomplicating and has poor explanations this seems like a harness/prompting issue?

      Fable is very steerable. For me it makes the "simplest thing that works" and - once I tuned it - very understandable explanations. (Unsteered, I agree that Fable's writing style is too cryptic.)

      I can let it loose on a 12+ hour task and it will test and verify everything autonomously. These days most of my work is done by letting Fable just run overnight, while my days consist of meetings/planning to figure out what I want Fable to build next.

Klaster_1

For the last week, I've been heavily immersed reverse engineering a device with help of GLM-5.3 and it surpassed all my expectations - I actually managed to achieve very way more than I thought I would. I never worked on such low level stuff, it would have taken me months to learn ARM assembly and how to find for and write exploits. Initially, I attempted this with Claude, but it blocked me on the very first message, so I got a refund and decided to try z.ai. The only downsides are that it's maybe a bit slower than my day job Opus and I had to pay ~200 EUR for a monthly plan in order not to bump into weekly limits in a couple of days. If this level of capability cost maybe 50 EUR, I'd strongly consider getting a long time subscription.

  • teravor

    GLM 5.2 and 5.3 are exceptional at reverse engineering. you just point them at an IDA Pro MCP and off they go doing whatever you want from them.

  • matheusmoreira

    > z.ai

    > The only downsides are

    The catch's in their revolting terms of service.

    • HighGoldstein

      Can you elaborate?

      • matheusmoreira

        Obnoxiously broad and perpetual license over inputs and outputs.

        Yes, Z.ai demands an unconditional, irrevocable, transferable, sublicensable, perpetual, worldwide license to use, modify, reproduce, adapt, publish, perform, distribute, and create derivative works from your prompts and outputs. The same license extends to your username and profile picture.

        Vague prohibitions on whatever may harm Z.ai’s "interests" or even the "national interests" of any country. Prohibitions on "disturbing" or "inappropriate" content, whatever that is. Professional use prohibitions.

        Discussing Z.ai is prohibited to the point even my posting this comment is against their terms of service.

        And they can of course ban you if you, in the "sole and absolute opinion" of Z.ai, have violated these broad terms, and if you paid for the discounted yearly plan you won't ever see that money ever again.

        Even the US companies aren't this bad.

      • 3asgf

        They are pretty broad and vague:

        https://chat.z.ai/legal-agreement/terms-of-service

        I don't even know if I could use code generated by them, because they claim the copyright.

        Better don't travel to Singapore (wise anyway because someone could slip drugs into your suitcase) or China if you use them.

CMay

> if you run one model, run glm-5.3

That is a horrible take-away from this, with only 28 tasks and a high pass rate for most models, it says almost nothing.

Test a model for your use case and use the fastest, smallest, cheapest model that 100% satisfies your use case.

Or, if you truly do need a model with strong generalized performance, definitely do not take a benchmark like this serious with such a limited task set.

  • ed-is-aiOP

    The point I'm making is that most models are good enough for most tasks, so choose on speed/cost.

    Definitely benchmark on your own tasks. GLM-5.3 is the winner on mine, on yours maybe not. I am not trying to be a universal benchmark, as these serve no one but the person doing the benchmark.

    My previous post sank like a stone, but all the evidence + code to run this + what you have todo to adapt it for your own use cases is all here https://github.com/ed-is-ai/featherbench

    Encourage everyone to eval like the devil

scottfits

my prediction is even if open source Chinese models are 90% as good (or even a bit better, which I don’t really believe because of benchmark hacking) enterprises will still pay for Claude / ChatGPT and the harness, integrations, and peace of mind versus using some Chinese cloud.

  • sschueller

    Enterprise's peace of mind is being able to use the model and not have the US government decide on a whim to block access.

    Additionally I may want to run attack simulations which requires the removal of safeguards. My only option is to use an open model I can run on my own hardware.

    • eightysixfour

      You think the US can’t, on a whim, decide US companies can’t use Chinese models? They already showed exactly how they would do it - designate it a supply chain risk and say anyone using it can’t be a provider to the government.

      • nicoburns

        For non-US companies, the supply chain risk is probably higher with US models. The US government has already removed access to some models (Fable). IIRC that particular incident also affected US companies.

      • sschueller

        How exactly is that going to be enforced? Anyway I am not actually talking about US based companies.

        • hyperpape

          > How exactly is that going to be enforced?

          Flow chart:

          1. Is your company run by fuckwits? If no: they will not try to trick the US government about whether they are using prohibited models when the government asks. Your CISO will block access. We're done. If yes, continue to #2.

          2. Are they the specific brand of fuckwits who would try to trick the US government about whether they are using prohibited models? If no: They will probably still not let you use those models, but maybe they'll be bad at enforcement. If yes: this is probably not the kind of company that it will serve your long term interests to work for, but have fun with the prohibited models.

        • eightysixfour

          > How exactly is that going to be enforced?

          How is anything enforced on B2G agreements? Contractually & legally, which turns into internal policy, which shuffles the risk on to the rogue dev deciding to use GLM instead of the mandated Grok subscription.

          This is literally the playbook they ran for Claude. I know folks who work for government contractors who were immediately going through the evals to get rid of Anthropic because it became a risk for them.

  • microtonal

    peace of mind versus using some Chinese cloud

    These are open weight models (GLM-5.3 soon too). You can run them on the Together AIs or Firework AIs of this world. Use OpenRouter or HF Inference Providers in between and you can effortlessly switch between models and providers.

    I have been using GLM and Kimi models the last few months mixed with the latest Anthropic models and for my daily work there is barely a difference anymore (except for pricing).

  • KronisLV

    On a tech level I’d say that Kimi and GLM 5.3 on Max reasoning are good enough for non-trivial planning and exploration and on High are good enough for various implementation tasks. They can easily replace Opus 5 for me and mostly even Fable (webdev with some ML and DevOps work on the side, as well as local software).

    All of that pretty much means nothing for the orgs that just want to do the AI equivalent of picking IBM.

  • nozzlegear

    > versus using some Chinese cloud.

    Just use Openrouter.

  • epolanski

    In the real enterprise world companies are running their processes writing Gemini "gems" or using copilot because they were already google/Microsoft customers.

    Am I the only one that knows people in industries like insurance, banking, consultancy, materials, etc? Cause none of them gives two damns about what the leading SOTA is, procurement and compliance matter.

Aeroi

ai;dr

can't take any generated benchmark seriously. if you produce actual results, then produce actual copy to go with it.

nylonstrung

The stealth model "Ox Alpha" has been crushing benchmarks and appears to be the next release in the GLM family

CamperBob2

The GLM-5.3 weights are not yet open, and they've said that the delay is due to the need to nerf them for "safety."

So I have a feeling a lot of these early claims are not going to pan out in the long run.

  • bigyabai

    I'm not super worried about "safety" nerfs. Most refusals can be finetuned out, it's usually not a dealbreaker for open-weight models.

svachalek

Nice to see the TTFT chart, wish aggregators like OpenRouter would track this. Matches my experience, the Deepseek models while fast overall can have a horrendous wait before they start responding, and Claude models are superbly responsive. It's particularly annoying that models like flash and luna, where you've explicitly chosen speed over quality, can still stall out before they even get started.

_joel

Sorry, just can't read that page, too AI spammy

garn810

American people are fed propaganda that Chinese are bad people with bad tech etc

In fact, it's often superior

Looks at cars, BYD, Geely etc...

Looks at neural networks benchmarks...

angoragoats

The writing in the first three sentences is so bad that I closed the page.

Please, bloggers, write with your own voice. Don’t let an LLM do it for you.

gilesvangruisen

All of the answers. None of the understanding.

Arcuru

Anecdotal, but from my personal usage I found GLM-5.3 was not as capable at performing autonomous tasks as Opus/Sol. Certainly competitive with the Sonnet/Terra level, but not with the Frontier.

I got their lowest subscription tier and burned a week of quota on trialling it.

silverwind

Those benchmarks don't tell much, they only check if a problem was solved, not how. Also there's surely a lot of benchmaxxing going on in the model training.

zero0529

Having used GLM-5.3 I honestly don't think it is better than 5.1. It is slower and the result is often overengineered, it is if it overthinks everything.

tamimio

What’s the best model for planning and architecture design, rather than solving problems in codes or an issue?

amazingamazing

Hopefully all of these models lead to manufacturing breakthrough so we can bring down prices of cards.

gosolozero

2 articles on how G 5.3 is the best in the top 10? Seems like a bit of astroturfing going on

  • ed-is-aiOP

    I had to google astroturfing. I liked the due diligence... I am a hacker news infant, so all my history is about this.

    Give it a proper look. The results all there, and code if you want to run the evals (or make your own - it's very easy) https://github.com/ed-is-ai/featherbench

    You might have read about Ox Alpha aka glm5.3-flash, well I updated to include that

visiondude

so the coding tasks illicit refusal by anthropic classifiers and instead of updating tasks to do similar things that don’t trigger classifiers their choice is to count those as fails? feels wrong given their stated task list.

  • Youden

    I think that's reasonable. The goal of the benchmark is to determine how the model performs on real-world tasks. If real-world tasks trigger Anthropic's classifiers, that's a failure on Anthropic's part.

    I feel that choosing tasks that don't trip the classifier would also be a form of bias towards Anthropic.

    In case it's not clear, the coding tasks are really benign things; there's nothing security-, health- or biology- related in there. The classifier being tripped is definitely unreasonable.

    • ed-is-aiOP

      Indeed; I find it really interesting there is so much disbelief and outright hostility to my post, in these comments. I am an AI Realists rather than AI hype-artist, say it as it is. For me, the evidence is clear that for our own use cases OpenAI, Anthropic models should not be the default choice any longer.

      Incidentally, if you want to look at this in more depth - my previous post about the eval framework got like zero response; that's where the results, methodology, code for running the evals, evals are all shared.

      Anyone can run this to verify it for themselves

      https://github.com/ed-is-ai/featherbench

rfgplk

In their current form open weight models are simply not worth running. Literally the amortized cost of hardware + electricity you need to operate them is >> than the cost of paying for subscriptions.

dainiusse

Is the beater in the room with us now?

pcwelder

The whole thing (article, benchmark) is a slop soup.

- Metaphor overload. We get it, it's just like car racing. Show some mercy on human readers.

- X, not Y

- A, never B

- Tasteless em dashes

- Hallucinated data, like model add date. The standings are so unbelievable that they border on laughable.

hereme888

Chinese propaganda. Both current top articles on HN are shilling for GLM-5.3

  • ed-is-aiOP

    It's good to check the facts.

    Time will tell, I'm not the only one saying these things. But you look at the raw data and test to your hearts content if you care enough - all results and the code use to run it is above. And you can add your own tests if this isn't enough for you...

    Policy of total transparency

    https://github.com/ed-is-ai/featherbench

couAUIA

haiku 4.5 top 7, that benchmark is absolute crap

sehw

open-source when?

pulkitsh1234

>Short version: if you run one model, run glm-5.3 — 100% pass, a 9.3 rubric, $0.28 for the lap, about a fifth of gpt-5.5's cost

... I mean wtf is this prose ? what rubric ? what lap ? I can definitely say this is Opus 5.

jacobgold

The entirety of the coding evals are just 7 trivial coding tasks in Python? This is a joke.

  • ed-is-aiOP

    Thanks for the feedback.

    It was limited to 7 to keep things balanced. On the premise most people don't just code. The 7 were an example of my realworld use cases - the point of this is to encourage people to run their own benchmarks and not just take what they read as gospel

    All the code and results are here for anyone who wants to delve, see if they can reproduce. That is the point of discourse https://github.com/ed-is-ai/featherbench

spiderfarmer

And now there's 0x Alpha, which is their (now free) new model.

  • ed-is-aiOP

    Thanks for point out - I updated the benchmark today to include. It is every bit as good as everyone says it is. The intelligence / $ is something else...

vatsachak

> reinvently

Slop

qwertox

I can't wait for Mistral to host these models. For those who don't know yet, Mistral is pivoting to also offer Chinese models in their own cloud.

  • ed-is-aiOP

    That's pretty cool. All available already in OpenCode if you're a developer, see for yourself!

Keyboard Shortcuts

j
Next item
k
Previous item
o / Enter
Open selected item
?
Show this help
Esc
Close modal / clear selection