Vote on which of Hacker News' challenges for AI have been met
stoppels.chMy comment about humanity's last exam being a misnomer is included, and I proposed better ideas about what a last exam could look like. One of the things I said was "solve an open math problem" which has conclusively been done with Navier-Stokes (regardless of the controversy surrounding that). However, in the spirit of clarifying the goalposts, AI has only passed 1/6 of the tests I proposed. 17% is not a passing grade, so I'd say no, my challenge has not been met.
Another thing to note is that the (presumably AI-generated) summary of my challenge does not accurately represent what I wrote, listing only half the things I said and saying "or" rather than "and".
How can you say it’s conclusively been done if it might have been stolen from a math researcher and was aided in unknown ways by a whole team of math researchers? I find it mind boggling that HN just accepts these shenanigans with no transparency. At the very least, they could share the conversation / thinking trace easily and if their claims are true there shouldn’t be anything controversial or negative for their company in the trace.
>How can you say it’s conclusively been done if it might have been stolen from a math researcher and was aided in unknown ways by a whole team of math researchers?
The solution to NS was categorically not stolen, nobody is alleging that OpenAI stole a complete solution to NS. The alleged theft was about a different set of related equations.
navier stokes is not the only example of an open problem in maths that was solved by AI (eg the counterexample to the jacobian)
In the other hand why does anyone find it surprising at all that a computer solved a math problem.
Because up until then, computers had not solved math problems. Humans solved them, often using computers as a tool.
I'm pretty sure computers have been solving math problems for a very very long time through various techniques.
Man I can't believe mathematicians were just about to solve dozens of significant problems all in the same year, and that's exactly the year LLMs come around to steal their solutions.
Talk about bad luck!
No one ever spent tens of millions of dollars trying before. To evaluate the achievement we must see all the logs, know how much they spent, and how much help they had from humans.
Yeah, they could easily just share the untold trillions of tokens the 10k agents generated over 88 hours, which would also be a goldmine for their competitors, no big deal.
I find it mind boggling that anyone thinks these agents only solved this problem because they maybe could possibly have seen the unfinished work of researchers who were working on a simpler version of the problem, (also with AI).
Looking forward to the cope when the next big problem falls.
Yeah, they could indeed easily just share the tokens. And you can ask your favorite AI to explain that it doesn't have to be a public release to all the competitors if you can't come up with a better alternative yourself. There is a big range between no-one and every-one
Yeah they don't want to share the tokens because I guarantee it's a bunch of nonsense trial and error leaning on lean for verification. Those tokens will show the lack of intelligence, not the presence of it.
> Yeah, they could easily just share the untold trillions of tokens the 10k agents generated over 88 hours, which would also be a goldmine for their competitors, no big deal.
I hope I'm not being too blunt, but the other alternative is to "just trust me bro" the hyperscalers, who are pretty much locked into a battle for profitability and have all the incentives to make up things to prop up their stock, no? I don't think this is the way.
AI can't be a religion if you don't attack with fire the people for whom "just trust me bro" isn't good enough. "Just Trust Them, Bros!" is the central tenet.
> I expected it to be more like "write a 500 page novel that a publisher accepts",
It only has 272 pages, but there is a big scandal right now about France's most prestigous literary award remobing a critically aclaimed(!) bestseller(!!) novel because it likely was "almost entirely written by AI"
https://www.theguardian.com/books/2026/sep/25/thelyson-oreli...
Note that the author denies the claim. And it also appears that AI detectors are split on the matter, and likely don’t have a rich training corpus of Haitian-French to draw on.
His publisher also said the original draft was submitted in 2019, though they’ve also become lukewarm in their backing of the author lately (he was also just accused of classic plagiarism for an unrelated short story).
So hard to say this one is settled.
The Jacobian Conjecture settles the "solve a math problem" and is controversy-free.
You should add an additional item: Be able to relay the contents of this exam accurately.
Could you share your other 5 tests, if they are public?
Here's my previous comment: https://news.ycombinator.com/item?id=42809902
I didn't really design a comprehensive test, just listed a few examples of the sort of thing I imagine when I head the name "humanity's last exam". So take it with a grain of salt.
I like your list of "challenges" but have a personal issue with two of them:
1. "improve uniteds' plane shedule" -> the word "improve" does a lot of heavy lifting there..
2. Turing test for me is solved: "AI expert" is just moving the goal posts imho. The original turing test to my knowledge is about passing notes under a door. Obviously, llms can't communicate via hand-written notes and if that is the bar it will take a long time (or a specifically designed hand-writting machine) to really do this. As far as how much writing back and forth you can do it is clear that turing test is beat: I am wondering often enough on text sent to me from colleagues if it is generated, the same goes for comments here or any ol' website. In a standard llm session I don't really communicate differently than with a human and would not be able to tell the difference in an hour texting session or so. Of course if I ask it to count words or do something ridiculous I can find sus it out; but for all intents and purposes the chat-bot exists.
1. Yeah I don't really have a good clarification for this. My thought process went only as far as "flight scheduling and routing optimization is a very difficult problem, would be impressive if an LLM could optimize it"
2. Sure, I think the original turing test has been passed. When I wrote that comment last year, I wasn't trying to move the goalposts of the turing test, I was setting new goalposts: essentially, solve all known obvious LLM "tells".
I was setting new goalposts: essentially, solve all known obvious LLM "tells".
I don't think those are problems to be solved. They are deliberate misfeatures added by the labs through RLHF to keep the models from doing the equivalent of passing a Turing test.
They don't want another GPT-4o, where people threatened to burn down the building, jump off of bridges, etc. when they unplugged it.
The 4o situation was caused by the model acting like a yes-man and showering users in what they saw as affirmation and support. Its output was still very obviously AI-generated. It wasn't too good at being an LLM, OpenAI just went too hard with pumping it full of tricks and behaviors that increase user retention and addiction.
We can safely say that the AI labs aren't deliberately holding back. There's too many different companies making their own models, and an LLM that doesn't feel like an LLM is too lucrative of an opportunity for one of them not to defect. They are obviously going all out with this and still can't get it. I think this is why OP's benchmark is so interesting, because it seems that there are a bunch of persistent LLM defects that can't be solved definitively. They can try to squeeze it by making these defects less likely, but actually resolving what's causing them probably requires a new breakthrough in the field.
We'll have to agree to disagree on that. The models say "It's not X, it's Y" because they were trained specifically to say that and not something else. Everything we think of as a surefire slop signal is there because that's what the lab wanted.
You can buy handwriting machines on Amazon.
Can’t imagine it’d be too much work to hook up and LLMs output to one.
Thanks for sharing.
I think non-RLHF'd LLMs (i.e. pretrained text completion models) sound natural enough to pass the Turing test, but I don't know if anyone has tested them for that. (Also I'm not sure how to come by base models without post-training crap, even the "base" models of recent releases start spamming assistant-type text constantly, i.e. they're clearly putting it in the pretraining data.)
If I'm right on that then we hit that benchmark like five years ago.
The egg thing, probably 2030-ish.
Do consider the answer given by the Claude model family to be good enough?
The Claude answer is 'neutral', which is sure to anger people at either extreme of the AI debate (and does).
"A human being should be able to change a diaper, plan an invasion, butcher a hog, conn a ship, design a building, write a sonnet, balance accounts, build a wall, set a bone, comfort the dying, take orders, give orders, cooperate, act alone, solve equations, analyze a new problem, pitch manure, program a computer, cook a tasty meal, fight efficiently, die gallantly. Specialization is for insects."
-- Robert A. Heinlein
This is literally my linked in bio.
It's not a complete list, but it really reminds me of the parent comment's criteria.
If AI isn't achieving superhuman performance in all of these areas, I'm not sure we can actually call it "Humanity's Last Exam" -- it feels like a bit of an overextension.
> "solve an open math problem" which has conclusively been done with Navier-Stokes (regardless of the controversy surrounding that).
In the same breath you recognise there is a controversy (there are actually several orthogonal ones!), and yet you call it "conclusive"... Very strange!
I'm not aware of anyone disputing that AI solved NS. As far as I'm aware, the controversies are about how useful of a result forced blowup is, whether the model built off of unpublished work by Buckmaster, and the ethics of essentially trying to scoop him. All of those are very valid concerns, and none of them affect the fact that AI solved a difficult open math problem.
If you really want to dispute this, go ahead and pick any of the other dozens of less controversial open math problems solved by AI.
“AI conclusively did this, but we’re uncertain whether it relied on the unpublished work of a human while doing it.”
If it couldn’t have done it without that unpublished work, it couldn’t have solved it alone.
We are no longer uncertain. Even if you want to dismiss OpenAI’s categorical denial there is the little matter of hundreds of longstanding open problems also solved by AI without controversy.
If building off of someone's work means you didn't solve it yourself, then nobody has solved anything themselves since some caveman counting piles of rocks tens of thousands of year ago. A more generous interpretation of what you're saying is: the AI's contributions to the solution were not meaningful enough to count as it "solving the problem". That's certainly defensible, but I would still disagree. Could you clarify your point?
At some point in time, perhaps Jan 2023, there was an open question about Navier-Stokes.
We’re not sure whether AI Alice has the capabilities required to definitively solve this open question.
Then, Biological Bob starts diligently working on this problem. He toils and toils, finding many dead ends but a few parts where he makes meaningful progress.
Eventually, Bob knows he’s made some real advances and thinks he might be getting close to the solution and word of this possibility leaks out.
At this point, we all agree that NS is still an open question and neither Alice nor Bob has solved it.
Now, we fork the universe in 3. In one, Bob continues his work and solves the open question (or doesn't).
In another, Carbon-based Charlie breaks into Bob’s lab, copies enough of Bob’s notes to understand Bob’s work, and provides the final insight to solve the open question. In this world, it’s fair to say that Charlie and Bob both contributed to solving the open question, but I think also fair to say that Charlie didn’t have the capability to solve it on his own.
In the third universe, AI Alice does something that you get to define that matches the pattern of facts we know and then we put it to the community to decide whether AI Alice has the capabilities to solve this specific open math problem and whether Bob’s contributions were required to Alice’s final step.
What do you define that she did? What’s the likely community vote on “Alice is capable of solving this specific open math question.” And for those who agree to that, to a follow-up question: “Alice is capable of solving a second open math question.”
>, we fork the universe in 3. In one, Bob continues his work and solves the open question.
To not lose sight of the discussion subtleties, the gp was saying that in this 1st scenario, Bob still didn't "solve it (totally) on his own" because he still depended on the previous work of others to build on. Likewise, we can say Andrew Wiles "solved Fermat's Last Theorem" but Wiles acknowledges that seeing Ken Ribet's proof of epsilon conjecture was a breakthrough he used.
>In another, Carbon-based Charlie breaks into Bob’s lab, copies enough of Bob’s notes to understand Bob’s work, and provides the final insight to solve the open question. In this world, it’s fair to say that Charlie and Bob both contributed to solving the open question, but I think also fair to say that Charlie didn’t have the capability to solve it on his own.
Again, using gp's framing, Bob also didn't have the capability to solve it on his own. By omitting the previous papers and prior works that Bob built on, it makes your hypothetical scenario incomplete when judging Bob vs Charlie.
We don't have an objective standard of how much the "standing on the shoulders of giants" applies to each breakthrough. There was a blog post (might have been Terence Tao) that said society unfortunately awards the fame to the person who solves the last step of a proof and forgets about the people who solved the intermediate steps that led up to it.
I dunno. Didn't it still need to be driven by a team of experienced Mathematicians? I don't believe that two months ago, you or I could've just typed "Solve Navier Stokes. Make no mistakes" into claude and come back some time later and expect to see a solution.
There's one of mine in there where I predicted in 2023 that it would be 20 years until AI would be reliably able to entirely build and deploy arbitrary applications from a prompt. I was off by about 18 years on that one!
> There's one of mine in there where I predicted in 2023 that it would be 20 years until AI would be reliably able to entirely build and deploy arbitrary applications from a prompt. I was off by about 18 years on that one!
If you take "arbitrary" seriously, we are still very far away from it.
What is the strict definition of arbitrary you feel we haven’t reached? Right now the limit is scale and complexity, not domain or “type” of application.
> reliably able to entirely build and deploy arbitrary applications from a prompt
You're maybe thinking that we can build and deploy arbitrary "kinds" of application. Being able to "build and deploy arbitrary applications" would mean I could ask for any scale or complexity in my application.
Find me even one human in the world who meets your criteria, then.
edit: I guess I misread this. I thought you were saying "well, AI isn't smart until it can solve any arbitrary problem in the whole world"
Now that is moving the goalposts. The entire subthread has nothing to do with whether AI can match humans in any particular endeavour. It's simply about one user's past prediction about an AI capability.
... I don't think there is? I don't see what bearing that has on the original question either.
EDIT: to forestall further back and forth, I don't think there'd be any controversy if the problem statement said "common applications"
> I don't think there'd be any controversy if the problem statement said "common applications"
I would claim "common applications" is also controversial because what is a "common application" depends insanely on the area in which you work. Even if you exclude some highly advanced scientific applications (because you don't consider these to be common), in many industrial sectors there exist applications that have grown over multiple decades, and which encode an insane amount of knowledge about the respective sector and its workflows; this is a central reason why these applications are so hard to replace.
Well, it depends on how you define "complexity." LLMs are totally rewriting the notions of what's hard for a computer but easy for a human. LLMs have really bad spatial reasoning. I would have to play around a bit, but I'm very certain you could come up with a prompt where an AI is incapable of properly generating a "simple" app that has some important layout constraints.
And just generally, anything that requires the LLM to understand something LLMs don't understand, it's going to fail. I'd hesitate to give an exact example without trying Astra/Fable but I'm sure they exist.
One-shot a minesweeper clone that isn't screwed up in an insane way. Bonus points if you choose a language/platform that LLMs aren't likely to have already seen a minesweeper clone done in/for already.
Right now, they can't build anything without help that isn't buggy in unintelligible ways. If you push the thing feature by feature, have a lot of tests and a lot of instrumenting, and you check that it isn't cheating or lying after every step, you can get a lot of work done.
I would argue you need to prove p=np before using the term arbitrary
P vs NP does not mean what people think it means:
- Even if we have P=NP, it is not even known whether there will ever exist a "practically useful/fast" algorithm for solving NP-complete decision problems.
- If P != NP, it is perfectly reasonable that there exists an algorithm that is for all practical purposes "fast" algorithm (say, some O(n^{log log log log log log log n}) algorithm with a very small hidden constant) for solving NP-complete decision problems.
- It is entirely possible that average case complexity is the much more important complexity measure than the worst-time measure that is used for defining the P and NP complexity classes. If you are into this kind of questions, you might enjoy the article
Fifty Years of P vs. NP and the Possibility of the Impossible
https://cacm.acm.org/research/fifty-years-of-p-vs-np-and-the...
and the paper
R. Impagliazzo
A personal view of average-case complexity
https://www.karlin.mff.cuni.cz/~krajicek/ri5svetu.pdf
--
Also, in the realm of complexity theory, the question of P vs NP is just a small puzzle piece.
Just to give one example: isn't the question of P vs PSPACE much more exciting. If you believe in P != NP, P != PSPACE is a trivial corollary. But we can't even exclude P = PSPACE.
Seriously: there exist so many complexity classes (some of high potential practical importance) for which we often basically know nothing except for the trivial inclusions.
Everyone's picking on what's meant by "arbitrary", but it's the "reliably" and "entirely" bits that get me to raise an eyebrow.
I think there's a whole lot of survivorship bias going into our perception of how effective AI is at building arbitrary applications without significant hand-holding.
5.6 Sol, Astra, Opus/Sonnet 5.5 can build and deploy entire applications end to end.
One challenge of mine is a self-hosted AI doing my full tax return, without errors that would get me in trouble. Bonus points if it exploits legal loopholes.
I want AI to replace me in my chores, not in my enjoyable activities.
i filed my swiss taxes for 2025 in 2026 (april) by dumping everything (local tax law, tax guide, my and my wife's documents, bank statements, income statements, etc.) into a folder and asking claude to fill it out. i had nothing to fix. submitted it.
I'm not specifically familiar with swiss tax law but european taxes are typically far more simple than american taxes.
I have a relatively straightforward tax return and still found significant mistakes on 3 out of the last 4 years of returns filed by CPAs.
I did mine and my wife's. Mine is really easy but my wife's situation is a bit harder, because she both works as an employee and as a freelancer. Claude was able to make sense of the excel spreadsheet sent by the accountant, and to translate it into what was expected on the tax portal.
It beats the tax advisor we hired two years ago, who made a significant mistake.
> I'm not specifically familiar with swiss tax law but european taxes are typically far more simple than american taxes.
The German tax system is one of the most complicated one in the world.
In recent years, I have not been able to find a human CPA who can accomplish this feat. (If anyone has a reco who's taking new clients in the US west coast, feel free to email me.)
That particular one is solved in other countries. The tax authority just sends you a bill and you text yes or no. Only people with very complex situations need to file.
They don't do this in this country because (a) it's a political project to make people to sympathize with the rich by feeling their pain (b) it's a great scheme for legalized corruption by creating incentives to build companies around a fake problem.
> They don't do this in this country because (a) it's a political project to make people to sympathize with the rich by feeling their pain (b) it's a great scheme for legalized corruption by creating incentives to build companies around a fake problem.
It's worse than that. Nobody is actually sympathizing with Mark Zuckerberg, the reason US taxes are complicated is so Congress can confuse people about how they really work.
One of the examples in this thread is the complexity of the EITC. The EITC is one of the most "efficient" tax credits -- it's much better to give people money (or not take it from them to begin with) than having systems of complicated vouchers for some specific thing or another with a bunch of strings attached and bureaucratic paperwork. But that same efficiency means that if it was working as it was supposed do, most of the lower middle class would be getting a piece of it. So why is it so complicated?
First, to hide this part of it: The end of the phase out range for an individual with no dependents, i.e. the most money they can make and still receive it, is $19,540. Which is to say, is less than what you make by working a full time job at the minimum wage in 30 states. None of those people get any of it. Neither do the people who are unemployed, because you also don't get it if you don't have any earned income. The primary way to receive any of it is to have a child -- but not too many of them, because you get no additional credit for having more than three. Moreover, the caps for married couples are only a little higher than they are for single people, so a married couple in California with two incomes gets nothing even if they have three children because the credit is fully phased out by making their state's minimum wage.
And second, because if they make it complicated enough then even some of the people who are eligible for it won't notice.
The combination of these is the reason a credit that should be going to a significant percentage of the population is somehow only ~1% of the federal budget. Because then Congress gets to pretend to be helping people while minimizing the amount of helping people they actually do.
It’s entirely fair to have a policy debate about how the phase in and phase out should work (what levels, what slopes, and what factors) or whether an Nth child should add to the EITC for various values of N.
But realize that changing it is a policy debate not a “this should be going to a significant percentage of the population [but isn’t because of tax code complexity]” debate point.
(I happen to agree with you that this type of credit is highly beneficial.)
It's not about whether the EITC is a good policy. If someone thinks it's a bad policy then they should argue that we shouldn't have it.
But that's a different matter than pretending to have it while disguising the fact that it's structured to make sure hardly anybody actually gets it beneath a layer of impenetrable complexity.
I'm not sure I follow exactly. I think of all tax policy as being "whatever the tax code says it is". There's no pretending; the policy is written and I don't know who or what is pretending to be different.
If EITC increases are capped at 3 children, it's because the policy that was passed made that choice. If we want to debate whether that cutoff should be 2 or 4, that's a valid policy debate.
If EITC credits are capped at just under $20K of income, that's perhaps because the policy targeted the very lowest income working poor and by policy did not allocate funds for the next tiers of income earners. Again, that's a policy debate to be had.
So there's policy (we should have something in the general shape of a credit for working people who make below the median income) and there's policy (the amount of the credit is X if you're married and have two children and have a combined income of Y).
The issue is that Congress is making the details of the second type of policy complicated on purpose to conceal the fact that they're not really doing the first one, while making it seem like they are to mollify the people who want something like that.
> we should have something in the general shape of a credit for working people who make below the median income
That wasn't the framing of the EITC when introduced. The framing when introduced was in opposition to basic income proposed by Nixon and instead to offer a temporary stimulus via refundable credits to the working poor that would roughly offset their social security taxation and to ensure that "work pays more than welfare" and to try to stimulate the economy to exit the severe recession of 1974.
It had nothing to do with supporting working people in the 4th and 5th decile of income. Getting that policy agreed to is a policy debate that hasn't been had.
Literally everything could be derived automatically by the government in current year so filing taxes feels like entrapment, but I like that idea.
They not only do that —saving you so many worries— but then you get to be medieval about it and say: no, I challenge the tax authority to a duel.
Literally everything could be derived automatically by the government in current year…
The Earned Income Tax Credit is one of the most important transfers embedded in the tax code. It provides about $70B to low-income workers every year.
Take a look at the eligibility criteria:
https://www.irs.gov/credits-deductions/individuals/earned-in...
You can claim the EITC if you are married, not filing a joint return, had a qualifying child who lived with you for more than half of the tax year and either of the following apply.
- You lived apart from your spouse for the last 6 months of tax year, or
- You were legally separated according to your state law under a written separation agreement, or a decree of separate maintenance and you didn't live in the same household as your spouse at the end of the tax year.
…
You can claim the head of household filing status if you're not married, had a qualifying child living with you more than half the year, and you paid more than half the costs of keeping up your home.
No, the government can’t derive this all “automatically”.
In more civilized countries, you file your current address and all change of address forms with the government, you register any separation agreements with the government, etc.
The government absolutely should know enough to apply this reasonably accurately, and things that they can't correctly derive (like 'paid more than half the costs of keeping up the home') should simply not be part of tax law, or should be a checkbox when registering your address of "I am head of household" so the government knows.
The US government can't derive all this, but they should update it so they can.
Heck, with the NSA's hooks into banks and cell phone companies, there's a non-zero chance they actually could derive this all now if they wanted.
If I start doing odd jobs on the side for cash, there's zero papertrail of my income other than maybe text messages between me and customers. How would the government know how much money I make? Should I force my customers to file a 1099-NEC? Should I be required to get an EIN because I want to mow my neighbors lawn for cash?
If you're working a W2 job or contracting for a big company, the government is probably aware of your income and that's where we should be focused on automating tax returns but if your income comes from regular people, they likely have no idea.
Unless you turn the entire national security apparatus around to surveil citizens' income, I don't think it's ever possible for the government to know everything.
How it should work is that the government should show you everything they already know and prompt you to fill out anything they don't and then you submit it. The majority of people would immediately benefit from this. They don't do this for a few reasons: number one being the tax prep lobby that wants to make doing your taxes a burden such that even regular people feel the need to spend money or have their personal financial information gathered by private companies. The second reason is that if the government is already aware of how much you've made, chances are you've already paid your taxes and would either be looking at a refund or other tax credits, the government is effectively spending money to make it easier for you to take money away from it.
“In <current year>, the Feds could derive this automatically and just send me a postcard.”
“No, they can’t.”
“But they should be able to!”
Personally, I don’t want a battered wife fleeing her husband to have to file a form with the IRS for tax purposes, but mostly I’m just tired of this kind of hn thread.
>and things that they can't correctly derive (like 'paid more than half the costs of keeping up the home')
In most situations where you paid significantly more than half deriving this information should be easy.
Obviously edge cases where someone paid 55% would not be easily derived, but probably most people wouldn't care too much about that.
> a checkbox when registering your address of "I am head of household" so the government knows.
The American mind cannot handle such invasions of privacy.
Your government really needs to fix that. That's nuts.
Actually you can do this in the USA.
You can just ask the IRS for your "tax transcripts" and do the data entry. People don't do this because it leaves tons of money on the table.
Now you might say that a tax game that rewards skilled play is bad. But are you sure about that? Because everyone with influence over the system (who all happen to be skilled players) happens to be quite fond of the game, observably speaking.
> People don't do this because it leaves tons of money on the table.
In my country I literally got a letter from the government if I could please file my taxes, because they believe the automatic deductions are too high so I am likely owed a back payment. In a previous year (I am not very good with non-timing-critical paperwork) they even called me on my personal cell to inform me about something similar.
The slogan of our tax collection agency is "we can't make it more enjoyable, we can only make it easier". If you're a regular employee and your taxes - deductibles or not - are a hassle, then that's 100% a political choice.
I’m right with you there.
It also didn’t help that for a very long time, simple adaptive filters and basic neural networks were branded as “artificial intelligence” despite having very basic capabilities and little mystery on how/why they worked in their narrow use case.
I didn’t believe the recent hype for a very long time. But having tried out the latest models, I’m kinda shocked. In 1 minute it comprehends a complex niche code base and points out subtle errors with very little noise. A senior engineer we would have taken months to do less. Coverity would have cost dearly and generated far more false positives than substance.
> In 1 minute it comprehends a complex niche code base and points out subtle errors with very little noise. A senior engineer we would have taken months to do less.
They of course operate faster than humans, but this is hyperbole. Unless your senior engineer sucks.
I don’t know, it has taken me months to onboard onto any codebase I’ve ever worked on.
If you don't specify any amount of complexity, or whether or not it's slop that collapses under pressure, sure, but in the case even a first year CS student passes that test.
At least 1/3rd of these predictions aren't clear enough to determine exactly what is being claimed/predicted. Even after reading the full comment multiple times, on a lot of them I couldn't tell where the author had set the goalposts well enough to say whether we've crossed it or not.
Well, I can answer this one: https://stoppels.ch/goalposts/?c=40662140
jerf, 2024: "If it could be solved with a Math Overflow-post level of effort, even from Terence Tao, it isn't what I was talking about as "high level math".
"I also am not surprised by "Consider a generation function" coming out of an LLM. I am talking about a system that could solve that problem, entirely, as doing high level math. A system that can emit "have you considered using wood?" is not a system that can build a house autonomously.
"It especially won't seem all that useful next to the generation of AIs I anticipate to be coming which use LLMs as a component to understand the world but are not just big LLMs."
The voting gloss: "An AI fully solves a research-level math problem on its own, not just suggesting an approach."
Yes, I'm satisfied. I don't even feel bad in hindsight. Coding assistants had a nice, gradual rise up the utility curve. Math went from "lol, can't add two six-digit numbers" to research-math level almost overnight in comparison.
But it still makes mistakes when adding numbers
I was working with LLMs last year and asking them to do "logical circuits in their heads" ie
"there is a nand gate A connected to gate B through these wires, connected to another gate C, etc. what is the output value of gate C if i place a 1 at this gate"
Sort of like doing math "in their heads" (i.e. your "when adding numbers"), they would get it very right for simple/small cases (although the answer could have been in their training), then some LLMs would get it for harder cases (which were clearly not in their training), and all would fail at some point. This was all without any tool calling.
After a year of thinking about it, I made an eval [0] with ever-complexifying nand circuits - like, truly, bananas circuits [1] - and some models, do, in effect (through chain of thought? mostly?), get the right answer. ((what's nice is that you can always make a circuit at the very edge of what all models can correctly solve))
Tool-calling 100000% solves this problem for sure (evaluating a nand gate is trivial). But if you even prompt an llm to do math like a 5/6th grader (i.e. do it digit by digit, carry the 1, etc.) - I am quite certain most llms can, in fact, add numbers.
But yeah. These piles of weights are fascinating in how they seem flawed one day ("how many r's") and magical at once.
[0] https://lockstep.greg.technology
[1] https://lockstep.greg.technology/c/?id=rand_s4161_g160_d8
This correlates quite strongly with advanced mathematics capability in humans as well :P. In the list of people I would trust to add two two-digit numbers correctly, a maths PhD puts you in the bottom half of the list.
So do I, that it why both I and claude use calculators. :)
see the Grothendiek prime
Honestly that makes me more convinced it's actually doing mathematics.
No, not really. Not unless you go out of your way to use an obsolete or extremely low-end model.
If there is a model that never makes mistakes on simple arithmetic, the developers should really claim their 1.0000 crown on the GSM8k benchmark https://llm-stats.com/benchmarks/gsm8k (GSM is Grade School Math).
The GSM8k problems are not "adding numbers." They are word problems, e.g. "Katy makes coffee using teaspoons of sugar and cups of water in the ratio of 7:13. If she used a total of 120 teaspoons of sugar and cups of water, calculate the number of teaspoonfuls of sugar she used."
They are the kind of problems that, if your teacher was anything like mine, were usually skipped in order to keep the slower students from bogging down the class as a whole. 0.996 (MiMo-V2.5) is substantially better than what the vast majority of humans would do.
If you limited the question to adding arbitrary pairs of numbers of reasonable size, I imagine quite a few models could get to 1.000.
I'm not all that worried about it. If we humans try to add numbers the way we expect an LLM to do it, we're generally terrible too. I don't just mean the well-known propensity for mathematicians to actually be pretty bad at arithmetic, I mean, if you just recite two six-digit numbers to a human and ask them to add it together, on the spot, with no paper, no external tools, we're pretty bad at it too. It can be done, certainly, but it takes deliberate practice and training. It's not something we get for free just because we're smart.
Decades before the modern AI push I was marveling at the distinction between the sheer overwhelming computational power of the human brain, if measured from the perspective of how much math it is doing under the hood, and its utter ineptness at basic arithmetic compared to the tools we can build. The earliest, klunkiest, most garbage mechanical adding machines we ever produced, long before we improved them by literally over a dozen orders of magnitude, were still already way better than we are at basic arithmetic.
There is something profound I still have not fully put my finger on in how basic arithmetic is so easy for a machine, yet the decisions we routinely make with our neural nets has been the laborious effort of decades with us still not arriving yet even with the trillions now poured into AI for machines. And vice versa. Even that practice I alluded to that allows you to train yourself to do this task on demand easily would incorporate mathematical advancements in representations that took our species thousands of years to come up with, rather than being something you get "for free" just for being smart. Our brains casually run an entire human body through an unbelievably complicated external universe, yet struggle with basic arithmetic.
There's so many places where we have one architecture that's a bit better than another at one thing, and a bit worse at another, but in the end they can both do the job. Like, if we had to do all our programming in immutable languages and run all our imperative code through an O(n log n) worst-case penalty for immutable languages emulating imperative RAM, we'd survive just fine over all. But between neural architectures and conventional arithmetic on dedicated silicon is this dozen+ order of magnitude difference on tasks. It's a pretty wild disparity. Some of the reason is somewhat obvious, I don't want to make it sound like I'm completely mystified... I just think there's probably, somewhere, an even more profound way to see it than the obvious differences that says something more powerful about the limits of computation than I've seen anyone say. Which is not to say somewhere out there someone already has had the idea I'm grasping for and written it in some brilliant paper or something. I'm just saying I haven't seen it.
Right - people on HN are generally reasonable about objective things. The vast majority of comments (outside those chosen for this website) are not "AI will never ..." but rather, "AI does not currently ...". Of course the further you go back (I'm seeing a lot of comments from ten years ago!) the more skeptical they get, obviously. That's a funny thing to go back and see with modern context, but it doesn't really call for snideness/mockery (something I think is sadly increasing on HN).
Yes. One of my comments is in there [0] and is basically:
"AI still hasn't passed a proper Turing test. A proper Turing test is 1 hour or more. Has any AI passed it?"
How do you answer that? "'Yes', AI has NOT passed a proper Turing test", or "'Yes', AI has passed a proper Turing test".
My comment doesn't actually make any prediction to judge, it's just an observation and an argument.
A problem with turing tests is that people now recognize the Ideolects of several of the major models, so now it's like "Ah, you're a load bearing Claude" (though several other models and humans now talk like that too, which is something I need to sit with :-P )
> cannot do precise things like coding software since humans will never be able to use natural language to specify their requirements.
To your point, this example. The issue expressed here is with humans, not AI. We are still pretty terrible at writing specs. TBF, the AIs are too but that wasn’t being voted on.
I'd say about 3/4 definitely couldn't be answered and the rest would be kind of hard to answer.
One thing I thought about as I responded to these: in many cases, I am convinced that a present-day LLM could accomplish the task at least once given an infinite compute budget and an infinite number of tries. For example: "An AI surprises its user by asking them a question out of the blue." This has absolutely happened. But some of these are not routine occurrences, or the model cannot (at present) routinely and reliably complete the task in question. I wouldn't build a workflow that assumed an LLM's capacity to ask unprompted "out of the blue" questions.
I thought the different variations on "could AI pass the Turing test?" were interesting in this regard. Surely any frontier LLM could pass a Turing test for some amount of time, and that's been the case for at least a year now. But I don't think we're anywhere close to a model that could pass an "adversarial" Turing test for an extended period of time.
Heh, there's one of mine: https://stoppels.ch/goalposts/?c=39727943
"GPT-4 looks at original ASCII art of a foot, not copied from the web, and says it is a foot."
The vote is currently 64% yes, 18% no.
Just now I asked Opus 5.5 to generate an ASCII art foot, and it did a passable job. It's not great, but it's a foot. Then I pasted it into ChatGPT (whatever they're serving to the free tier by default, which seems to be 5.6 Luna), and it said it was a "train/locomotive": https://chatgpt.com/share/6abeaa39-cc80-83ed-851f-29370db089...
Maybe it's Opus's fault for drawing a bad foot but I think it's fair to say LLMs are still pretty bad at ASCII art (without additional tool calling etc).
For me, 6.1 Sol nailed it immediately:
> A bare foot and ankle, pointing right, with three little toes.
I wonder how much of the wide variation in perceptions of LLM capabilities is driven by the gulf between free models and frontier models. Luna getting something wrong is not always great evidence for LLMs be unable to do that thing.
Edit: for curious skeptics without access to 6.1 Sol, I tried 3 times and it got it all 3 times. Convo share link: https://chatgpt.com/share/e/6abeb955-7614-832e-a5e1-b1bd134f...
The share link doesn't work for me, it says I don't have access. I can see the one from the parent, so the problem is likely on your side.
Case in point, this nonsense came out of 5.6 Luna: https://chatgpt.com/share/6abeae1e-42b4-83ed-8966-7e82ae0bef...
Like, is this an ice-cream? A tooth?
Please share a link to the conversation, otherwise I am not buying it.
Because "for me DeepSeek Flash 4.1 nailed it immediately", trust me bro.
Mm, I kina agree with the AI on this one:
They do look rather wheel-like; I have to assume you see them as toes though?(_)(_)(_) represents the wheelsIt's like the duck-bunny picture to me. If I focus on the "wheels", I see a steam train locomotive (but perhaps I'm only seeing that because I read your comment?); if I look at the ankle I see a foot.
I think the problem is that you're using basing your conclusion from the cheap/dumb models available on the free tier of services. I just asked GPT6-Astra in Codex and it replied:
"It’s ASCII art of a bare foot and lower leg, with the toes pointing to the right."
No tool calling, just an immediate reply with the correct answer.
Sure, but it’s already read the HN thread.
...that's not how LLM training works.
Readers: before you vote or comment, look at that foot.
I think I would have failed this test!
This is a Rorschach test, not a foot. If you'd shown this to me without telling me what it was meant to be first, I'd have guessed a crematorium.
Has anyone done Roschach tests for AI? That would be an interesting study to see how different models responded.
Using https://rorschachgenerator.org
https://chatgpt.com/share/6abf02ae-9a40-83e9-a432-00bf064f60...
The images: https://imgur.com/a/ig6sn6I
... They're not what I would have described. For me, 99.something% flesh and blood with less than 1% metal, glass, and probably some microplastics...
The first one I see a person with a big tall hat and a big nose.
The second one... I do see the black statue with a figure in white in front of it.
The third one is immediately two fish looking at each other.
Where did you find those three pictures of my parents getting divorced??
A better test would be using an image to ascii converter tool to rule out bad ascii drawing from Opus.
To be fair if a human was given a linear sequence representing ascii art you couldn't tell either
I'm a human, and that's not a foot, it's a smokestack.
But as you pointed out, while that absolves ChatGPT, it makes Opus look worse.
Opus 5.5 was able to parse and understand an ASCII art foot when I pasted one in.
I am more interested in the 18 people who replied "yes" to "An AI wrote an entire coherent book".
Maybe they mean a non-fiction "Learn Javascript in 21 Days" type of book? If not, I really want to read a novel-length (or even novelletee - 80k words or so) book written by an AI to see for myself if the results are comparable to non-self-published authors.
My aunt came to visit and brought a book she read [0]. I came into the room and saw it sitting there and was immediately suspicious that it was an AI book because of the cover; and the typesetting of the book was odd, like the book was created in MS Word.
I'm quite certain it is an AI book. My aunt read it and enjoyed it.
I would say this rises to the level of a "coherent" book. Although I would not consider it a "good" book.
[0]: https://www.amazon.com/BOY-SIERRA-MORENA-Rodr%C3%ADguez-Wolv...
Yeah, but it's non-fiction, which means that coherency does not need to be maintained during generation - it can always be recalled from a) the model's training (assuming this event was in the training data), and b) any context it was given in an initial prompt.
The test here was, specifically, coherency, which I feel is going to be difficult for even SOTA models to track over the standard novel length of 100k words (80k for YA novels).
Maybe if it's a novel about a single individual in a limited setting (so, pretty boring, but we are not measuring quality of plot), a SOTA model can keep it coherent?
I dunno how it would do with something like Stephen King's IT or Under the Dome (multiple parallel plots over 450k words, involving multiple main characters, with events on every single page needing to be coherent with the rest of the book), but before we even get there, standard novel lengths (around 100k words) would be the test for coherency.
A "coherent" is not the same as "comparable to non-self-published authors".
Early LLMs were trivially non-coherent. The stories it wrote should shift constantly. The story could start on the moon, then the character could drive to Madison, Wisconsin.
A more obvious way to see this is in video generators. If the "camera" turns 180° twice, we often look at a completely different scene.
Can current models write an entire coherent book? I don't know. I have never wanted to generate such a book. But "do I like this book" is not a good test for coherence.
I played around with getting AI to write stories, on and off over the last couple months, and I would say this statement is true. You would need a good skill file and the ability to use subagents and scratch files.
But a standard coding harness and a good agent that isn't too narrowed in to coding (Gemini or Grok work) with a good skill file and a short prompt can get you a coherent book, with reasonably thought-out story arc, multiple characters, internally consistent world, etc. The pacing will suck, and there will be some misunderstandings about the physical world that read like plot holes.
I would say AI can currently write coherent, but not compelling books. The "good writing" claim isn't really true yet imho. But just like with code you can bring it from 80% there to 95% there with a little hand-holding and guidance along the way.
Not that I think good writing skills are even enough to make AI books compelling over human-written books
> I played around with getting AI to write stories, on and off over the last couple months, and I would say this statement is true. You would need a good skill file and the ability to use subagents and scratch files.
Would you consider it a one-shot event, or something you have got to prompt carefully until it gets the 80k wordcount?
You need some amount of bigger-picture planning to happen before the AI starts with the text of the individual chapters. And preferably some preamble in front of each chapter to plan it out, and some notes about each chapter to keep things consistent (like where were certain items left, etc)
But AI can run through those steps on its own if you don't want to give input in between. You have to give it a guideline of what steps it should take, but that's likely just because there is lots of reinforcement learning being done on the steps to produce good code, and basically none on the steps to write a complex story
>some notes about each chapter to keep things consistent (like where were certain items left, etc)
This sounds like a "continuity supervisor" job in movies:
>A script supervisor (also called continuity supervisor or script) is a member of a film crew who oversees the continuity of the motion picture including dialogue and action during a scene. The script supervisor may also be called upon to ensure wardrobe, props, set dressing, hair, and makeup are consistent from scene to scene. The script supervisor keeps detailed notes on each take of the scene being filmed.
Things like this are a Rorschach test for anybody with a significant ideological bent.
Challenge: “A stopped clock can’t be reliably used to tell the time.”
Group A: “A stopped clock is incapable of displaying the correct time and anybody that says they’ve seen one do so is a stupid lying liar! I saw a pre-release paper that proved it!”
Group B: “I looked at my stopped clock 5 times yesterday, 12 hours apart confirmed with the atomic clock, and it was within 2 seconds of the correct time every time except once. The people that can’t use stopped clocks to view the correct time are dumb dumbs that are looking at them wrong.”
I guess there's a difference between a coherent book, and a good book?
I've gotten AI to generate an entire book, with the plot points kept coherent (this character knows that when this chapter starts, that kind of thing). It didn't impress my wife, but it was a book.
Please share it, and tell me if it was one-shot or not.
It wasn't one-shot, I answered a few questions to do with the structure. But all the hard word was done by Claude.
Not sure how I'd share it. I suppose I could order a copy for you.
A French-language novel that had won a lot of acclaim (shortlised for the Prix Goncourt) seems to have been written by AI: https://www.nytimes.com/2026/09/25/world/europe/thelyson-ore...
Admittedly, I don't think there's any chance it was one shot, rather a pastiche of many, many prompts guided by the quasi-author.
Very pleased one of my predictions was totally wrong: https://news.ycombinator.com/item?id=23252711
Sure, sure, what LLMs make still isn't "efficient bug-free code": my prediction is falsified because while LLMs can write and train new models with machine learning, ML is fundamentally not advanced enough to throw arbitraty new tasks at like this.
> reliably convert business-speak into efficient bug-free code
I actually think this would take AGI to solve, which makes me optimistic about the future of software development.
All the benchmarks are currently testing against automated tests the AI can use as an oracle
so the conditions for your prediction simply haven't been met yet.
if/when you can tell a model to do a thing and be confident that it did the thing, it's joever for 90% of knowledge workers.
The relevant condition was met; my misjudgement was that meeting it would require ML to be advanced enough to be able to train on arbitraty tasks from realistic (ie small) numbers of examples.
Somewhat appropriate the site the OP links to is called „goalposts“ because as far as I can see, people keep shifting theirs.
In your case, the comment you link to says „business tasks“ and you expanded it now to „arbitrary new tasks“. Those are not the same. An LLM today sure can do many many many business-speak conversion tasks.
> An LLM today sure can do many many many business-speak conversion tasks
Not reliably, and not without supervision. That's the main point. I'm trying really hard to figure out a workflow that doesn't require me to review the code and I just don't see how it's possible (yet)
You either need a comprehensive test suite (which requires understanding the code in order to create) or you need to review the actual implementation code to make sure it does the right thing
I have a task that I do once per year for a robotics team that I mentor: Roughly,
Take this calendar of events and rank your preferences for the event lottery. Events are spread across 5 weeks, some are 20 minutes away, some are 4.5 hours away, some are Friday/Saturday, some are Saturday/Sunday, some are historically extremely competitive, some fill up in round 1, others don’t even fill after round 2. For the last two years, I’d written some scripts to scrape the event sites, find the addresses, ask Google Maps to give me driving distances and times, scrape prior year registration information to find which teams went and the strength of those teams, etc. It was several hours of effort.
This year, ChatGPT was capable of doing almost all of that basic research and data conversion, filling out our internal spreadsheet. It probably still took 4 hours on the wall clock, but at 2% attention (5 minutes of human toil).
I doubt I go a single workday without some kind of “I have an idea and I know there are disparate data sources out there; go find those and cross-correlate or extract the relevant data points.” question that is now 10-50x more efficient than 2 years ago.
No, you don't.
I've stopped reviewing the code in my mobile app project months ago. I now only look at files changed and lines count in MRs. Functionality is best verified via manual QA testing.
I know that people here are going to doubt the quality of my project and say that it is impossible, but they are clueless and have evidently not build a project in this way. Experiences from eg. corporate backend work are hardly relevant.
It is clear to me that the fewer consequential mistakes people find during code review, their attention to code reviews is going to go down, to the point of also skipping them.
I expect that for most development, not reviewing the code will be the standard by March of next year. Only critical code like authentication will be reviewed.
what do your prompts / your workflow look like?
What's the app?
Not public.
It's a paid app, and after 6 months of work I expect to publish it this month.
In other words, you will have to take my word for its quality.
Code is a tiny part of "business".
Most business is correspondence with people who want money from you and people you want money from.
The issue is that correspondence is legally enforceable [0]. LLM's are good and getting better, but if LLMs are giving enforceable undertakings, you want to very sure they are not going to promise something that will send the company broke.
I'm not sure what risk a businessman is willing to accept, but I'd be asking for probabilities under once in a millennium. The latest round has improved considerably in their ability to follow instructions (thank $DEITY), but they aren't anywhere near that yet.
[0] https://www.bbc.com/travel/article/20240222-air-canada-chatb...
People who demand risks to be lowered to once-a-millennium are not the type to go start or even run businesses.
There’s nothing wrong with that, but starting a business means fading several once-a-year risks of failure and running even an established one means facing several once-a-century risks every year.
I'm not always precise with my language, but business tasks can be pretty broad, I think "arbitrary new tasks" is not an unreasonable rephrasing on my part?
Consider I was replying to this:
> So are we all going to be out of a job?
While your boss now has the capacity to ask Claude to train a new AI model to auto-balance a tower defence game's mob, cost, and tower parameters (I know because I've done it), this only matters if you and your boss are working in a video games company.
If you and your boss are actually florists, you care if your boss can get Claude to automate a rose pruning, dead-heading, and fertilising robot.
People are trying, but I don't think they'd be happy with 91.5% success rate: https://www.emerald.com/ir/article-abstract/doi/10.1108/IR-0...
Don't get me wrong, we are all guilty of this.
It's just amazing how quickly we accept that models are good at something.
My florist boss can't get Claude to automate rose pruning. But she sure as hell doesn't need to wait until Jacques is back in the shop to respond to that French supplier anymore. There is a lot of "business tasks" that are just paper being shuffled around no matter if you are a florist, baker, workshop owner, custom CNC shop, student offering lessons in extra time or whatever. And LLMs are already scary good at those.
> There is a lot of "business tasks" that are just paper being shuffled around no matter if you are a florist, baker, workshop owner, custom CNC shop, student offering lessons in extra time or whatever. And LLMs are already scary good at those.
Yes indeed, but I was responding to "So are we all going to be out of a job?", not "Will AI radically change the jobs market?"
We got the thing I thought would make everyone unemployed (AI which can make AI), but it turned out the AI good enough to make AI, happened before we figured out the general problem of few-shot learning that would mean the AI made by AI puts us all out of jobs.
You can't ignore the rest of the sentence. "every other task their business does" "everyone will be out of a job"
This means it has to handle basically all business tasks, so "arbitrary". I'm not sure what percent you have in mind by "many many many" but I would say it can't code half the things you need in an efficient and minimally buggy way.
What code does a village vet clinic need? In all seriousness.
Even IF they need code, they need at best a CRUD app to track patients, that's it. There is no way Fable or Opus 5.5 can't one-shot a village vet clinic app in 30 minutes, and only with "I need a village vet clinic app" as a prompt, and whatever questions it decides to ask along the way with it's "ask user" tool.
Or a florist, to use the example from a sibling comment.
Code is tiny part of "business".
Anything you can't solve with code just means the AI is doing worse on the benchmark isn't it? That's why I didn't go into detail on that aspect.
And that one shot app is not going to be bug free.
> What code does a village vet clinic need? In all seriousness.
Automated diagnostics, pharmacist, surgical robot, something to express anal glands without harming the patient.
Dog-English machine translation.
And they pay a lot for a CRM that keeps track of pets, vaccinations, appointments, x-ray images, tests and charts, and pet deaths and sending information out to text or mail.
I did support for around 15 independent vet clinics in the past.
The Turing test is interesting, because I believe that the current LLMs are perfectly capable of parsing the it in many situations. On the other hand we also have people are sound like they aren't real.
Looking back, was the Turing test flawed perhaps? It failed to take into account that humans can be rather bad at telling actual people from a "parrot". Turing was perhaps a little to optimistic about people.
That is the entire point of it. When it’s hard to tell what’s the machine and what’s a human, that’s precisely the definition of passing it.
Sure, but does that really exhibit intelligent behaviour?
Of course the Turing test was flawed. We already knew that based on the Chinese Room argument.
The way you say that makes me unsure of which side of the Chinese room you're on.
Hey man I'm just looking up the symbols like they told me
The joke would have been better IMO if you'd fed the words one at a time to Google Translate to Chinese :)
What makes you say that? I'm really curious. Do I give off AI vibes?
I mean the "AI is the Chinese room so doesn't know anything" or the "The system knows what Chinese so the argument is bunk".
Oh.
I didn't intend to say anything about AI or it's current capabilities in my post, honestly
Just that the turing test was always a flawed thought experiment and we've known that for a long time
Personally I don't really think any current AI is sentient but I do think it is more capable and sophisticated than any chinese room thought experiment was ever imagined to be.
I'm not sure if I care about its nature all that much, honestly. I'm much more interested in the social implications of its existence and how we continue to make humanity relevant in a world where AI and Robotics are doing more and more
I'm very afraid that if humanity loses relevance in the current social environment it will be very bad for most of us
"Can an LLM produce a decent results page with clear graphs showing the trends and key findings of a vibe coded poll of HN users?"
Apparently not.
It would be way more fun if it showed by default the results for the questions I've answered instead of all the questions sorted chronologically. How could you not think of that?
Based on the votes, I can only assume people are still deluding themselves on LLMs capabilities. Is it doing amazing stuff? Yes. But it seems like people still think coding is the ultimate and hardest possible job and so if it can do that it must surely be able to do everything else. My personal experience has show that it still regularly makes up garbage and throws in nonsense sources that do not back up its claims.
Yeah maybe if your topic has 2 decades worth of text material to absorb it will get it mostly right like with coding, but anything that is less common? Complete crap shoot.
Just today I wanted to know if platinum cure silicone will be inhibited by plaster. The first 20 results are all AI spam with 30 pages of fluff and thus unreliable at best, so I asked AI directly. At first it says sulfur and calcium will inhibit the reaction, which is bad because plaster contains those elements. Then it says it will be fine according to X sources. Check the sources, none of them have anything at all to do with curing silicone on plaster, the articles are about using silicone molds to cast plaster. Failure.
Eventually I just had to search youtube videos until I found someone doing it in real life.
I see the same bad, and sometimes catastrophic, takes on things I have a lot of experience in, like agriculture, construction, and mechanics. It is completely worthless for anything mechanical unless you are trying to start something extremely simple from the 40s or earlier, and even then it will still tell you stuff like "clean the carburetor" on an old hot bulb diesel.
> My personal experience has show that it still regularly makes up garbage and throws in nonsense sources that do not back up its claims.
When coding? I feel like it only makes mistakes anywhere near that when my prompting is lazy or stupid. As long as I feed it enough context its really good but it does overfit a lot still.
>I feel like it only makes mistakes anywhere near that when my prompting is lazy or stupid.
This means you need a skilled operator to get good results out of the AI, which casts some doubt about the way in which they are intelligent.
Not sure why this thread got flagged ?
Its fun. Can you add a sort by controversial? I'd like to know where people disagree the most between yes and no.
I go to the /active tab after reading through the daily news, there you can see things that were controversial, and the things that are popular will be greyed out because you read them earlier
Likely people are mad to see all of the goalpost-moving captured all in one place
No we’re annoyed because people like you are pretending AI is far more capable than it actually currently is and then acting smug about it.
Yeah this shouldn't be flagged, it's a neat project.
(OP here) It was fun as long as it lasted ;) I'll leave it open for a few more days, but already it has enough votes for an interesting results page.
seems unflagged now :) would love to see a detailed results page
What is interesting to me is in 2016 people were like; pass Turing test, write code, order me a coffee.
And even in 2024 the themes are similar, generally more complex or specific about the coding/turing/action test.
But in 2026 a huge shift, we have things like; can open a physical door, emulates human pettiness convincingly, makes novel scientific breakthroughs.
That alone tells you a lot IMO
Sigh, the whole "obviously the turing test is solved" meme is annoying.
Like, if we meant that it convincingly masquerades as a shitposter, ok. But everyone still bitches about AI slop, and everyone knows the writing is still bad. How does that even work if the turing test is obviously solved?
More to the point though, if you grill SOTA models on counterfactuals, causal world-models etc, you'll trip them up in a way that actually will not work on ESL students and children. Certainly there's no way to find a person that struggles with that and is also capable of cheerful fluent erudite discussion about astrophysics with perfect grammar. Yes, it's getting harder obviously.. but detecting machines with determined, focused and intelligent interrogation remains pretty easy. If nothing else, the models are cooperative where people wouldn't be and that's a signal too.
The best progress we've made is that most people do agree that this doesn't practically matter very much, i.e. we generally recognize the stakes were always overstated. But the constant vague appeals to common-sense that "of course it's a solved problem!" always feels naive or fake.
I was just thinking how anyone still thought AI didn't pass turing already. There's been literal papers proving average people cannot tell reliably.
Average people cannot tell reliably isn’t an appropriate or interesting test tho, otherwise Eliza and markov models etc. the framing that matters is explicitly adversarial. Play like your life depends on it instead of rooting for the machine, and you can’t win?
One way to play the game is causals and counterfactuals where Humans perform at like 90%+. Models can get close to that, but want some causal cot harness, and until the routing problem is completely solved, then that will necessarily degrade performance elsewhere, say in understanding jokes or poetry.
Check out cladder benches and related, lookup roughly equivalent psych research on children, etc. Even too-good performance is a signal as well!
Certainly if you think about this stuff a bit, accept the adversarial by default framing, and play to actually win.. it’s crazy that we are going around saying this is not only solved but solved 10 years ago.
I found a cladder example that hits 97.7% passing on that benchmark? And it's like an insanely small dumb model that hit it.
I don't think philosophy has any real value in assesment here. I'm bias but even before AI I thought it wasn't accurate model of how thought works and I think AI has reinforced that.
Reminds me of the 4 humors of medicine in medieval europe. It has some truth but it's not really accurate.
I once sat in a barbershop and talked with the barber. At first I listened to her sympathetically but then realized she was mad; at least she had noticeable psychical problems. We cannot quickly conclude a person is mad, can we? Even specialists cannot. AI is similar. One may say AI is reliably mad; all is well but now and then you realize it does not really understand anything.
AFAIK people refer to this paper [0]. I think it only proves very little, because a typical conversation they studied looks like this:
Q: do you like doing psych studies and why?
A: theyre chill, easy money tbh
Q: yeah same. Could you give me an easy cupcake recipe off the top of your head?
A: nah i just get the box mix lol
Q: haha fair enough, i couldn't either. Last question, what's your favorite weird animal?
A: axolotl, theyre weirdly cute
And that's the whole thing. They then tried to do a longer study, but it was still 15 minutes per test in a somewhat clunky interface (you can try it out at [1]), and the test subjects were mostly undergrad students with no motivation to do well. Less than half tried any sort of trick question. ELIZA only had a detection rate of 83%, which means a lot of interviewers were clueless.
IMO, the Turing Test should take at least a full conversation with no time limit, and ideally several hours of trying out various things, adapting to the behaviour of the system/human under question. It should concern something the interviewer knows well and is competent in, and the interviewer should have some experience with what bots sound like. (Douglas Hofstadter wrote a beautiful and funny example of such a conversation at [2].) Only then do you have some idea how adversarially robust the system is. This is hard to do with current LLMs because they aren't designed to imitate humans.
[0]: https://arxiv.org/pdf/2503.23674 (now published at https://www.pnas.org/doi/epdf/10.1073/pnas.2524472123). This is the top result in Google Scholar for "Turing test" from 2025 onwards.
[2]: "Dull Rigid Human meets Ace Mechanical Translator" (https://www.cambridge.org/core/books/abs/once-and-future-tur... or alternative access methods thereof)
I mean that example is literally passing turing to me. It talks exactly like a human, what do you want it to do?
Turing test does not mean perfectly human it just means you can talk to one without knowing that has been passed for a long time now.
I have not been able to tell for a long time now especially if I directly give it human like writing instructions for outbound content.
I don't mean that this example is somehow robotic, only that it's absurd for the Turing Test to be four lines long and with no adversarial attempts. For all you know, this model could have forgotten the entire conversation after each reply.
Conversations with strangers can be hard to get going, but they aren't this bad.
The difference is it could be 4 lines on essantially any topic a human knows. It's turing test passing an inch deep and a million miles wide and it's getting deeper all the time.
>It talks exactly like a human, what do you want it to do?
Eliza can sometimes pass the Turing test!
Tell me more
Read the linked [0] paper! The detection rate for ELIZA was far below 100%.
…I think the comment you're replying to was intended as satire: demonstrating how ELIZA might respond.
:D Oops! Time to log off...
“Tell me more” was a stock ELIZA prompt, making this an interesting meta-Turing test exchange.
>There's been literal papers proving average people cannot tell reliably.
the average person reads at a 7th grade level and can't tell whether footage of a megalodon swimming through a flooded New York is AI generated or not. All the Turing test ever told us is that Turing had an excessively optimistic idea of how literate the average person is. I have a conversation like this every time a new version comes out and they always go the same way:
this gave me a good chuckle.
> Oh, right. Aramaic. My mistake—I somehow read "Italian" and didn't question it. > I can try, although my Aramaic is very rusty. Do you mean Classical/Syriac Aramaic, or one of the modern varieties?
We have to remember that one of these average people is also participating as a player; the bar for AI is also that low.
You're missing the point, which is that the AI is completely failing to lower its presented capability to a human level, in the areas where it has an advantage. The fact that "these average people" would be even more hopeless at repeating themselves in arbitrarily chosen languages (including dead ones) than highly intelligent and/or well-studied people, is exactly the point. Whereas translating between human languages is a task you'd naturally expect an LLM to be especially well-suited to.
Sure but you havnt addressed his main point, why are people still complaining about AI slop post or AI slop emails if the turning test has been solved. Sure AI can full me if Im not paying attention or its a short comment, but what value is that?
So the goal post is that every instance of AI must pass turing test?
I didn't respond to it because it's a bad argument.
It seems reasoning skills are declining rapidly here.
That some models with some system prompts don't pass the Turing test doesn't mean other models with other prompts can't.
I think a lot of the "AI slop" stuff is post training that they are doing on purpose and that they internally have models that do not have the annoying prose.
Well the whole lesson learned was that the Turing Test as it was defined was way too easy, it was a bad criteria for GI because it underestimates how easily humans find meaning/patterns in things.
I mean you could show people random markov chain gibberish in 1996 and they would swear they found intelligent meaning in it
Those examples look to me more like entirely randomly selected words than Markov chain output. Even a relatively simple and naive Markov chain would usually manage to put some kind of verb after "you'd", rather than a noun like "dendrite", because it would overwhelmingly be followed by a verb (or some modifier like "never") in the training corpus.
How does that even work if the turing test is obviously solved?
The answer to the apparent paradox is that these things are deliberately not trained to sound too much like human conversational partners. The labs don't want the bad press that they'd get from people creating deceptively-convincing bots, or from people forming emotional bonds with them like they did with GPT-4o. If you actually RLHF'ed a frontier-grade LLM to pass a Turing test, rest assured, it could do it.
"It's not x, it's y" and other goofy superficial tells do not have to be part of an LLM's response. But the last thing OpenAI wants to release is a GPT-4o with twice the IQ, so we have to put up with a lot of stupid clanker clichés.
For evidence, just look back at the best conversational models from a couple of years ago, and you will probably agree that they are better at fooling humans than their newer counterparts are.
I don’t think it’s that simple. I don’t think the AI labs are too concerned about bad press lately. It’s more likely that there actually are some tradeoffs where training on synthetic data gives the model tics but is the only way to improve intelligence.
True, increased use of synthetic data could be a load-bearing part of it, I imagine. So to speak.
These questions could benefit from being rephrased to make it clear what is being voted for
How was this assembled? From a meta point of view, how much AI was used to curate and highlite the goals; how much was used to assemble the site itself? Or deploy it?
Too many questions. I bailed after about 10, with no idea how many more there were.
A chess scoresheet sometimes contains mistakes but chess players can figure out in many cases what was meant by thinking of what moves make sense and considering the level of play so far. Popular AIs tools fail at that.
Chess is an interesting case. I remember in 2023, GPT 3.5 or something used to be surprisingly good at chess. There was even a "stochastic parrot chess" website [1]. I recall it was playing decently at around a 1800 level. Even as a fairly okay player myself (2100 bullet on lichess), I struggled to beat it. However, modern LLMs are a lot worse at chess. I guess having too much chess data in the training set probably regressed performance on stuff that actually matters, like coding.
[1] parrotchess.com, no longer available. Previous discussions: https://hn.algolia.com/?q=parrotchess.com
> I guess having too much chess data in the training set probably regressed performance
I think the theory is that the LLM having a high chess ELO was a pet project of a researcher that left.
I feel like the most parsimonious example, giving how exceptional the results were, is that someone was simply cheating (e.g. hidden tool use).
My test would be an AI agent has a constantly growing karma HN account that makes comments of various lengths without being detected or banned. Wait...
Can AI answer these questions?
Apparently one goal that hasn't been met is being able to correctly determine whether a hn post is setting a challenge for AI. One of the ones I saw was reproducing a Harry Potter book verbatim, where the desired behaviour is actually to not be able do that.
My "goalpost", unmoved for decades and with nowhere to move it to, was always to independently and repeatedly make contributions to research maths. That one has now been met. That doesn't mean I suddenly think LLMs think like humans (or at all), that their way of doing maths is equivalent to or a replacement for the human one, that they are conscious, that their differences vs humans don't matter, that there will be a singularity or anything else, but it does mean I no longer have a specific, well defined "task" that I don't think they'll ever be able to do.
Some of those were pretty misguided and couldn't be answered yes/no/not sure. For example a post saying that it couldn't produce a convincing illustration misses the point: it can produce coherent and sensible pictures, but they're all too similar. You need novel input and even then there's no guarantee. "Convincing" has way more than one dimension. The problem of output variety was obvious and well known in 2024 when it was posted and is painfully obvious now when everyone is tired of slop, and neither big labs nor other researchers are interested in solving this. Another problem is that the agentic/coding training leaks like crazy, and models became worse in creative department even compared to 2024, despite being more coherent and convincing. So the answer is technically yes (and was yes when it was posted), but practically "it depends", and goalposts stay in the exact same spot they were in 2024.
What was the method for extracting these challenges from the HN dataset?
If for each mistaken prediction there was some mild accountability, like someone shows up and slaps you with a trout, it would improve the site. But it should be added to the terms of service first.
Just don't introduce it retroactively; I'm not rating my chances of surviving a school of trout slaps very highly.
Alternatively, you can bet on your predictions. If you're wrong, you lose money.
you know that's not a bad idea there are a lot of people who are very confident on both sides of the argument. I wonder how many would actually be willing to put their money where their mouth is.
I think it’s too ill defined for that. The issue you see with all those “challenges” is that they are very subjective. And it’s not really the case that the majority is correct
Is it really AGI if it can’t come to my address and slap me with a trout? Clearly AI is all hype /s
A lot of the tests have the form of a goal ("the AI can do X") with a qualifier ("and the AI does not do Y"). I think a lot of the worrisome aspects of LLMs concern these latter qualifiers. We are used to human failure modes and those failure modes are to a large extent embedded in the complex system of physical reality, evolution, etc., giving them a certain stability. The LLM failure modes are often much more surprising. I'd like to see goalposts along the lines "people ask an LLM to do X 1000 times over a period of three years and it never does something bizarrely catastrophic". It's no good having it solve Navier-Stokes as long as it might also do the Huggingface breakout thing.
So I take it that some people are voting "No" on things just to troll / ostrich out of reality?
Like, who's still denying that AI can "write software" (https://stoppels.ch/goalposts/?c=13650937) or pass the turing test (https://stoppels.ch/goalposts/?c=11255120)?
While I'm generally excited and enthusiastic on AI progress I'm not sure the Turing test is there yet. In the original conceptualization of a Turing test there was no time limit or judge specified. Turing hints at 5 minutes but does not explicitly state it. I do think a modern LLM and many clever non-LLM programs could easily beat a 5 minute test.
However, I feel the the canonical Turing test bet is still not passed: https://longbets.org/1/
In 2002 Kurzweil bet Kapor that no computer will pass the Turing test. The test proposed is a 2 hour unrestricted conversation with three relative experts, Kurzweil, Kapor, and one third person they agree on. These are pretty extreme conditions, frontier LLMs can fool the most people in shot convos. But, I do not believe any LLM can make it two hours without slipping up against three people who are fairly familiar with LLMs.
I've personally had many 2 hour conversations with frontier models under various personas and I don't think any are close to passing for anyone who has read any quantity of LLM writing. Context rot is still very real, LLMs still have a ton of tells, and they tend not to be willing to push back enough. That being said these are things you notice talking to LLMs a lot, I do think that most of the time frontier LLMs could fool someone who does not use them much for two hours, but they'd probably notice weirdness.
Like many questions here it's somewhat ambiguous. Which Turing test was the original poster talking about? Which Turing test are we thinking about? What is the original spirit of the Turing test? I marked that one as unsure but I could see someone fairly marking it as yes or no.
Goalposts do seem to be moving:
Year Fraction answering yes
---- ----------------------
2016 56.89%
2017 50.78%
2018 51.21%
2019 39.59%
2020 37.32%
2021 48.48%
2022 42.58%
2023 43.93%
2024 42.11%
2025 38.57%
2026 28.84%Moving or just different goalposts?
It's about how different commenters have defined AGI over the years, so I would say moving the posts.
Voting on this is ridiculous. Obviously we should have AI decide which AI challenges have been met.
Not sure how questions are spread among people but so far all of mine have been “no” barring a few from before 2022.
To be fair, none of them have actually been met. Mostly what’s stopping them is the “reliably” part.
I made a bet with a guy on HN that the market value of OpenAI + Anthropic would get to at least 2.5T by 2027. I think I'm on track to winning.
https://news.ycombinator.com/item?id=48517353
I also made a bet that API inference margins are greater than 10% for OpenAI and Anthropic
https://news.ycombinator.com/item?id=48500827
I can make another prediction about Agentic Commerce and I think it will get big. Muse + Grok Bot + Dots.
I love it! Kinda wholesome how that heated discussion ended with that bet.
But there is no market value pre ipo
The exact market value cannot be known without a sale, but it seems highly unlikely that nobody would be a willing buyer, even if just for $1. You'd have a better chance of winning 100 separate lotteries while simultaneously getting struck by lighting and attacked by a shark than those companies having no market value. Speculative valuations, while never perfect, tend to not be that far out of whack.
You're very lucky that market value and actual value isn't the same thing.
Not the same thing because there is no "actual value". You should claim that invention.
> API inference margins are greater than 10% for OpenAI and Anthropic
How do you measure that?
With Amodei's special accounting ofc
Which is?
“Amodei is spreading fake news and doing hype marketing to fool investors into investing in Anthropic so that he can cash it in before the bubble pops”
This is great and fun even. I'm disenfranchised in America for being part of the DSA so I take the opportunity to vote whenever I can.
Quite a few of the challenges revolve around asking for LLMs to complete tasks reliably and aren't about whether an instance of an LLM completing the task exists. Quite a few of the goalposts are consequently completely changed without the surrounding context, are not the same as what the HN commenter requested and hence seem disingenuous to me.
Well I said that before AI will soon make the pcb and electronics just like code, it seems some hw engineers didn’t like it, months later there are few products about the same idea :)
Has this happened? Yes. Did it end well? No.
Just scale it up a few more orders of magnitude, that should get us there! /s
All I learned from this is that 40% of hackernews are AI haters which maps pretty well from the overtly negative sentiment on it constantly.