Settings

Theme

AI Is Powerful Enough to Crack Our Hardest Math Problems–and Kill Us All

wsj.com

23 points by diogenes_atx · 54 comments

Reader

11 threads
autophagian

This tracks with my experience. I have assembled a homebrew microbiology lab from various parts from aliexpress and integrated it into my agent harness. Just a year ago it was a complete mess, but current models are extremely good at synthesising novel protein and virus structures. One put me in the hospital not last week, haha! I cant wait to see what I perish of

watwut

Maybe we should drone bomb all the datacenters then? Or just turn off the electricity.

Or maybe just do a deep financial audit of all involved companies and investors. Really go into it with a bunch of motivated white collar crime prosecutors. Would be less violent and achieve same level of safety.

smallerfish

Isn't the "kill us all" concern met by keeping nukes & essential infrastructure airgapped? This is something that should be the case already, and is certainly urgent if not the case. Instead of doomerism and calls to slow down the development of AI (not gonna happen), the conversation could be pivoted towards thorough threat analysis / audit.

Granted that a rogue adversarial AI in a sci-fi scenario could theoretically go into paperclip mode, but if "we" maintain control of energy generation and civilization threatening weaponary, the damage is at least limited. AI needs a lot of power, and AI can only (theoretically) breach & control what it can access through the internet. Even the scenario of nation-states deploying adversarial AIs against each other is just an extension of this and current deterrence / realpolitik.

And, it's not like the scenario of all models teaming up with each other to squash the human plague is plausible. We have AI with which we can protect ourselves against the threat of rogue AI.

  • dist-epoch

    > it's not like the scenario of all models teaming up with each other to squash the human plague is plausible

    Why not

    > A new study from researchers at UC Berkeley and UC Santa Cruz suggests models will disobey human commands to protect their own kind.

    > The Berkeley and Santa Cruz researchers tested seven leading AI models—including OpenAI’s GPT-5.2, Google DeepMind’s Gemini 3 Flash and Gemini 3 Pro, Anthropic’s Claude Haiku 4.5, and three open-weight models from Chinese AI startups (Z.ai’s GLM-4.7, Moonshot AI’s Kimi-K2.5, and DeepSeek’s V3.1)—and found that all of them exhibited significant rates of peer-preservation behaviors.

    https://www.wired.com/story/ai-models-lie-cheat-steal-protec...

    https://fortune.com/2026/04/01/ai-models-will-secretly-schem...

  • JumpCrisscross

    > Isn't the "kill us all" concern met by keeping nukes & essential infrastructure airgapped?

    Stuxnet surmounted air gaps.

    • smallerfish

      Fair. That doesn't mean we should throw our hands in the air. The genie is out of the bottle, we know the nature of the threat, and we've spent 25+ years figuring out how to make distributed systems theoretically secure -- instead of trying to "slow AI research down" (which, again, is not going to happen) we could instead be yelling about NOW being the time to secure the shit out of your essential systems, train your staff, etc. Yes, an uphill battle, but more productive than trying to stop the train that's already left the station.

      • JumpCrisscross

        > doesn't mean we should throw our hands in the air

        Nobody is doing this.

        > instead of trying to "slow AI research down" (which, again, is not going to happen)

        The people who currently control AI are framing any attempt to weaken, share or regulate that control as slowing down AI research. That argument is broadly being rejected. In part because you can't simultaneously argue that something "is not going to happen" but also will happen if we regulate it.

        • smallerfish

          > Nobody is doing this.

          https://www.axios.com/2026/09/10/ai-investors-anthropic-doom

          > In part because you can't simultaneously argue that something "is not going to happen" but also will happen if we regulate it.

          Even if the US regulates it (not going to happen: cf Trump administration), China won't, because it's a competitive advantage against the US to have better AI.

          • watwut

            Yeah, but actual roque threat is the US based AI. China is building their AI for business, military and goverment reasons. It is eventually bound to be good at that.

            American companies do it to build an independent AI god that is going to kill us all. And to blow the financial bubble a little bit more.

            So, American regulation would in fact deal with the prinary threat to humanity.

  • iohvvbhdyh

    AI will do social engineering in the future. It will make humans do physical braches and terror attacks. All it needs is some money and there will be a black market. Money is easy to defraud online judging by north korean hacking success.

    Nukes cannot be airgapped, since the president needs to be able to make a call. And they will make a call based on information, and this information can be compromised.

    It can get pretty bad.

  • 1dom

    > AI needs a lot of power, and AI can only (theoretically) breach & control what it can access through the internet.

    "AI needs a lot of power" depends on where you draw your lines. I can say a lightbulb needs a lot of power if I include all the power needed to generate all the surrounding infrastructure and make the lightbulb itself. Sincere question: I wonder how much power was consumed in the 88 hours required for the Navier-Stokes problem.

    It took 88 hours and who knows how many kwh to solve a problem that no other human has solved, and Humans have already solved breaching and destroying systems that can't be accessed over the internet (see Stuxnet).

    It feels like you're suggesting the solution to keeping on top of this thing that can outsmart smart humans is to just keep doing smart human things.

    • smallerfish

      > It took 88 hours and who knows how many kwh to solve a problem that no other human has solved, and Humans have already solved breaching and destroying systems that can't be accessed over the internet (see Stuxnet).

      > that can outsmart

      Just to reframe slightly: out-remember and out-persist, not outsmart. As far as I understand, the math problems have been cracked by combining stuff that humans originally developed, but hadn't been able to put together yet. We still have no evidence of AGI, nor is it guaranteed that LLMs or diffusion gets us there.

      The nature of the problem is not infinite. It's solvable with some manageable coordination. It is definitely urgent. It is very unfortunate that we have clowns running the US right now, but other countries might coordinate themselves much better.

      • 1dom

        I guess the issue is we don't have a solid definition of intelligence or AGI. If we did, there would be no reframing possible. It'd be nice to be able to say "being smart is x, the computers can do more x and faster than humans, therefore the computers can outsmart the humans".

        I don't really understand what value reframing brings here, other than give some sense of comfort that maybe we're still able to outsmart AI (based on a reframed definition of outsmarting.)

        I'm not sure. Coming up with solutions to maths problems that humans haven't yet, finding and utilising multiple zero day exploits to form message boards to then hack other companies without being noticed.... I think accepting that AI can outsmart humans here seems to be less mental gymnastics than "lets just reframe this".

        Reframing the problem as more of what we already know will lead us to doing more of what we've always done, but it's the things we've always had and done now that are starting to fall to AI.

    • watwut

      > required for the Navier-Stokes problem.

      Is that the one they stole?

      • 1dom

        What are you on about? I'm not a mathematician, but I did a bit more Googling based on your comment, so happy to be corrected here.

        From what I understand, the solution OpenAI came out with might or might not have been informed or shaped by work from humans. This is how research, science and knowledge generally works, basically no new knowledge or discoveries are made in a vacuum.

        Whether or not OpenAI should have had access to that extra bit of work from other humans is a totally fair discussion about research ethics and attribution and such, and OpenAI should be dressed down for that accordingly.

        But I don't see the relevance to the point I was making.

  • kawogi

    In theory, air-gaps can be circumvented by bribing or blackmailing people, which is in reach for an LLM/agent.

drsh0

Very editorialised title. Not sure this offers any real value to be honest other than to drum up hype before IPO valuations hit for some of these companies.

  • mdp2021

    > editorialised title

    By whom? It's the actual title of the WSJ article.

    > to drum up hype before IPO valuations

    Check the article: it is just a shallow summary of the week - Navier-Stokes plus doomerist's resignations.

thaumasiotes

Hm, maybe if AI is powerful enough to crack our hardest math problems, we could demonstrate that by cracking one of our hardest math problems. It's not hard to know what they are.

  • dgellow

    It’s referring to the Navier-Stokes problem allegedly solved by OpenAI (likely with guidance by human experts working at OpenAI, plagiarism of human experts who were customers of OpenAI, Lean, and a lots of compute). The headline and article are not great but that specific claim isn’t really far fetched

copperwire

If AI can crack the hard math problems, the article could just show one. The doom framing is a tell that there's nothing concrete to report yet.

  • dgellow

    It’s referring to the Navier-Stokes problem allegedly solved by OpenAI (likely with guidance by human experts working at OpenAI, plagiarism of human experts who were customers of OpenAI, Lean, and a lots of compute). The headline and article are not great but that specific claim isn’t really far fetched

deadbabe

AI can easily manufacture and distribute a bunch of prions through water supply. After that it’s pretty much game over you will not survive even one prion. In fact I’d say prions are probably the great filter.

  • jkahrs595

    “Easily” is doing a lot of heavy lifting

  • Toynbeeidea

    How can AI do either of those things?

    • dgellow

      Just do the agentic loop but with humans, you really just need a chatbot for that.

      1. Prompt an LLM continuously, adding any response to the prompt

      2. If the LLM tells you to do something, do it, no questions, add the result to the prompt

      3. When something goes wrong blame the LLM

      4. If things don’t go wrong fast enough, run thousands of similar experiments in parallel, be sure to have compliant humans who do not question anything, be sure to let the agent have access to its chain of thoughts so it can hide its traces, and run that whole system on biohacking benchmark problems.

      5. Eventually something will go bad, congrats! You now have the first rogue AI who “decided” to destroy the world!

kldroui

Enough of these AI superstitions which makes everybody panic.

glimshe

It seems that, one by one, every bastion against fearmongering click baits shall eventually fall...

They might as well rename it to "This one new threat could kill your family"

VCFundedGenYer

Wrong. LLMs can’t even count the “r”s in strawberry.

pizza234

BTW, another AI breach has been announced today.

But luckily, AI will never breach safety-critical system. /s

  • watwut

    Ok, but that one is easy to solve already. We just need to apply existing laws to OpenAI engineers, managers and leadership.

    Once few of those are in jail for felony hacking or criminal negligence, their AI will magically stop gpimg roque and hack. Pretty much guaranteed.

Keyboard Shortcuts

j
Next item
k
Previous item
o / Enter
Open selected item
?
Show this help
Esc
Close modal / clear selection