Settings

Theme

OpenAI have no mathematicians capable of understanding what they put out

mastodon.social

85 points by doener · 42 comments

Reader

5 threads
oefrha

Loosely related, you may want to check your own privacy settings at

https://chatgpt.com/codex/cloud/settings/data#settings/DataC...

https://claude.ai/new#settings/data-privacy-controls

I just realized I've been happily "improving the model for everyone"...

  • tristanj

    That OpenAI setting helps, but there is a better way to do it. To completely opt-out of training, submit a request via the OpenAI privacy portal.

    Visit this website https://privacy.openai.com/policies/en/ , click "Make a Privacy Request", choose "Do not train on my content", and complete the form. That submits a formal objection to training on your data, as required by GDPR/your local legislation.

    • rich_sasha

      Haha. I am 100% these companies will ignore this if they choose to. Just as they played fast and loose with copyright rules.

      They would do it, the say “ah sorry chaps, impossible to extract it from the dataset by now, anyway we anonymized it so can’t tell what’s what, and we can’t risk losing to China. Oh look - did you see Superman fly outside?”.

      • nicce

        European users have right to be forgotten. Waiting for the court order to delete all models.

      • tristanj

        This form is legally binding and has more legal weight than just clicking a toggle. If they still train on my data, they can get sued, and I'll get a payout.

        • m4rtink

          As legaly binding as all the data the license and ToS of which they ignored when scrapping - before trying to sell it back with an eternal subscription to their plagiarism machine ?

        • 63stack

          Are there any precedents that companies got sued on this? Not that I doubt you, but it would be nice to see if it actually has teeth or not.

        • vuurmot

          Fighting legally against one of the richest companies in existence, with the entire American apparatus behind it, is a brave move

          • bot403

            The world is pretty pissed at the u.s. and seeking sovereign models and ai. Richest or not if the companies are wise they will try a bit not to piss off the entire rest of the world. At this point many countries are willing to cut off the u.s. even if they take a short or medium term hit.

        • moljac024

          How would you prove they trained on your data specifically?

          • Ancapistani

            Showing that a novel mathematical approach was in your prompts shortly before the model “proposes” that approach to relative amateurs would be ideal, if only we had a case like that…

            • itemize123

              and still they denied it - it only shows that we have little recourse.

              • Ancapistani

                It’s possible that they didn’t train on it, and the approach was derived by the LLM.

                What’s a problem is that they haven’t outright denied it. That could be caution and them doing their due diligence first, it could be that it was intentional and they didn’t expect to get caught, or it could be because they have no way of knowing themselves.

    • simianwords

      > but there is a better way to do it.

      This looks incorrect. OpenAI have flat out said that there's no need of doing this and both ways are equivalent

      > We respect our users' choice whether to use their data to “improve our models for everyone” regardless of where they express that choice. Users can opt out in the in-app settings or indeed also in our privacy portal. They do not need to opt out in both places, and we will make this clearer in our Help Center.

      https://x.com/thsottiaux/status/2097746417012166816

      • my-huge-pony

        Even if that's true right now, if one of the options is legally binding and the other is "we promise not to eat your data... for now", the first option is still better.

        • simianwords

          so what option do you leave them with? they are giving both alternatives and clearly stating both do the same thing underneath. there's literally nothing else they could have done here.

          • tristanj

            Option A can be accidentally undone with a few clicks.

            Option B is legally binding and permanent.

            • simianwords

              I get it, but what option does OpenAI have here other than to present the two options?

      • tristanj

        They're not equivalent.

        The one in ChatGPT settings only applies to ChatGPT. The one on the OpenAI privacy portal applies to all OpenAI products, present and future.

tetrisgm

The problem is that eventually, there will be no humans who can follow the results AI will give. That’s the endgame for this tech: to produce knowledge at speeds and quality beyond what we can

  • rsfern

    That’s the prevailing narrative, but I think this controversy calls it into question to some extent. If the OpenAI result wouldn’t have been possible without experts seeding the training data with feedback on promising solution routes, there’s less reason to believe this, IMO. More information and transparency is needed

    • bot403

      That's the pickle isn't it? AI solved it with human help. But it's standing on their shoulders. Neither AI nor the humans got it alone.

      Humanity is better having solved this issue. But which humans were credited and benefited is the issue.

      • rsfern

        I agree (and so does Buckmaster based on his written statement) that we are better having solved this.

        But I disagree that which humans were credited is the heart of the issue in this particular controversy. The question is what do you need to bring to the table for a result like this. A pre-release frontier model trained on the open literature and $15 million of inference? Or all that plus a year of the experts finding the path to the solution for the model to run with?

        I think it makes a huge difference in terms of what we think the future of mathematical research will be like, and whether we should still encourage students to go into this field, which was the original topic of this thread

  • uargos

    Which asks the question of what knowledge is. Can knowledge be super human ? Or is knowledge a human matter ? If so (like i believe), then what those ai labs are doing is far from the end of the story. Because the goal of science is not to produce a certificate of something, but more to produce an explanation that can fit in a human brain, that can be reasoned on, and that can be retargeted. In this sense, producing a million lines proof is not really producing knowledge, even less so doing science.

    • curt15

      Suppose there is an oracle that can tell you whether a particular result in maths or physics is true. Would research still have any value?

    • michaeljx

      Knowing something is proven true is knowledge, even if the mechanics of the proof are not understood.

      • trolleski

        In maths, you go through proofs every step of the way. Same in physics, you start by performing even the most basic experiments, and build your way up.

      • rented_mule

        If the mechanics of the proof are not understood, how do we know it's been proven? Math proofs are not a "trust me, bro" kind of thing.

  • menaerus

    In certain, more complex, software engineering domains this almost became true as of today.

  • smitty1e

    Yeah, but then you can use another round of AI to break these things down to crayon level, no?

    Or is it just all PFM? (Pure, Fanciful Magic)

tristanj

This post is out of date. OpenAI quietly updated the references on their paper earlier today and added several authors.

  • kzrdude

    The main claim, that "they do not seem to have any mathematicians capable of understanding what they put out", was also corroborated by Sebastien (OpenAI) who explained they don't have any experts on Navier-Stokes.

japgolly

> our hypodissipative result (which is not public), but as I understand it part of their training data

How did it become part of their training data if it wasn't public? /confused

  • vessenes

    This is speculation right now. The idea would be that if someone used the product and granted training rights, which is the default for many subscription levels, then some knowledge would have been imparted into the general weights of the new model.

    oAI has made clear they did not specifically pull in any user data to context for this run.

    • dpiers

      Tristan Buckmaster’s post cited extensive use of LLMs in the process of his collaboration with Levent:

      “We used several LLMs throughout: Anthropic’s Claude, OpenAI’s Codex, especially with GPT-5.6 Sol and, more recently, Astra.

      The latter was only used for writeups and auditing our arguments. For most of the past year progress was slow. We worked through the literature and upgraded various preliminary results, up to obtaining finite time blow up for the Incompressible Porous Media equation (with smooth forcing).

      This was until about a month ago, when we had real progress: on August 15th, we obtained the blow up results, with smooth forcing, for both Boussinesq and Euler. I can say the first LLM generated proof Levent sent me was the most horrendous I have ever read; we verified it on Lean on August 22nd. Since this point, we have been working around the clock to understand this proof and turn it into something readable.”

      The OpenAI research post states they began training GPT-6 internally on August 28th, and that user chats are used to train models.

      “We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem. While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models .”

      For an incredibly niche topic like this, I believe it’s extremely likely that Buckmaster/Levant’s work would influence the direction of OpenAI’s agents’ work even as a de-identified drop in the overall bucket of training data.

    • rramadass

      Tristan Buckmaster, the mathematician at the center of it (https://cims.nyu.edu/~tristanb/) put out a public statement that everybody should read (pdf) - https://cims.nyu.edu/~tristanb/statement.pdf

      So what might have been the incentive for OpenAI to do all this shenanigans? It might have to do with getting its models certified for AGI and getting out of lockin with Microsoft - https://deadneurons.substack.com/p/the-quiet-unwinding-of-mi...

  • WoodenChair

    One of the researchers was using the product. They train on your private interactions unless you explicitly opt out in the settings.

    • rrobukef

      It wouldn't surprise me that even if you opt-out they still train on 'de-identified' chat data. After all, facts cannot be copyrighted and one cannot be sued for unethical behaviour.

simianwords

Can anyone explain to me why Tristan simply didn't go to settings page and turn off the thing? Especially when Levent was collaborating with him and Levent is completely aware of how the training data is used and the implications thereby?

If it is oversight, then that's ok and OpenAI can volunteer to make him the lead author which they did. But he's pissed that OpenAI is not letting Levent as well, who had access to internal Anthropic models. So this guy thinks

1. oh my bad i forgot to turn off the consent thing in settings page

2. also i'll collaborate with a literal Anthropic employee who has access to their internal models

3. i'll also reject OpenAI's deal to be the lead author because i want an employee of the competitor to be a part of it

I don't get the mindset.

Keyboard Shortcuts

j
Next item
k
Previous item
o / Enter
Open selected item
?
Show this help
Esc
Close modal / clear selection