Live data from Hacker News

More questions about whether researchers can trust OpenAI with unpublished math

mathstodon.xyz

121–130 of 845 posts

Re: More questions about whether researchers can trust OpenAI with unpublished math

#121

Earlier quoted context omitted.

AI is America's last chance to salvage its empire. Nothing will be allowed to impede it.

But I don't see how. AI is going to be a commodity in short order and best case the US will be a temporary leader in the supply of tokens. Meanwhile AI is going to destroy much of the Service and Software industry that make up most of the US economy. And the US is betting every last cent to bring about this future. It does make sense for Trump since this might be a sugar high that lasts till the end of his term.

In USA there is surprisingly little state involvement in the whole llm mania. Who needs the state with 800 lbs gorillas like Google, Amazon, Nvidia, etc

In China, it is the principal obsession of the entire communist party which eg funds the whole infrastructure without a single NIMBY peep.

The strange emphasis in China on humanoid robotic constructions is due to the CCP realization that with the cataclysmic fertility collapse they will increasingly have no one to rule.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#123
post #98
post #62

I wonder what’s more valuable in our prompts: the raw data or the feedback system that drives the exchange towards a goal. For a long time it was clearly the former, but now I think it is the latter. The models have enough knowledge (orders of magnitude more than a human could ever learn) but are now getting better at what to do with it thanks to learning from the decisions that we make in conversations with AI agent…

I think so too. The value is in the entire conversation. IMO, "domain experts" don't run LLMs blindly and hands free. This does not work for top level work (e.g., mathematical proofs, coding anything more complex than yet another slop game or website). Experts have long sessions where they prompt and guide LLM in response to what it produces. This is the discovery process. And frontier labs definitely train on that.…

Last year we were saying there must be a human-in-the-loop (HitL), but anyone who is the HitL exhibits the “HitL skill” to the agent.

There might be no books about human intuition but we teach it to LLMs by interacting with them

Re: More questions about whether researchers can trust OpenAI with unpublished math

#124
post #52

Doing some research and at this point doing it very much in the open with dates on GitHub so if any AI Lab says they re-discover my exact work it will be obvious that the AI used or was trained on my work. I am guessing anyone in a similar situation is now thinking about how they date their existing work if the math is done, but the proses are not.

If only prompts could also be watermarked.

The session data could be cryptographically signed. Probably easier in an open harness?

Re: More questions about whether researchers can trust OpenAI with unpublished math

#125
post #26

Earlier quoted context omitted.

AI is also trained on your HN posts. And lots of other things you post on the internet.

Public posts on the internet are acceptable (to me). For my (private) prompts, I need a warning telling me they may be used for training.

> Public posts on the internet are acceptable (to me).

Everybody needs to rethink this again.

Before LLMs the barrier to entry for building a character profile based on your various public posts was quite high. Remember "Psychographics" (https://en.wikipedia.org/wiki/Psychographics) and the infamous "Cambridge Analytica"?

Earlier it involved data mining, data cleaning, structuring data, building models, running algorithms and then evaluating the results for semantic information. Now it is straight to unfiltered semantic inference using a single sentence prompt (eg. point it to your HN profile and see what you get).

I actually did this on my HN profile and found it troubling. There were many unwarranted/hallucinated inferences due to the fact that it requires "commonsense reasoning" (https://en.wikipedia.org/wiki/Commonsense_reasoning), understanding human motivations and behaviour, context, assumptions, societal knowledge etc. which LLMs are bad at.

PS: You can cut-and-paste the above paras into a LLM prompt and ask it to elaborate for further details. The system itself will explain to you the problems/deficiencies which are quite scary.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#127

In this domain, an apparent single unique piece of work is often composed of several breakthroughs. For example, when Andrew Wiles proved Fermat's Last Theorem, he had to develop multiple new pieces of mathematical technology to get there. The claim here seems to be that the human mathematicians, working with AI, developed technology to go A->B->C. By training on those conversations, OpenAI was then able to encourage…

Also, wasn't their B+C research private at the time, with them only releasing those details publicly after this blew up?

If they had published B+C, I think that would lean more towards fair game, as that is how research works and is improved on over time. But it seems like unpublished/private B + C may have been used by the model to hint it into working out how to get from A->D.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#128
post #34

I'm genuinely surprised that more people - including this mathematician in particular - don't untick the "improve the model for everyone" box. Unless the suggestion is that OpenAI ignore this preference?

Even if you've ticked that box, the conversation can still be trained on if you:

(a) Click thumbs-up/down in the conversation [1]

(b) Have the conversation flagged for potential safety concerns

[1]: https://help.openai.com/en/articles/5722486-how-your-data-is....

Re: More questions about whether researchers can trust OpenAI with unpublished math

#129
post #9

Some mathematicians I know who've been following this have realized that they'd all gotten some emails from people they now know to be affiliated with OpenAI/Anthropic asking questions about their research in a way that seemed like scooping attempts. Also, a lot of my mathematicians buddies have reported students basically asking if it's worth ever doing grad school for pure math, and even very motivated students are…

One has nothing to do with the other. It was long predicted that math and software developments would be the first domain where AI was going to do major damage. If OpenAI and Anthropic didn't get into math result dick measuring, Internet anons would have in their place, 6 months later when it got cheaper.

I think you’re missing an important distinction. “Major damage” to the talent pipeline because models become capable of original end-to-end mathematics is what the community has been discussing. But if the models rely on sniping nearly complete work then this damage is antisocial without a lot of upside, it would be destroying a talent pipeline that would still necessary for continued progress.

Which is it? I don’t think OpenAI is being transparent enough for us to really understand whether these results would have been possible without relying on unpublished information from the solution strategies of the experts

Re: More questions about whether researchers can trust OpenAI with unpublished math

#130
I've been wondering whether AI really is improving rapidly at open problems or we're being fooled.

- OpenAI invites researchers to use their models, in fact giving at least 100,000 researchers free access[1], but there are also those that pay

- Internal OpenAI models are reportedly solving open problems at a surprisingly fast rate[2]

- But researchers will typically work on open problems. A researcher who is using Codex to make progress on open problems will be feeding it fresh training data on precisely the problems the internal models are evaluated on.

- So while it looks like the new models are suddenly solving lots of open problems, they could be significantly piggybacking on human progress, with models "inspired" by the work of researchers from all around the world?

This theory predicts that there'll be many more researchers coming forward just like TFA, as sOpenAI announces more solutions. It doesn't assume all of AI progress is a mirage, just that there's plagiarism.

[1]: https://openai.com/index/chatgpt-for-academic-researchers/

[2]: https://xcancel.com/OpenAI/status/2097374643518640382#m

Post reply on HN