Live data from Hacker News

Terence Tao on O1

mathstodon.xyz

41–50 of 527 posts

Re: Terence Tao on O1

#41
post #5

Once GPT is tuned more heavily on Lean (proof assistant) -- the way it is on Python -- I expect its usefulness for research level math to increase. I work in a field related to operations research (OR), and ChatGPT 4o has ingested enough of the OR literature that it's able to spit out very useful Mixed Integer Programming (MIP) formulations for many "problem shapes". For instance, I can give it a logic problem like "…

Why do you expect GPT being tuned on Lean will help it for research-level math?

Re: Terence Tao on O1

#42
post #5

Once GPT is tuned more heavily on Lean (proof assistant) -- the way it is on Python -- I expect its usefulness for research level math to increase. I work in a field related to operations research (OR), and ChatGPT 4o has ingested enough of the OR literature that it's able to spit out very useful Mixed Integer Programming (MIP) formulations for many "problem shapes". For instance, I can give it a logic problem like "…

I entirely agree about their utility. HN, and the internet in general, have become just an ocean of reactionary sandbagging and blather about how "useless" LLMs are. Meanwhile, in the real world, I've found that I haven't written a line of code in weeks. Just paragraphs of text that specify what I want and then guidance through and around pitfalls in a simple iterative loop of useful working code. It's entirely a lea…

LLMs are certainly not useless.

But "lines of code written" is a hollow metric to prove utility. Code literacy is more effective than code illiteracy.

Lines of natural language vs discrete code is a kind of preference. Code is exact which makes it harder to recall and master. But it provides information density.

> by just knuckling down and learning how to do the work?

This is the key for me. What work? If it's the years of learning and practice toward proficiency to "know it when you see it" then I agree.

Re: Terence Tao on O1

#43
post #5

Once GPT is tuned more heavily on Lean (proof assistant) -- the way it is on Python -- I expect its usefulness for research level math to increase. I work in a field related to operations research (OR), and ChatGPT 4o has ingested enough of the OR literature that it's able to spit out very useful Mixed Integer Programming (MIP) formulations for many "problem shapes". For instance, I can give it a logic problem like "…

I entirely agree about their utility. HN, and the internet in general, have become just an ocean of reactionary sandbagging and blather about how "useless" LLMs are. Meanwhile, in the real world, I've found that I haven't written a line of code in weeks. Just paragraphs of text that specify what I want and then guidance through and around pitfalls in a simple iterative loop of useful working code. It's entirely a lea…

> Meanwhile, in the real world, I've found that I haven't written a line of code in weeks. Just paragraphs of text that specify what I want and then guidance through and around pitfalls in a simple iterative loop of useful working code.

could it be that you are mostly engaged in "boilerplate coding", where LLMs are indeed good?

Re: Terence Tao on O1

#44

>could not generate conceptual ideas of their own Is the most important part imo. A big goal should be some ai system coming up with its own discovery and ideas. Really unclear how we can get from the current paradigm to it coming up with something like general relativity, like Einstein. Does it require embodiment?

we don't know how to reliably produce humans who produce GR-level ideas, this might be biting off a lot more than we can chew

Re: Terence Tao on O1

#45
post #5

Once GPT is tuned more heavily on Lean (proof assistant) -- the way it is on Python -- I expect its usefulness for research level math to increase. I work in a field related to operations research (OR), and ChatGPT 4o has ingested enough of the OR literature that it's able to spit out very useful Mixed Integer Programming (MIP) formulations for many "problem shapes". For instance, I can give it a logic problem like "…

I entirely agree about their utility. HN, and the internet in general, have become just an ocean of reactionary sandbagging and blather about how "useless" LLMs are. Meanwhile, in the real world, I've found that I haven't written a line of code in weeks. Just paragraphs of text that specify what I want and then guidance through and around pitfalls in a simple iterative loop of useful working code. It's entirely a lea…

> HN, and the internet in general, have become just an ocean of reactionary sandbagging and blather about how "useless" LLMs are.

This is cult like behaviour that reminds me so much of the crypto space.

I don't understand why people are not allowed to be critical of a technology or not find it useful.

And if they are they are somehow ignorant, over-reacting or deficient in some way.

Re: Terence Tao on O1

#46
post #19

Earlier quoted context omitted.

The original text is: Ἠιτήσθω ἀπὸ παντὸς σημείου ἐπὶ πᾶν σημεῖον εὐθεῖαν γραμμὴν ἀγαγεῖν. Roughly: let it be required that from any point to any point it is possible to draw a straight line. Both gpt4o and o1 roughly know the correct original text, so prompting, the model’s background memory, or random chance may influence your outcomes, though hopefully (in an improved model) you should never get you incorrect info.…

it's definitely wrong though, "exactly one" straight line between two points is a different postulate and a stronger one. Euclid has been translated, restated, and re-presented in enough books and textbooks that I'd expect a big-enough LLM to have actually memorized this correctly tbh

Agreed. That is what the original poster said. I didnt manage to reproduce the error on my end, but I dont know the full context or maybe the memory on my end changes the output.

Re: Terence Tao on O1

#47
post #9

I checked the links and I think it's amazing and it answers with Latex formatted notation. But I was curious and I asked something very simple, Euclid's first postulate and I got this answer: Euclid's Postulate 1: "Through any two points, there is exactly one straight line." In fact Euclid's Postulate 1 is "To draw a straight line from any point to any point." http://aleph0.clarku.edu/~djoyce/java/elements/bookI/book…

Euclid's Elements is less pervasive on the Internet then content produced for Liberal Arts math courses. As those courses tend to emphasize critical thinking and problem-solving math over pure theory and advanced concepts, they tend to be far more common and tend to win compared to more domain specific meanings.

Examples:

https://en.wiktionary.org/wiki/Euclidean_geometry

https://www.cerritos.edu/dford/SitePages/Math_70_F13/Postula...

Problems with polysemy across divergent, more advanced theories has been one of my biggest challenges in probing some of my areas of intrest.

Funny enough, one of my pet areas of obscure interest, riddled basins, is constantly muddied not by math, but LSAT questions, specifically non-math content directed at a reading comprehension test: "September 2006 LSAT Section 1 Question 26"

IMHO a lot of the prompt engineering you have to do with these highly domain specific problems is avoiding the most common responses in the corpus.

LLM responses will tend to reflect common usage, not academic terminology unless someone cares enough to change that for a specific case.

Re: Terence Tao on O1

#48
post #40

My experience with O1 has been very different. I wouldn't even say it's performing at a "good undergrad" level for me. For example, I asked a pretty simple question here and it got completely confused: https://moorier.com/math-chat-1.png https://moorier.com/math-chat-2.png https://moorier.com/math-chat-3.png (Full chat should be here: https://chatgpt.com/share/66e5d2dd-0b08-8011-89c8-f6895f3217... )

Anecdata, but I've been finding O1 to be worse than 4o & Claude 3.5 Sonnet. To add insult to injury, it's slower & chattier.

And sometimes it just bugs out and doesn't give any response? Faced that twice now, it "thought" for like 10-30s then no answer and I had to click regenerate and wait for it again.

Re: Terence Tao on O1

#49

Earlier quoted context omitted.

I entirely agree about their utility. HN, and the internet in general, have become just an ocean of reactionary sandbagging and blather about how "useless" LLMs are. Meanwhile, in the real world, I've found that I haven't written a line of code in weeks. Just paragraphs of text that specify what I want and then guidance through and around pitfalls in a simple iterative loop of useful working code. It's entirely a lea…

> I've found that I haven't written a line of code in weeks Which is great until your next job interview. Really, it's tempting in the short run but I made a conscious decision to do certain tasks manually only so that I don't lose my basic skills.

Job interview? You might be surprised at the number of us who don’t code for a job.

Re: Terence Tao on O1

#50
post #3

"with even the latest tools the effort put in to get the model to produce useful output is still some multiple (but not an enormous multiple now, say 2x to 5x) of the effort needed to properly prompt and verify the output. However, I see no reason to prevent this ratio from falling below 1x in a few years, which I think could be a tipping point for broader adoption of these tools in my field" Given the log scale on c…

The y axis is also log scale (log likelihood). It’s a power law, not an exponential law.

I was referring to the o1 AIME accuracy figure (x log scale compute, y is % (not log)) and similar https://openai.com/index/learning-to-reason-with-llms/
Post reply on HN