Live data from Hacker News

Terence Tao on O1

mathstodon.xyz

31–40 of 527 posts

Re: Terence Tao on O1

#31
post #19
post #9

I checked the links and I think it's amazing and it answers with Latex formatted notation. But I was curious and I asked something very simple, Euclid's first postulate and I got this answer: Euclid's Postulate 1: "Through any two points, there is exactly one straight line." In fact Euclid's Postulate 1 is "To draw a straight line from any point to any point." http://aleph0.clarku.edu/~djoyce/java/elements/bookI/book…

The original text is: Ἠιτήσθω ἀπὸ παντὸς σημείου ἐπὶ πᾶν σημεῖον εὐθεῖαν γραμμὴν ἀγαγεῖν. Roughly: let it be required that from any point to any point it is possible to draw a straight line. Both gpt4o and o1 roughly know the correct original text, so prompting, the model’s background memory, or random chance may influence your outcomes, though hopefully (in an improved model) you should never get you incorrect info.…

[deleted]

Re: Terence Tao on O1

#32
post #7
post #2

I wonder how long it took for each of the responses it gave

It varies a lot. If it's a simple question, it just does 3-4 sections of "thinking & reflection" but for more complicated ones I think I've seen something like 10 or more. Maybe 3-4 seconds per section on average I'd guess.

It's unclear if Terence is referring to "GPT-o1... a prototype version of the model that I was granted access to" as in "he was given access to GPT-o1 by the research team" or as in "he is using o1-preview". The differences in scale and quality between his shared output and the answer I get trying the same prompt from o1-preview suggest perhaps the former (otherwise luck). I haven't actually seen any examples of how long o1 "full" will think about this kind of question, though I expect it's somewhere in the same ballpark given the thought expansion still only has one real concept in it.

Re: Terence Tao on O1

#33
post #5

Once GPT is tuned more heavily on Lean (proof assistant) -- the way it is on Python -- I expect its usefulness for research level math to increase. I work in a field related to operations research (OR), and ChatGPT 4o has ingested enough of the OR literature that it's able to spit out very useful Mixed Integer Programming (MIP) formulations for many "problem shapes". For instance, I can give it a logic problem like "…

It also doesn't help that Lean has had so many breaking changes in such little time. When I tried using GPT-4 for it, it mostly rendered old code that would fail to run unless you already knew the answer and how to fix it, which basically made it entirely unhelpful.

Re: Terence Tao on O1

#34
post #5

Once GPT is tuned more heavily on Lean (proof assistant) -- the way it is on Python -- I expect its usefulness for research level math to increase. I work in a field related to operations research (OR), and ChatGPT 4o has ingested enough of the OR literature that it's able to spit out very useful Mixed Integer Programming (MIP) formulations for many "problem shapes". For instance, I can give it a logic problem like "…

> people who complain on HN that (paid/good - only Sonnet 3.5 and GPT4o are in this category)

Correction: I complain that the only decent model in "Open"AI's arsenal, that is GPT-4, has been replaced by a cheaper GPT-4o, which gives subpar answers to most of my question (I don't care it does it faster). As they moved it to "old, legacy" models, I expect they will phase it out, at which point I'll cancel my OpenAI subscriptions and Sonnet 3.5 will become the clear leader for my daily tasks.

Kudos to Anthropic for their great work, you guys are going in the right direction.

Re: Terence Tao on O1

#35
post #24
post #9

I checked the links and I think it's amazing and it answers with Latex formatted notation. But I was curious and I asked something very simple, Euclid's first postulate and I got this answer: Euclid's Postulate 1: "Through any two points, there is exactly one straight line." In fact Euclid's Postulate 1 is "To draw a straight line from any point to any point." http://aleph0.clarku.edu/~djoyce/java/elements/bookI/book…

Regardless of the wording being exact or not, ChatGPT’s answer is incorrect in its contents. The statement “exactly one” requires the parallel postulate, since otherwise it’s not necessarily true. Specifically in spherical geometry, which is considered to be consistent with Euclid’s first four postulates (i.e. without the parallel postulate). The bottom line is, you can’t take any single LLM statement at face value,…

it's the point-line postulate, you can use it as part of a set of axioms equivalent to Euclid but it definitely not one of Euclid's

https://en.wikipedia.org/wiki/Point%E2%80%93line%E2%80%93pla...

Re: Terence Tao on O1

#36
Is there a list of discoveries or siginficant works/constructions made by people collaborating with LLM's? I mean as opposed to specific deep networks like Alphafold or Graphcast?

Re: Terence Tao on O1

#37
post #19
post #9

I checked the links and I think it's amazing and it answers with Latex formatted notation. But I was curious and I asked something very simple, Euclid's first postulate and I got this answer: Euclid's Postulate 1: "Through any two points, there is exactly one straight line." In fact Euclid's Postulate 1 is "To draw a straight line from any point to any point." http://aleph0.clarku.edu/~djoyce/java/elements/bookI/book…

The original text is: Ἠιτήσθω ἀπὸ παντὸς σημείου ἐπὶ πᾶν σημεῖον εὐθεῖαν γραμμὴν ἀγαγεῖν. Roughly: let it be required that from any point to any point it is possible to draw a straight line. Both gpt4o and o1 roughly know the correct original text, so prompting, the model’s background memory, or random chance may influence your outcomes, though hopefully (in an improved model) you should never get you incorrect info.…

it's definitely wrong though, "exactly one" straight line between two points is a different postulate and a stronger one.

Euclid has been translated, restated, and re-presented in enough books and textbooks that I'd expect a big-enough LLM to have actually memorized this correctly tbh

Re: Terence Tao on O1

#38
post #5

Once GPT is tuned more heavily on Lean (proof assistant) -- the way it is on Python -- I expect its usefulness for research level math to increase. I work in a field related to operations research (OR), and ChatGPT 4o has ingested enough of the OR literature that it's able to spit out very useful Mixed Integer Programming (MIP) formulations for many "problem shapes". For instance, I can give it a logic problem like "…

I entirely agree about their utility. HN, and the internet in general, have become just an ocean of reactionary sandbagging and blather about how "useless" LLMs are. Meanwhile, in the real world, I've found that I haven't written a line of code in weeks. Just paragraphs of text that specify what I want and then guidance through and around pitfalls in a simple iterative loop of useful working code. It's entirely a lea…

> I've found that I haven't written a line of code in weeks

Which is great until your next job interview. Really, it's tempting in the short run but I made a conscious decision to do certain tasks manually only so that I don't lose my basic skills.

Re: Terence Tao on O1

#39
My experience with O1 has been very different. I wouldn't even say it's performing at a "good undergrad" level for me.

For example, I asked a pretty simple question here and it got completely confused:

https://moorier.com/math-chat-1.png https://moorier.com/math-chat-2.png https://moorier.com/math-chat-3.png

(Full chat should be here: https://chatgpt.com/share/66e5d2dd-0b08-8011-89c8-f6895f3217...)

Re: Terence Tao on O1

#40

My experience with O1 has been very different. I wouldn't even say it's performing at a "good undergrad" level for me. For example, I asked a pretty simple question here and it got completely confused: https://moorier.com/math-chat-1.png https://moorier.com/math-chat-2.png https://moorier.com/math-chat-3.png (Full chat should be here: https://chatgpt.com/share/66e5d2dd-0b08-8011-89c8-f6895f3217... )

Anecdata, but I've been finding O1 to be worse than 4o & Claude 3.5 Sonnet. To add insult to injury, it's slower & chattier.
Post reply on HN