Live data from Hacker News

Terence Tao on O1

mathstodon.xyz

21–30 of 527 posts

Re: Terence Tao on O1

#21
post #5

Once GPT is tuned more heavily on Lean (proof assistant) -- the way it is on Python -- I expect its usefulness for research level math to increase. I work in a field related to operations research (OR), and ChatGPT 4o has ingested enough of the OR literature that it's able to spit out very useful Mixed Integer Programming (MIP) formulations for many "problem shapes". For instance, I can give it a logic problem like "…

> side Or (4) LLMs simply do not work properly for many use cases in particular where large volumes of trained data doesn't exist in its corpus. And in these scenarios rather than say "I don't know" it will over and over again gaslight you with incoherent answers. But sure condescendingly blame on the user for their ignorance and inability to understand or use the tool properly. Or call their criticism low-effort.

What's the difference between (3) and (4), shouldn't the former contain the latter?

Re: Terence Tao on O1

#22
post #3

"with even the latest tools the effort put in to get the model to produce useful output is still some multiple (but not an enormous multiple now, say 2x to 5x) of the effort needed to properly prompt and verify the output. However, I see no reason to prevent this ratio from falling below 1x in a few years, which I think could be a tipping point for broader adoption of these tools in my field" Given the log scale on c…

The y axis is also log scale (log likelihood). It’s a power law, not an exponential law.

Re: Terence Tao on O1

#23
post #9

I checked the links and I think it's amazing and it answers with Latex formatted notation. But I was curious and I asked something very simple, Euclid's first postulate and I got this answer: Euclid's Postulate 1: "Through any two points, there is exactly one straight line." In fact Euclid's Postulate 1 is "To draw a straight line from any point to any point." http://aleph0.clarku.edu/~djoyce/java/elements/bookI/book…

“Exact wording” would be Ancient Greek. Euclid did not even write in English. You’re checking whether the model matches a specific translation, which is not valuable. If you search around you’ll find many sources that choose a more intelligible phrasing.

Re: Terence Tao on O1

#24
post #9

I checked the links and I think it's amazing and it answers with Latex formatted notation. But I was curious and I asked something very simple, Euclid's first postulate and I got this answer: Euclid's Postulate 1: "Through any two points, there is exactly one straight line." In fact Euclid's Postulate 1 is "To draw a straight line from any point to any point." http://aleph0.clarku.edu/~djoyce/java/elements/bookI/book…

Regardless of the wording being exact or not, ChatGPT’s answer is incorrect in its contents. The statement “exactly one” requires the parallel postulate, since otherwise it’s not necessarily true. Specifically in spherical geometry, which is considered to be consistent with Euclid’s first four postulates (i.e. without the parallel postulate).

The bottom line is, you can’t take any single LLM statement at face value, even in seemingly easy to answer cases like this.

Re: Terence Tao on O1

#25

>could not generate conceptual ideas of their own Is the most important part imo. A big goal should be some ai system coming up with its own discovery and ideas. Really unclear how we can get from the current paradigm to it coming up with something like general relativity, like Einstein. Does it require embodiment?

Why should that be a big goal? It's difficult, it's not what they are good at, and they can get a lot better at assisting in other ways through incremental improvements. I'm happy to leave this part to the humans, at least for now, especially when there's so much more improvement still possible in other directions.

It also seems like one of those things where we ought to ask whether we should, before asking whether we could. Why not focus on areas that are easier, more beneficial, and less problematic from a "should" perspective?

Re: Terence Tao on O1

#26
post #5

Once GPT is tuned more heavily on Lean (proof assistant) -- the way it is on Python -- I expect its usefulness for research level math to increase. I work in a field related to operations research (OR), and ChatGPT 4o has ingested enough of the OR literature that it's able to spit out very useful Mixed Integer Programming (MIP) formulations for many "problem shapes". For instance, I can give it a logic problem like "…

I entirely agree about their utility.

HN, and the internet in general, have become just an ocean of reactionary sandbagging and blather about how "useless" LLMs are.

Meanwhile, in the real world, I've found that I haven't written a line of code in weeks. Just paragraphs of text that specify what I want and then guidance through and around pitfalls in a simple iterative loop of useful working code.

It's entirely a learned skill, the models (and very importantly the tooling around them) have arrived at the base line they needed.

Much Much more productive world by just knuckling down and learning how to do the work.

edit: https://aider.chat/ + paid 3.5 sonnet

Re: Terence Tao on O1

#27

I'm so excited in anticipation of my near-term return to studying math, as an independent curiosity hobby. It's going to be epically fun this time around with LLM's to lean on. Coincidentally like Terence Tao, I've also been asking complex analysis queries* of LLM's, things I was trying to understand better in my working through textbooks. Their ability to interpret open-ended math questions, and quickly find distant…

How will you know if its answers are correct or not?

Re: Terence Tao on O1

#28

I'm so excited in anticipation of my near-term return to studying math, as an independent curiosity hobby. It's going to be epically fun this time around with LLM's to lean on. Coincidentally like Terence Tao, I've also been asking complex analysis queries* of LLM's, things I was trying to understand better in my working through textbooks. Their ability to interpret open-ended math questions, and quickly find distant…

How will we even measure this? Benchmarks are gamed/trained on and there is no way that there is much signal in the chatbot arena for these types of queries? I think in just a few month the average user will not be able to tell the difference in performance between the major models

[deleted]

Re: Terence Tao on O1

#29
post #5

Once GPT is tuned more heavily on Lean (proof assistant) -- the way it is on Python -- I expect its usefulness for research level math to increase. I work in a field related to operations research (OR), and ChatGPT 4o has ingested enough of the OR literature that it's able to spit out very useful Mixed Integer Programming (MIP) formulations for many "problem shapes". For instance, I can give it a logic problem like "…

Give it a few months. ChatGPT will be recommending GPTs to use or do it automatically.

Nothing is static in the way things are moving.

Re: Terence Tao on O1

#30

I'm so excited in anticipation of my near-term return to studying math, as an independent curiosity hobby. It's going to be epically fun this time around with LLM's to lean on. Coincidentally like Terence Tao, I've also been asking complex analysis queries* of LLM's, things I was trying to understand better in my working through textbooks. Their ability to interpret open-ended math questions, and quickly find distant…

How will you know if its answers are correct or not?

Because I'm verifying everything by hand, as is the whole point of studying pure mathematics.
Post reply on HN