Once GPT is tuned more heavily on Lean (proof assistant) -- the way it is on Python -- I expect its usefulness for research level math to increase. I work in a field related to operations research (OR), and ChatGPT 4o has ingested enough of the OR literature that it's able to spit out very useful Mixed Integer Programming (MIP) formulations for many "problem shapes". For instance, I can give it a logic problem like "…
> side Or (4) LLMs simply do not work properly for many use cases in particular where large volumes of trained data doesn't exist in its corpus. And in these scenarios rather than say "I don't know" it will over and over again gaslight you with incoherent answers. But sure condescendingly blame on the user for their ignorance and inability to understand or use the tool properly. Or call their criticism low-effort.
Terence Tao on O1
21–30 of 527 posts
Re: Terence Tao on O1
#22"with even the latest tools the effort put in to get the model to produce useful output is still some multiple (but not an enormous multiple now, say 2x to 5x) of the effort needed to properly prompt and verify the output. However, I see no reason to prevent this ratio from falling below 1x in a few years, which I think could be a tipping point for broader adoption of these tools in my field" Given the log scale on c…
Re: Terence Tao on O1
#23I checked the links and I think it's amazing and it answers with Latex formatted notation. But I was curious and I asked something very simple, Euclid's first postulate and I got this answer: Euclid's Postulate 1: "Through any two points, there is exactly one straight line." In fact Euclid's Postulate 1 is "To draw a straight line from any point to any point." http://aleph0.clarku.edu/~djoyce/java/elements/bookI/book…
Re: Terence Tao on O1
#24I checked the links and I think it's amazing and it answers with Latex formatted notation. But I was curious and I asked something very simple, Euclid's first postulate and I got this answer: Euclid's Postulate 1: "Through any two points, there is exactly one straight line." In fact Euclid's Postulate 1 is "To draw a straight line from any point to any point." http://aleph0.clarku.edu/~djoyce/java/elements/bookI/book…
The bottom line is, you can’t take any single LLM statement at face value, even in seemingly easy to answer cases like this.
Re: Terence Tao on O1
#25>could not generate conceptual ideas of their own Is the most important part imo. A big goal should be some ai system coming up with its own discovery and ideas. Really unclear how we can get from the current paradigm to it coming up with something like general relativity, like Einstein. Does it require embodiment?
It also seems like one of those things where we ought to ask whether we should, before asking whether we could. Why not focus on areas that are easier, more beneficial, and less problematic from a "should" perspective?
Re: Terence Tao on O1
#26Once GPT is tuned more heavily on Lean (proof assistant) -- the way it is on Python -- I expect its usefulness for research level math to increase. I work in a field related to operations research (OR), and ChatGPT 4o has ingested enough of the OR literature that it's able to spit out very useful Mixed Integer Programming (MIP) formulations for many "problem shapes". For instance, I can give it a logic problem like "…
HN, and the internet in general, have become just an ocean of reactionary sandbagging and blather about how "useless" LLMs are.
Meanwhile, in the real world, I've found that I haven't written a line of code in weeks. Just paragraphs of text that specify what I want and then guidance through and around pitfalls in a simple iterative loop of useful working code.
It's entirely a learned skill, the models (and very importantly the tooling around them) have arrived at the base line they needed.
Much Much more productive world by just knuckling down and learning how to do the work.
edit: https://aider.chat/ + paid 3.5 sonnet
Re: Terence Tao on O1
#27I'm so excited in anticipation of my near-term return to studying math, as an independent curiosity hobby. It's going to be epically fun this time around with LLM's to lean on. Coincidentally like Terence Tao, I've also been asking complex analysis queries* of LLM's, things I was trying to understand better in my working through textbooks. Their ability to interpret open-ended math questions, and quickly find distant…
Re: Terence Tao on O1
#28I'm so excited in anticipation of my near-term return to studying math, as an independent curiosity hobby. It's going to be epically fun this time around with LLM's to lean on. Coincidentally like Terence Tao, I've also been asking complex analysis queries* of LLM's, things I was trying to understand better in my working through textbooks. Their ability to interpret open-ended math questions, and quickly find distant…
How will we even measure this? Benchmarks are gamed/trained on and there is no way that there is much signal in the chatbot arena for these types of queries? I think in just a few month the average user will not be able to tell the difference in performance between the major models
Re: Terence Tao on O1
#29Once GPT is tuned more heavily on Lean (proof assistant) -- the way it is on Python -- I expect its usefulness for research level math to increase. I work in a field related to operations research (OR), and ChatGPT 4o has ingested enough of the OR literature that it's able to spit out very useful Mixed Integer Programming (MIP) formulations for many "problem shapes". For instance, I can give it a logic problem like "…
Nothing is static in the way things are moving.
Re: Terence Tao on O1
#30I'm so excited in anticipation of my near-term return to studying math, as an independent curiosity hobby. It's going to be epically fun this time around with LLM's to lean on. Coincidentally like Terence Tao, I've also been asking complex analysis queries* of LLM's, things I was trying to understand better in my working through textbooks. Their ability to interpret open-ended math questions, and quickly find distant…
How will you know if its answers are correct or not?