> We have found that the performance of o1 consistently improves with more reinforcement learning (train-time compute) and with more time spent thinking (test-time compute). Wow. So we can expect scaling to continue after all. Hyperscalers feeling pretty good about their big bets right now. Jensen is smiling. This is the most important thing. Performance today matters less than the scaling laws. I think everyone has…
Learning to Reason with LLMs
61–70 of 1001 posts
Re: Learning to Reason with LLMs
#62I like the idea too that they turbocharged it by taking the limits off during the "thinking" state -- so if an LLM wants to think about horrible racist things or how to build bombs or other things that RLHF filters out that's fine so long as it isn't reflected in the final answer.
Re: Learning to Reason with LLMs
#63Using Claude 3 Opus I noticed it performs and while browsing the web for me. I don't guess that's a change in the model for doing reasoning.
Re: Learning to Reason with LLMs
#64> we are releasing an early version of this model, OpenAI o1-preview, for immediate use in ChatGPT Awesome!
Re: Learning to Reason with LLMs
#65Re: Learning to Reason with LLMs
#66A lot of skepticism here, but these are astonishing results! People should realize we’re reaching the point where LLMs are surpassing humans in any task limited in scope enough to be a “benchmark”. And as anyone who’s spent time using Claude 3.5 Sonnet / GPT-4o can attest, these things really are useful and smart! (And, if these results hold up, O1 is much, much smarter.) This is a nerve-wracking time to be a knowled…
Is it? They talk about 10k attempts to reach gold medal status in the mathematics olympiad, but zero shot performance doesn't even place it in the upper 50th percentile. Maybe I'm confused but 10k attempts on the same problem set would make anyone an expert in that topic? It's also weird that zero shot performance is so bad, but over a lot of attempts it seems to get correct answers? Or is it learning from previous a…
Re: Learning to Reason with LLMs
#67Congrats to OpenAI for yet another product that has nothing to do with the word "open"
And Apple's product line this year? Phones. Nothing to do with fruit. Almost 50 years of lying to people. Names should mean something!
Re: Learning to Reason with LLMs
#68A lot of skepticism here, but these are astonishing results! People should realize we’re reaching the point where LLMs are surpassing humans in any task limited in scope enough to be a “benchmark”. And as anyone who’s spent time using Claude 3.5 Sonnet / GPT-4o can attest, these things really are useful and smart! (And, if these results hold up, O1 is much, much smarter.) This is a nerve-wracking time to be a knowled…
I cannot, in fact, attest that they are useful and smart. LLMs remain a fun toy for me, not something that actually produces useful results.
If you're not seeing value in them, maybe it's because you're not looking at the right problems. Or maybe you're just not using them correctly. Either way, dismissing an entire field of research because it doesn't fit your narrow use case is pretty short-sighted.
FWIW, I've been using LLMs to generate production code and it's saved me weeks if not months. YMMV, I guess