Live data from Hacker News

Learning to Reason with LLMs

openai.com

121–130 of 1001 posts

Re: Learning to Reason with LLMs

#121
post #39

A lot of skepticism here, but these are astonishing results! People should realize we’re reaching the point where LLMs are surpassing humans in any task limited in scope enough to be a “benchmark”. And as anyone who’s spent time using Claude 3.5 Sonnet / GPT-4o can attest, these things really are useful and smart! (And, if these results hold up, O1 is much, much smarter.) This is a nerve-wracking time to be a knowled…

[deleted]

Re: Learning to Reason with LLMs

#122
post #54

Sounds great, but so does their "new flagship model that can reason across audio, vision, and text in real time" announced in May. [0] [0] https://openai.com/index/hello-gpt-4o/

Yep, all these AI announcements from big companies feel like promises for the future rather than immediate solutions. I miss the days when you could actually use a product right after it was announced, instead of waiting for some indefinite "coming soon."

Re: Learning to Reason with LLMs

#123
post #13

after weighing multiple factors including user experience, competitive advantage, and the option to pursue the chain of thought monitoring, we have decided not to show the raw chains of thought to users.

[deleted]

Re: Learning to Reason with LLMs

#124
My first interpretation of this is that it's jazzed-up Chain-Of-Thought. The results look pretty promising, but i'm most interested in this:

> Therefore, after weighing multiple factors including user experience, competitive advantage, and the option to pursue the chain of thought monitoring, we have decided not to show the raw chains of thought to users.

Mentioning competitive advantage here signals to me that OpenAI believes there moat is evaporating. Past the business context, my gut reaction is this negatively impacts model usability, but i'm having a hard time putting my finger on why.

Re: Learning to Reason with LLMs

#125

> We have found that the performance of o1 consistently improves with more reinforcement learning (train-time compute) and with more time spent thinking (test-time compute). Wow. So we can expect scaling to continue after all. Hyperscalers feeling pretty good about their big bets right now. Jensen is smiling. This is the most important thing. Performance today matters less than the scaling laws. I think everyone has…

[deleted]

Re: Learning to Reason with LLMs

#127
post #8

The model performance is driven by chain of thought, but they will not be providing chain of thought responses to the user for various reasons including competitive advantage. After the release of GPT4 it became very common to fine-tune non-OpenAI models on GPT4 output. I’d say OpenAI is rightly concerned that fine-tuning on chain of thought responses from this model would allow for quicker reproduction of their resu…

That's unfortunate. When an LLM makes a mistake it's very helpful to read the CoT and see what went wrong (input error/instruction error/random shit)

Re: Learning to Reason with LLMs

#128
post #67
post #34

Earlier quoted context omitted.

And Apple's product line this year? Phones. Nothing to do with fruit. Almost 50 years of lying to people. Names should mean something!

Did Apple start their company by saying they will be selling apples?

What's the statement that OpenAI are making today which you think they're violating? There very well could be one and if there is, it would make sense to talk about it.

But arguments like "you wrote $x in a blog post when you founded your company" or "this is what the word in your name means" are infantile.

Re: Learning to Reason with LLMs

#130
post #13

after weighing multiple factors including user experience, competitive advantage, and the option to pursue the chain of thought monitoring, we have decided not to show the raw chains of thought to users.

Saying "competitive advantage" so directly is surprising.

There must be some magic sauce here for guiding LLMs which boosts performance. They must think inspecting a reasonable number of chains would allow others to replicate it.

They call GPT 4 a model. But we don't know if it's really a system that builds in a ton of best practices and secret tactics: prompt expansion, guided CoT, etc. Dalle was transparent that it automated re-generating the prompts, adding missing details prior to generation. This and a lot more could all be running under the hood here.

Post reply on HN