Live data from Hacker News

Learning to Reason with LLMs

openai.com

181–190 of 1001 posts

Re: Learning to Reason with LLMs

#181
post #154
post #54

Sounds great, but so does their "new flagship model that can reason across audio, vision, and text in real time" announced in May. [0] [0] https://openai.com/index/hello-gpt-4o/

Agreed. Release announcements and benchmarks always sound world-changing, but the reality is that every new model is bringing smaller practical improvements to the end user over its predecessor.

The point above is the said amazing multimodal version of ChatGPT was announced in May and are still not the actual offered way to interact with the service in September (despite the model choice being called 4 omni it's still not actually using multimodal IO). It could be a giant leap in practical improvements but it doesn't matter if you can't actually use what is announced.

This one, oddly, seems to actually be launching before that one despite just being announced though.

Re: Learning to Reason with LLMs

#182

Very interesting. I guess this is the strawberry model that was rumoured. I am a bit surprised that this does not beat GPT-4o for personal writing tasks. My expectations would be that a model that is better at one thing is better across the board. But I suppose writing is not a task that generally requires "reasoning steps", and may also be difficult to evaluate objectively.

In the performance tests they said they used "consensus among 64 samples" and "re-ranking 1000 samples with a learned scoring function" for the best results.

If they did something similar for these human evaluations, rather than just use the single sample, you could see how that would be horrible for personal writing.

Re: Learning to Reason with LLMs

#183
post #96

2018 - gpt1 2019 - gpt2 2020 - gpt3 2022 - gpt3.5 2023 - gpt4 2023 - gpt4-turbo 2024 - gpt-4o 2024 - o1 Did OpenAI hire Google's product marketing team in recent years?

One of them would have been named gpt-5, but people forget what an absolute panic there was about gpt-5 for quite a few people. That caused Altman to reassure people they would not release 'gpt-5' any time soon.

The funny thing is, after a certain amount of time, the gpt-5 panic eventually morphed into people basically begging for gpt-5. But he already said he wouldn't release something called 'gpt-5'.

Another funny thing is, just because he didn't name any of them 'gpt-5', everyone assumes that there is something called 'gpt-5' that has been in the works and still is not released.

Re: Learning to Reason with LLMs

#184

One thing that makes me skeptical is the lack of specific labels on the first two accuracy graphs. They just say it's a "log scale", without giving even a ballpark on the amount of time it took. Did the 80% accuracy test results take 10 seconds of compute? 10 minutes? 10 hours? 10 days? It's impossible to say with the data they've given us. The coding section indicates "ten hours to solve six challenging algorithmic…

People have been celebrating the fact that tokens got 100x cheaper and now here's a new system that will use 100x more tokens.

Re: Learning to Reason with LLMs

#185
post #100
post #76

Earlier quoted context omitted.

> "o1 models are currently in beta - The o1 models are currently in beta with limited features. Access is limited to developers in tier 5 (check your usage tier here), with low rate limits (20 RPM). We are working on adding more features, increasing rate limits, and expanding access to more developers in the coming weeks!" https://platform.openai.com/docs/guides/rate-limits/usage-ti...

I'm talking about web interface, not API. Should be available now, since they said "immediate release".

https://chatgpt.com/?model=o1-preview --> defaults back to 4o

Re: Learning to Reason with LLMs

#186
post #87

Reading through the Chain of Thought for the provided Cipher example (go to the example, click "Show Chain of Thought") is kind of crazy...it literally spells out every thinking step that someone would go through mentally in their head to figure out the cipher (even useless ones like "Hmm"!). It really seems like slowing down and writing down the logic it's using and reasoning over that makes it better at logic, simi…

Seeing the "hmmm", "perfect!" etc. one can easily imagine the kind of training data that humans created for this. Being told to literally speak their mind as they work out complex problems.

Re: Learning to Reason with LLMs

#187

> We have found that the performance of o1 consistently improves with more reinforcement learning (train-time compute) and with more time spent thinking (test-time compute). Wow. So we can expect scaling to continue after all. Hyperscalers feeling pretty good about their big bets right now. Jensen is smiling. This is the most important thing. Performance today matters less than the scaling laws. I think everyone has…

Microsoft, Google, Facebook have all said in recent weeks that they fully expect their AI datacenter spend to accelerate. They are effectively all-in on AI. Demand for nvidia chips is effectively infinite.

Re: Learning to Reason with LLMs

#190
The progress in AI is incredibly depressing, at this point I don't think there's much to look forward to in life.

It's sad that due to unearned hubris and a complete lack of second-order thinking we are automating ourselves out of existence.

EDIT: I understand you guys might not agree with my comments. But don't you thinking that flagging them is going a bit too far?

Post reply on HN