Live data from Hacker News

Learning to Reason with LLMs

openai.com

791–800 of 1001 posts

Re: Learning to Reason with LLMs

#791

Just added o1 to https://double.bot if anyone would like to try it for coding. --- Some thoughts: * The performance is really good. I have a private set of questions I note down whenever gpt-4o/sonnet fails. o1 solved everything so far. * It really is quite slow * It's interesting that the chain of thought is hidden. This is I think the first time where OpenAI can improve their models without it being immediately dis…

Trying out Double now.

o1 did a significantly better job converting a JavaScript file to TypeScript than Llama 3.1 405B, GitHub Copilot, and Claude 3.5. It even simplified my code a bit while retaining the same functionality. Very impressive.

It was able to refactor a ~160 line file but I'm getting an infinite "thinking bubble" on a ~420 line file. Maybe something's timing out with the longer o1 response times?

Re: Learning to Reason with LLMs

#792
post #8

The model performance is driven by chain of thought, but they will not be providing chain of thought responses to the user for various reasons including competitive advantage. After the release of GPT4 it became very common to fine-tune non-OpenAI models on GPT4 output. I’d say OpenAI is rightly concerned that fine-tuning on chain of thought responses from this model would allow for quicker reproduction of their resu…

I see chain of thought responses in chatgpt android app.

Tested cipher example, and it got it right. But "thinking logs" I see in the app looks like a summary of actual chain of thought messages that are not visible.

Re: Learning to Reason with LLMs

#793
Reminder that it's still not too late to change the direction of progress. We still have time to demand that our politicians put the breaks on AI data centres and end this insanity.

When AI exceeds humans at all tasks humans become economically useless.

People who are economically useless are also politically powerless, because resources are power.

Democracy works because the people (labourers) collectivised hold a monopoly on the production and ownership of resources.

If the state does something you don't like you can strike or refuse to offer your labour to a corrupt system. A state must therefore seek your compliance. Democracies do this by given people want they want. Authoritarian regimes might seek compliance in other ways.

But what is certain is that in a post-AGI world our leaders can be corrupt as they like because people can't do anything.

And this is obvious when you think about it... What power does a child or a disable person hold over you? People who have no ability to create or amass resources depend on their beneficiaries for everything including basics like food and shelter. If you as a parent do not give your child resources, they die. But your child does not hold this power over you. In fact they hold no power over you because they cannot withhold any resources from you.

In a post-AGI world the state would not depend on labourers for resources, jobless labourers would instead depend on the state. If the state does not provide for you like you provide for your children, you and your family will die.

In a good outcome where humans can control the AGI, you and your family will become subjects to the whims of state. You and your children will suffer as the political corruption inevitably arises.

In a bad outcome the AGI will do to cities what humans did to forests. And AGI will treat humans like humans treat animals. Perhaps we don't seek the destruction of the natural environment and the habitats of animals, but woodland and buffalo are sure inconvenient when building a super highway.

We can all agree there will be no jobs for our children. Even if you're an "AI optimist" we probably still agree that our kids will have no purpose. This alone should be bad enough, but if I'm right then there will be no future for them at all.

I will not apologise for my concern about AGI and our clear progress towards that end. It is not my fault if others cannot see the path I seem to see so clearly. I cannot simply be quiet about this because there's too much at stake. If you agree with me at all I urge you to not be either. Our children can have a great future if we allow them to have it. We don't have long, but we do still have time left.

Re: Learning to Reason with LLMs

#794
post #87

Reading through the Chain of Thought for the provided Cipher example (go to the example, click "Show Chain of Thought") is kind of crazy...it literally spells out every thinking step that someone would go through mentally in their head to figure out the cipher (even useless ones like "Hmm"!). It really seems like slowing down and writing down the logic it's using and reasoning over that makes it better at logic, simi…

Seriously. I actually feel as impressed by the chain of thought, as I was when ChatGPT first came out. This isn't "just" autocompletion anymore, this is actual step-by-step reasoning full of ideas and dead ends and refinement, just like humans do when solving problems. Even if it is still ultimately being powered by "autocompletion". But then it makes me wonder about human reasoning, and what if it's similar? Just fo…

Again its not reasoning.

Reasoning would imply that it can figure out stuff without being trained in it.

The chain of thought is basically just a more accurate way to map input to output. But its still a map, i.e forward only.

If an LLM coud reason, you should be able to ask it a question about how to make a bicycle frame from scratch with a small home cnc with limited work area and it should be able to iterate on an analysis of the best way to put it together, using internet to look up available parts and make decisions on optimization.

No LLM can do that or even come close, because there are no real feedback loops, because nobody knows how to train a network like that.

Re: Learning to Reason with LLMs

#795
post #8

The model performance is driven by chain of thought, but they will not be providing chain of thought responses to the user for various reasons including competitive advantage. After the release of GPT4 it became very common to fine-tune non-OpenAI models on GPT4 output. I’d say OpenAI is rightly concerned that fine-tuning on chain of thought responses from this model would allow for quicker reproduction of their resu…

It'd be helpful if they exposed a summary of the chain-of-thought response instead. That way they'd not be leaking the actual tokens, but you'd still be able to understand the outline of the process. And, hopefully, understand where it went wrong.

Exactly that I see in the Android app.

Re: Learning to Reason with LLMs

#796

I had trouble in the past to make any model give me accurate unix epochs for specific dates. I just went to GPT-4o (via DDG) and asked three questions: 1. Please give me the unix epoch for September 1, 2020 at 1:00 GMT. > 1598913600 2. Please give me the unix epoch for September 1, 2020 at 1:00 GMT. Before reaching the conclusion of the answer, please output the entire chain of thought, your reasoning, and the maths…

No need for llms to do that

ruby -r time -e 'puts Time.parse("2020-09-01 01:00:00 +00:00").to_i'

Re: Learning to Reason with LLMs

#797

Advanced reasoning will pave the way for recursive self-improving models & agents. These capabilities will enable data flywheels, error-correcting agentic behaviors, & self-reflection (agents understanding the implications of their actions, both individually & cooperatively). Things will get extremely interesting and we're incredibly fortunate to be witnessing what's happening.

this is completely illogical. this is like gambling your life savings and as the die are rolling you say "i am incredibly fortunate to be witnessing this." like, you need to know the outcome before you know whether it was fortunate or unfortunate... this could be the most unfortunate thing that has ever happened in history.

Re: Learning to Reason with LLMs

#798

Earlier quoted context omitted.

Well, the fact that you typed this question makes me think that you're in the top X% of students. That's your reason. Those in the bottom (100-X)% may be better off partying it up for a few years, but then again the same can be said for other AI-affected disciplines. Masseurs/masseuses have nothing to worry about.

I am pretty sure there is a VC funded startup making massage robots

Point taken, but I'm still pretty sure masseurs/masseuses have nothing to worry about.

Re: Learning to Reason with LLMs

#799

This is incredible. In April I used the standard GPT-4 model via ChatGPT to help me reverse engineer the binary bluetooth protocol used by my kitchen fan to integrate it into Home Assistant. It was helpful in a rubber duck way, but could not determine the pattern used to transmit the remaining runtime of the fan in a certain mode. Initial prompt here [0] I pasted the same prompt into o1-preview and o1-mini and both c…

The screenshot [1] is not readable for me. Chrome, Android. It's so blurry that I cant recognize a single character. How do other people read it? The resolution is 84x800.

Direct link to full resolution: https://i.postimg.cc/D74LJb45/SCR-20240912-sdko.png

Re: Learning to Reason with LLMs

#800

First shot, I gave it a medium-difficulty math problem, something I actually wanted the answer to (derive the KL divergence between two Laplace distributions). It thought for a long time, and still got it wrong, producing a plausible but wrong answer. After some prodding, it revised itself and then got it wrong again. I still feel that I can't rely on these systems.

Look where you were 3 years ago, and where you are now. And then imagine where you will be in 5 more years. If it can almost get a complex problem right now, I'm dead sure it will get it correct within 5 years

Getting complex problem = having the solution in some form in the training dataset.

All we are gonna get is better and better googles.

Post reply on HN