Live data from Hacker News

Learning to Reason with LLMs

openai.com

451–460 of 1001 posts

Re: Learning to Reason with LLMs

#451
post #429

This is incredible. In April I used the standard GPT-4 model via ChatGPT to help me reverse engineer the binary bluetooth protocol used by my kitchen fan to integrate it into Home Assistant. It was helpful in a rubber duck way, but could not determine the pattern used to transmit the remaining runtime of the fan in a certain mode. Initial prompt here [0] I pasted the same prompt into o1-preview and o1-mini and both c…

second is very blurry

When you click on the image it loads a higher res version.

Re: Learning to Reason with LLMs

#452

This is incredible. In April I used the standard GPT-4 model via ChatGPT to help me reverse engineer the binary bluetooth protocol used by my kitchen fan to integrate it into Home Assistant. It was helpful in a rubber duck way, but could not determine the pattern used to transmit the remaining runtime of the fan in a certain mode. Initial prompt here [0] I pasted the same prompt into o1-preview and o1-mini and both c…

Thanks for sharing this, incredible stuff.

Re: Learning to Reason with LLMs

#453

Student here. Can someone give me one reason why I should continue in software engineering that isn't denial and hopium?

If you are not a software engineer, you can't judge the correctness of any LLM answer on that topic, nor you know what are the right questions to ask.

From all my friends that are using LLMs, we software engineers are the ones that are taking the most advantage of it.

I am in no way fearful I am becoming irrelevant, on the opposite, I am actually very excited about these developments.

Re: Learning to Reason with LLMs

#454
post #439

The "safety" example in the "chain-of-thought" widget/preview in the middle of the article is absolutely ridiculous. Take a step back and look at what OpenAI is saying here "an LLM giving detailed instructions on the synthesis of strychnine is unacceptable, here is what was previously generated vs our preferred, neutered content " What's this obsession with "safety" when it comes to LLMs? "This knowledge is perfectly…

There are two basic versions of “safety” which are related, but distinct: One version of “safety” is a pernicious censorship impulse shared by many modern intellectuals, some of whom are in tech. They believe that they alone are capable of safely engaging with the world of ideas to determine what is true, and thus feel strongly that information and speech ought to be censored to prevent the rabble from engaging in wr…

Third version is "brand safety" which is, we don't want to be in a new york times feature about 13 year olds following anarchist-cookbook instructions from our flagship product

Re: Learning to Reason with LLMs

#455

How could it fail to solve some maths problems if it has a method for reasoning through things?

It's using tree search (tree of thoughts), driven by some RL-derived heuristics controlling what parts of the practically infinite set of potential responses to explore.

How good the responses are will depend on how good these heuristics are.

Re: Learning to Reason with LLMs

#456

This is incredible. In April I used the standard GPT-4 model via ChatGPT to help me reverse engineer the binary bluetooth protocol used by my kitchen fan to integrate it into Home Assistant. It was helpful in a rubber duck way, but could not determine the pattern used to transmit the remaining runtime of the fan in a certain mode. Initial prompt here [0] I pasted the same prompt into o1-preview and o1-mini and both c…

Wow, that is impressive! How were you able to use o1-preview? I pay for ChatGPT, but on chatgpt.com in the model selector I only see 4o, 4o-mini, and 4. Is o1 in that list for you, or is it somewhere else?

Available on ChatGPT Plus signature or only using the API?

Re: Learning to Reason with LLMs

#457

Generating more "think out loud" tokens and hiding them from the user... Idk if I'm "feeling the AGI" if I'm being honest. Also... telling that they choose to benchmark against CodeForces rather than SWE-bench.

Why not? Isn't that basically what humans do? Sit there and think for a while before answering, going down different branches/chains of thought?

Yes but with concepts instead of tokens spelling out the written representation of those concepts.

Re: Learning to Reason with LLMs

#458
post #54

Sounds great, but so does their "new flagship model that can reason across audio, vision, and text in real time" announced in May. [0] [0] https://openai.com/index/hello-gpt-4o/

That is in chatgpt now and it greatly improves chatgpt. What are you on to now?

Re: Learning to Reason with LLMs

#459

Student here. Can someone give me one reason why I should continue in software engineering that isn't denial and hopium?

I went from economics dropout waiter who built a app startup with $0 funding and $1M a year in revenue by midway through year 1, sold it a few years later, then went to Google for 7 years, and last year I left. I'm mentioning that because the following sounds darn opinionated and brusque without the context I've capital-S seen a variety of people and situations.

Sit down and be really honest with yourself. If your goal is to have a nice $250K+ year job, in a perfect conflict-free zone, and don't mind Dilbert-esque situations...that will evaporate. Google is full of Ivy Leaguers like that, who would have just gone to Wall Street 8 years ago, and they're perennially unhappy people, even with the comparative salary advantage. I don't think most of them even realize because they've always just viewed a career as something you do to enable a fuller life doing snowboarding and having kids and vacations in the Maldives, stuff I never dreamed of and still don't have an interest in.

If you're a bit more feral, and you have an inherent interest and would be doing it on the side no matter what job you have like me, this stuff is a godsend. I don't need to sit around trying to figure out Typescript edge functions in Deno, from scratch via Google, StackOverflow, and a couple books from Amazon, taking a couple weeks to get that first feature built. Much less debug and maintain it. That feedback loop is now like 10-20 minutes.

Re: Learning to Reason with LLMs

#460
post #439

Earlier quoted context omitted.

There are two basic versions of “safety” which are related, but distinct: One version of “safety” is a pernicious censorship impulse shared by many modern intellectuals, some of whom are in tech. They believe that they alone are capable of safely engaging with the world of ideas to determine what is true, and thus feel strongly that information and speech ought to be censored to prevent the rabble from engaging in wr…

Third version is "brand safety" which is, we don't want to be in a new york times feature about 13 year olds following anarchist-cookbook instructions from our flagship product

Very good point, and definitely another version of “safety”!
Post reply on HN