Earlier quoted context omitted.
The screenshot [1] is not readable for me. Chrome, Android. It's so blurry that I cant recognize a single character. How do other people read it? The resolution is 84x800.
Click on it for full resolution
Learning to Reason with LLMs
921–930 of 1001 posts
Re: Learning to Reason with LLMs
#922Earlier quoted context omitted.
The screenshot [1] is not readable for me. Chrome, Android. It's so blurry that I cant recognize a single character. How do other people read it? The resolution is 84x800.
When you open on phone, switch to "desktop site" via browser three dots menu
Re: Learning to Reason with LLMs
#923Advanced reasoning will pave the way for recursive self-improving models & agents. These capabilities will enable data flywheels, error-correcting agentic behaviors, & self-reflection (agents understanding the implications of their actions, both individually & cooperatively). Things will get extremely interesting and we're incredibly fortunate to be witnessing what's happening.
this is completely illogical. this is like gambling your life savings and as the die are rolling you say "i am incredibly fortunate to be witnessing this." like, you need to know the outcome before you know whether it was fortunate or unfortunate... this could be the most unfortunate thing that has ever happened in history.
Re: Learning to Reason with LLMs
#924The "safety" example in the "chain-of-thought" widget/preview in the middle of the article is absolutely ridiculous. Take a step back and look at what OpenAI is saying here "an LLM giving detailed instructions on the synthesis of strychnine is unacceptable, here is what was previously generated vs our preferred, neutered content " What's this obsession with "safety" when it comes to LLMs? "This knowledge is perfectly…
Re: Learning to Reason with LLMs
#925Earlier quoted context omitted.
It's using tree search (tree of thoughts), driven by some RL-derived heuristics controlling what parts of the practically infinite set of potential responses to explore. How good the responses are will depend on how good these heuristics are.
That doesn't sound like a method for reasoning.
They are only vaguely describing the process:
"Similar to how a human may think for a long time before responding to a difficult question, o1 uses a chain of thought when attempting to solve a problem. Through reinforcement learning, o1 learns to hone its chain of thought and refine the strategies it uses. It learns to recognize and correct its mistakes. It learns to break down tricky steps into simpler ones. It learns to try a different approach when the current one isn’t working. This process dramatically improves the model’s ability to reason."
Re: Learning to Reason with LLMs
#926In practice, this implementation (through the Chat UI) is scary bad. It actively lies about what it is doing. This is what I am seeing. Proactive, open, deceit. I can't even begin to think of all the ways this could go wrong, but it gives me a really bad feeling.
How do you mean?
Re: Learning to Reason with LLMs
#927Re: Learning to Reason with LLMs
#928Earlier quoted context omitted.
The screenshot [1] is not readable for me. Chrome, Android. It's so blurry that I cant recognize a single character. How do other people read it? The resolution is 84x800.
Direct link to full resolution: https://i.postimg.cc/D74LJb45/SCR-20240912-sdko.png
Re: Learning to Reason with LLMs
#929This is incredible. In April I used the standard GPT-4 model via ChatGPT to help me reverse engineer the binary bluetooth protocol used by my kitchen fan to integrate it into Home Assistant. It was helpful in a rubber duck way, but could not determine the pattern used to transmit the remaining runtime of the fan in a certain mode. Initial prompt here [0] I pasted the same prompt into o1-preview and o1-mini and both c…
second is very blurry
Re: Learning to Reason with LLMs
#930One thing that makes me skeptical is the lack of specific labels on the first two accuracy graphs. They just say it's a "log scale", without giving even a ballpark on the amount of time it took. Did the 80% accuracy test results take 10 seconds of compute? 10 minutes? 10 hours? 10 days? It's impossible to say with the data they've given us. The coding section indicates "ten hours to solve six challenging algorithmic…
People have been celebrating the fact that tokens got 100x cheaper and now here's a new system that will use 100x more tokens.