Earlier quoted context omitted.
I would claim that o1 -> o3 is evidence of exactly that, and supposedly in half a year we will have even better reasoning models (further complexity horizon), so what could that be besides what I am describing.
Is there some breakthrough in reasoning between o1 and o3 that we are all missing. And no one cares what we may have in the future. OpenAI etc already have an issue with credibility.
A Knockout Blow for LLMs?
31–40 of 49 posts
Re: A Knockout Blow for LLMs?
#32Oh, another LLM skepticism paper from Apple. This paper from last year doesn't age well due to rapid proliferation of reasoning models. https://machinelearning.apple.com/research/gsm-symbolic
Is apple putting out these papers just to justify their seeming inability to properly integrate them into their software?
While everyone learned the bitter lesson, apple chose to focus on small on-device models even after the explosion of chatgpt.
Re: A Knockout Blow for LLMs?
#33Re: A Knockout Blow for LLMs?
#34The first figure in the paper with Accuracy vs Complexity makes the whole point moot. The authors find that the performance of Claude 3.7 collapses around complexity 3 while Claude 3.7 thinking collapsed around complexity 7. A massive improvement in the complexity horizon that can be dealt with. It's real, it's quantitative, so what's the point of philosophical atguments about whether it is truly "reasoning" or not.…
> bigger models trained on bigger data with bigger reasoning posttraining and better distillation will push the horizons further and further There is no evidence this is the case. We could be in an era of diminishing returns where bigger models do not yield substantial improvements in quality but instead they become faster, cheaper and more resource efficient.
The scaling laws themselves advertise diminishing returns, something like a natural log. This was never debated by AI optimists, so it's odd to suggest otherwise as if it contradicts anything the AI optimists have been saying.
The scaling laws are kind of a worst case scenario, anyway. They assume no paradigm shift in methodology. As we saw when the test-time scaling law was discovered, you can't bet on stasis here.
Re: A Knockout Blow for LLMs?
#35The authors speculate that this pattern is a consequence of reasoning models actually solving these puzzles by way of pattern-matching to training data, which covers some puzzles at greater depth than others.
Great. That's one possible explanation. How might you support it?
- You could systematically examine the training data, to see if less representation of a puzzle type there reliably correlates with worse LLM performance.
- You could test how successfully LLMs can play novel games that have no representation in the training data, given instructions.
- Ultimately, using mechanistic interpretability techniques, you could look at what's actually going on inside a reasoning model.
This paper, however, doesn't attempt any of these. People are getting way out ahead of the evidence in accepting its speculation as fact.
Re: A Knockout Blow for LLMs?
#36Playbook:
1) you want to "disprove" some version of AI. Doesn't really matter what.
Take a problem humans face. For example, an almost total inability to follow simple rules for a long time to make a calculation. It's almost impossible to get a human to do this.
Check if AI algorithms, which are algorithms made to imitate humans have this same problem. Now of course, in practice if they indeed have that problem, that is actually a success: algorithm made to imitate humans ... imitates humans succesfully, strengths and weaknesses! But of course, if you find it, you describe it as total proof this algorithm is worthless.
An easy source for these problems is of course computers. Anything humans use computers for ... it's because humans suck at doing it themselves. Keeping track of history or facts. Exact calculation. Symbolic computation. Logic (ie. exactly correct answers). More generally math and even positive sciences as a whole are an endless supply of such problems.
2) you want to "prove" some version of AI.
Find something humans are good at. Point out AIs do this. Humans are social animals so how about influencing other humans? From convincing your boss, or on a larger scale using a social network to win an election, right up to actual seduction. Use what humans use to do it, of course (ie. be inaccurate, lie, ...)
Point out what a great success this is. How magical it is that machines can now do this.
3) you want to make a boatload of money
Take something humans are good at but hate, have an AI do it for money.
Re: A Knockout Blow for LLMs?
#37Earlier quoted context omitted.
Is apple putting out these papers just to justify their seeming inability to properly integrate them into their software?
They seem to be very skeptical against Large models. While everyone learned the bitter lesson, apple chose to focus on small on-device models even after the explosion of chatgpt.
"See this is why we can't build with transformers and had to use JEPA and look how much better it is!"