Live data from Hacker News

The One-Step Trap (In AI Research)

incompleteideas.net

1–10 of 13 posts

Re: The One-Step Trap (In AI Research)

#2
Ha, interesting. I wasn't aware of Sutton's blog post, but if I might make a shameless plug, we demonstrated [1] exactly this problem (see section 4.4.3), and how multi-step world models (using diffusion models as the substrate) could be one potential answer.

Since then, I have come to like temporally-abstract models more and more. Rolling out in time -- either step-by-step or many steps at once -- suffers from the tyranny of the specific. For long horizon planning with agents, I care (often only approximately) about where I can end up, and seldom about exactly when I end up there. Successor features, GVFs, Forward-Backward representations, and the like seem like they have an elegant approach for structuring thinking at a "high level", instead of generating exponentially large search trees by rolling out microscopic world models.

[1] https://arxiv.org/abs/2410.05364 (funnily, from around the same time / few months after Sutton's blog post)

Re: The One-Step Trap (In AI Research)

#4
This is the same reasoning behind why Yann Lecun thought test-time scaling would not work for LLMs: compounding error.

Instead, the more tokens LLMs use, the better their performance on many tasks. LLMs can self-correct, evidenced by the power of getting models to question themselves by emitting "Wait," in S1. https://arxiv.org/abs/2501.19393

Re: The One-Step Trap (In AI Research)

#6
post #4

This is the same reasoning behind why Yann Lecun thought test-time scaling would not work for LLMs: compounding error. Instead, the more tokens LLMs use, the better their performance on many tasks. LLMs can self-correct, evidenced by the power of getting models to question themselves by emitting "Wait," in S1. https://arxiv.org/abs/2501.19393

You wouldn't believe the amount of reasoning I saw these past few months that was correct until the stochastic parrot decided that a "wait" token should now be used and everything steered off a cliff.

Re: The One-Step Trap (In AI Research)

#7
post #2

Ha, interesting. I wasn't aware of Sutton's blog post, but if I might make a shameless plug, we demonstrated [1] exactly this problem (see section 4.4.3), and how multi-step world models (using diffusion models as the substrate) could be one potential answer. Since then, I have come to like temporally-abstract models more and more. Rolling out in time -- either step-by-step or many steps at once -- suffers from the t…

What do you mean by tyranny of the specific?

Re: The One-Step Trap (In AI Research)

#8
I'm not sure I follow what one step means exactly. Aren't all models some f(x) = y? Is the suggestion instead that we should be doing f(x) = g(h(x)) = y?

What would the difference be?

Re: The One-Step Trap (In AI Research)

#9
post #4

This is the same reasoning behind why Yann Lecun thought test-time scaling would not work for LLMs: compounding error. Instead, the more tokens LLMs use, the better their performance on many tasks. LLMs can self-correct, evidenced by the power of getting models to question themselves by emitting "Wait," in S1. https://arxiv.org/abs/2501.19393

Yeah came here to comment exactly this. And this is generally why I dislike/avoid this type of first principle analysis: it can make very convincing arguments that are just totally wrong due to some misleading assumption
Post reply on HN