If chain of thought acts as a scratch buffer by providing the model more temporary "layers" to process the text, I wonder if making this buffer a separate context with its own separate FNN and attention would make sense; in essence, there's a macroprocess of "reasoning" that takes unbounded time to complete, and then there's a microprocess of describing this incomprehensible stream of embedding vectors in natural lan…
Here's a paper your idea reminds me of. https://arxiv.org/abs/2501.19201 It's also so not far from Meta's large concept model idea.
[41 comments, 166 points] https://news.ycombinator.com/item?id=42919597