Live data from Hacker News

Bag of words, have mercy on us

experimental-history.com

131–140 of 362 posts

Re: Bag of words, have mercy on us

#131
post #66

Earlier quoted context omitted.

> When you have a thought, are you "predicting the next thing" Yes. This is the core claim of the Free Energy Principle[0], from the most-cited neuroscientist alive. Predictive processing isn't AI hype - it's the dominant theoretical framework in computational neuroscience for ~15 years now. > much of our experience of the world does not entail predicting things Introspection isn't evidence about computational archit…

How does the free energy principle align with system dynamics and the concept of emergence? Yes, our brain might want to optimize for lack of surprise, but that does not mean it can fully avoid emergent or chaotic behavior stemming from the incredibly complex dynamics of the linked neurons?

FEP doesn't conflict with complex dynamics, it's a mathematical framework for explaining how self-organizing behavior arises from simpler variational principles. That's what makes it a theory rather than a label.

The thing you're doing here has a name: using "emergence" as a semantic stopsign. "The system is complex, therefore emergence, therefore we can't really say" feels like it's adding something, but try removing the word and see if the sentence loses information.

"Neurons are complex and might exhibit chaotic behavior" - okay, and? What next? That's the phenomenon to be explained, not an explanation.

This was articulated pretty well 18 years ago [0].

[0]: https://www.lesswrong.com/posts/8QzZKw9WHRxjR4948/the-futili...

Re: Bag of words, have mercy on us

#132

> “Bag of words” is a also a useful heuristic for predicting where an AI will do well and where it will fail. “Give me a list of the ten worst transportation disasters in North America” is an easy task for a bag of words, because disasters are well-documented. On the other hand, “Who reassigned the species Brachiosaurus brancai to its own genus, and when?” is a hard task for a bag of words, because the bag just doesn…

When sensitivity analysis of ordinary least-squares regression became a thing it was also a "retrospective narrative". That seems reasonable for detecting fundamental issues with statistical models of the world. This point generalizes even if the concrete example falls down.

Re: Bag of words, have mercy on us

#133
post #19

I am unsure myself whether we should regard LLMs as mere token-predicting automatons or as some new kind of incipient intelligence. Despite their origins as statistical parrots, the interpretability research from Anthropic [1] suggests that structures corresponding to meaning do exist inside those bundles of numbers and that there are signs of activity within those bundles of numbers that seem analogous to thought. T…

Amanda Askell studied under David Chalmers at NYU: the philosopher who coined "the hard problem of consciousness" and is famous for taking phenomenal experience seriously rather than explaining it away. That context makes her choice to speak this way more striking: this isn't naive anthropomorphizing from someone unfamiliar with the debates. It's someone trained by one of the most rigorous philosophers of consciousne…

[deleted]

Re: Bag of words, have mercy on us

#134

> “Bag of words” is a also a useful heuristic for predicting where an AI will do well and where it will fail. “Give me a list of the ten worst transportation disasters in North America” is an easy task for a bag of words, because disasters are well-documented. On the other hand, “Who reassigned the species Brachiosaurus brancai to its own genus, and when?” is a hard task for a bag of words, because the bag just doesn…

I could not tell you who reassigned the species Brachiosaurus brancai to its own genus, and when, because of all the words I've ever heard, the combination of words that contains the information has not appeared.

GIGO has an obvious Nothing-In-Nothing-Out trivial case.

Re: Bag of words, have mercy on us

#135
The article is actually about the way we humans are extremely charitable when it comes to ascribing a ToM (theory of mind) and goes on to the Gym model of value. Nice. The comments drop back into the debate I originally saw Hinton describe on The Newyorker: do LLMs construct models (of the world) - that is do they think the way we think we think - or are they "glorified auto complete". I am going for the GAF view. But glorified auto complete is far more useful than the name suggests.

Re: Bag of words, have mercy on us

#136
post #34

Earlier quoted context omitted.

Yea bag of words isn’t helpful at all. I really do think that “superpowered sentence completion” is the best description. Not only is it reasonably accurate it is understandable, everyone has seen autocomplete function, and it’s useful. I don’t know how to “use” a bag of words. I do know how to use sentence completion. It also helps explains why context matters.

I've been recently using a similar description, referring to "AI" (LLMs) as "glorified autocomplete" or "luxury autocomplete".

I think I first heard "spicy autocomplete" two or three years ago...

Re: Bag of words, have mercy on us

#137

Best quote from the article: > That’s also why I see no point in using AI to, say, write an essay, just like I see no point in bringing a forklift to the gym. Sure, it can lift the weights, but I’m not trying to suspend a barbell above the floor for the hell of it. I lift it because I want to become the kind of person who can lift it. Similarly, I write because I want to become the kind of person who can think.

I don't really like the assumption that anyone who uses AI to, say, write an essay, is not the "kind of person who can think." And using AI to replace things you find recreational is not the point. If you got paid $100 each time you lifted a weight, would you see a point in bringing a forklift to the gym if it's allowed? Or will that make you a person who is so dumb that they cannot think, as the author is implying?

The same person could use a forklift at work, and lift weights manually at the gym.

Just pick the right tool for the job: don't take the forklift into the gym, and don't try to overhead press thousands of pounds that would fracture your spine.

Re: Bag of words, have mercy on us

#138

Earlier quoted context omitted.

> Everyone is out here acting like "predicting the next thing" is somehow fundamentally irrelevant to "human thinking" and it is simply not the case. Nobody is. What people are doing is claiming that "predicting the next thing" does not define the entirety of human thinking, and something that is ONLY predicting the next thing is not, fundamentally, thinking.

I claim that all of thinking can be reduced to predicting the next thing. Predicting the next thing = thinking in the same way that reading and writing strings of bytes is a universal interface, or every computation can be done by a Turing machine.

People can claim whatever they like. That doesn't mean it's a good or reasonable hypothesis (especially for one that is essentially unfalsifible like predictive coding).

Re: Bag of words, have mercy on us

#139

> “Bag of words” is a also a useful heuristic for predicting where an AI will do well and where it will fail. “Give me a list of the ten worst transportation disasters in North America” is an easy task for a bag of words, because disasters are well-documented. On the other hand, “Who reassigned the species Brachiosaurus brancai to its own genus, and when?” is a hard task for a bag of words, because the bag just doesn…

I tested this with ChatGPT-5.1 and Gemini 3.0. Both correctly (according to Wikipedia at least) stated that George Olshevsky assigned it to its own genus in 1991.

This is because there are many words about how to do web searches.

Re: Bag of words, have mercy on us

#140

I think a better metaphor is the Library of Babel. A practically infinite library where both gibberish and truth exist side by side. The trick is navigating the library correctly. Except in this case you can’t reliably navigate it. And if you happen to stumble upon some “future truth” (i.e. new knowledge), you still need to differentiate it from the gibberish. So a “crappy” version of the Library of Babel. Very impre…

It's like a highly compressed version of the Library. You're basically trying to discern real details from compression artifacts.
Post reply on HN