Live data from Hacker News

Bag of words, have mercy on us

experimental-history.com

121–130 of 362 posts

Re: Bag of words, have mercy on us

#121
A lot of the confusion comes from forcing LLMs into metaphors that don’t quite fit — either “they're bags of words” or “they're proto-minds.” The reality is in between: large-scale prediction can look useful, insightful, and even thoughtful without being any of those things internally. Understanding that middle ground is more productive than arguing about labels.

Re: Bag of words, have mercy on us

#122
> “Bag of words” is a also a useful heuristic for predicting where an AI will do well and where it will fail. “Give me a list of the ten worst transportation disasters in North America” is an easy task for a bag of words, because disasters are well-documented. On the other hand, “Who reassigned the species Brachiosaurus brancai to its own genus, and when?” is a hard task for a bag of words, because the bag just doesn’t contain that many words on the topic

It is... such a retrospective narrative. It's so obvious that the author learned about this example first than came with the reasoning later, just to fit in his view of LLM.

Imaging if ChatGPT answered this question correctly. Would that change the author's view? Of course not! They'll just say:

> “Bag of words” is a also a useful heuristic for predicting where an AI will do well and where it will fail. Who reassigned the species Brachiosaurus brancai to its own genus, and when?” is an easy task for a bag of words, because the information has appeared in the words it memorizes.

I highly doubt this author has predicted that "bag of Words" can do image editing before OpenAI released that.

Re: Bag of words, have mercy on us

#123
post #66

Earlier quoted context omitted.

When you have a thought, are you "predicting the next thing"—can you confidently classify all mental activity that you experience as "predicting the next thing"? Language and society constrains the way we use words, but when you speak, are you "predicting"? Science allows human beings to predict various outcomes with varying degrees of success, but much of our experience of the world does not entail predicting things…

> When you have a thought, are you "predicting the next thing" Yes. This is the core claim of the Free Energy Principle[0], from the most-cited neuroscientist alive. Predictive processing isn't AI hype - it's the dominant theoretical framework in computational neuroscience for ~15 years now. > much of our experience of the world does not entail predicting things Introspection isn't evidence about computational archit…

How does the free energy principle align with system dynamics and the concept of emergence? Yes, our brain might want to optimize for lack of surprise, but that does not mean it can fully avoid emergent or chaotic behavior stemming from the incredibly complex dynamics of the linked neurons?

Re: Bag of words, have mercy on us

#124
post #57

The bag of words reminds me of the Chinese room. "The machine accepts Chinese characters as input, carries out each instruction of the program step by step, and then produces Chinese characters as output. The machine does this so perfectly that no one can tell that they are communicating with a machine and not a hidden Chinese speaker. The questions at issue are these: does the machine actually understand the convers…

Chinese room has been discussed to death of course. Here's one fun approach (out of 100s) : What if we answer the Chinese room with the Systems Reply [1]? Searle countered the systems reply by saying he would internalize the Chinese room. But at that point it's pretty much exactly the Cartesian theater[2] : with room, homunculus, implement. But the Cartesian theater is disproven, because we've cut open brains and the…

It just seemed like relevant background that the author might not have been aware of, adjacent and substantial enough to warrant a mention.

I think there is some validity to the Cartesian theater, in that the whole of the experience that we perceive with our senses is at best an interpretation of a projection or subset of "reality."

Re: Bag of words, have mercy on us

#125
post #19

I am unsure myself whether we should regard LLMs as mere token-predicting automatons or as some new kind of incipient intelligence. Despite their origins as statistical parrots, the interpretability research from Anthropic [1] suggests that structures corresponding to meaning do exist inside those bundles of numbers and that there are signs of activity within those bundles of numbers that seem analogous to thought. T…

> the interpretability research from Anthropic [1] suggests that structures corresponding to meaning do exist inside those bundles of numbers and that there are signs of activity within those bundles of numbers that seem analogous to thought I did a simple experiment - took a photo of my kid in the park, showed it to Gemini and asked for a "detailed description". Then I took that description and put it into a generat…

> How did these 2 models do it if not actually using language like a thinking agent?

By having a gazillion of other, almost identical pictures of kids in parks in their training data.

Re: Bag of words, have mercy on us

#126
I think a better metaphor is the Library of Babel.

A practically infinite library where both gibberish and truth exist side by side.

The trick is navigating the library correctly. Except in this case you can’t reliably navigate it. And if you happen to stumble upon some “future truth” (i.e. new knowledge), you still need to differentiate it from the gibberish.

So a “crappy” version of the Library of Babel. Very impressive, but the caveats significantly detract from it.

Re: Bag of words, have mercy on us

#127

> “Bag of words” is a also a useful heuristic for predicting where an AI will do well and where it will fail. “Give me a list of the ten worst transportation disasters in North America” is an easy task for a bag of words, because disasters are well-documented. On the other hand, “Who reassigned the species Brachiosaurus brancai to its own genus, and when?” is a hard task for a bag of words, because the bag just doesn…

Your conclusion seems super unfair to the offer, particularly your assumption, without reason as far as I can tell, that the author would obstinately continue to advocate for their conclusion in the face of new, contrary evidence.

Re: Bag of words, have mercy on us

#128
post #28

As usual with these, it helps to try to keep the metaphor used for downplaying AI, but flip the script. Let's grant the author's perception that AI is a "bag of words", which is already damn good at producing the "right words" for any given situation, and only keeps getting better at it. Sure, this is not the same as being a human. Does that really mean, as the author seems to believe without argument, that humans ne…

Her argument really only works if you institute new economic systems where humans don’t need to labor in order to eat or pay rent.

"Her"->"the"? (Or, who is "she" here?)

Either way, in what way is this relevant? If the human's labor is not useful at any price point to any entity with money, food or housing, then they presumably will not get paid/given food/housing for it.

Re: Bag of words, have mercy on us

#129

Best quote from the article: > That’s also why I see no point in using AI to, say, write an essay, just like I see no point in bringing a forklift to the gym. Sure, it can lift the weights, but I’m not trying to suspend a barbell above the floor for the hell of it. I lift it because I want to become the kind of person who can lift it. Similarly, I write because I want to become the kind of person who can think.

I don't really like the assumption that anyone who uses AI to, say, write an essay, is not the "kind of person who can think."

And using AI to replace things you find recreational is not the point. If you got paid $100 each time you lifted a weight, would you see a point in bringing a forklift to the gym if it's allowed? Or will that make you a person who is so dumb that they cannot think, as the author is implying?

Re: Bag of words, have mercy on us

#130

> “Bag of words” is a also a useful heuristic for predicting where an AI will do well and where it will fail. “Give me a list of the ten worst transportation disasters in North America” is an easy task for a bag of words, because disasters are well-documented. On the other hand, “Who reassigned the species Brachiosaurus brancai to its own genus, and when?” is a hard task for a bag of words, because the bag just doesn…

Your conclusion seems super unfair to the offer, particularly your assumption, without reason as far as I can tell, that the author would obstinately continue to advocate for their conclusion in the face of new, contrary evidence.

I literally pasted the sentence as a prompt to the free version of ChatGPT "Who reassigned the species Brachiosaurus brancai to its own genus, and when?"

and got ths correct reply from the "Bag of Words"

The species Brachiosaurus brancai was reassigned to its own genus by Michael P. Taylor in 2009 — he transferred it to the new genus Giraffatitan. BioOne +2 Mike Taylor +2

How that happened:

Earlier, in 1988, Gregory S. Paul had proposed putting B. brancai into a subgenus as Brachiosaurus (Giraffatitan) brancai, based on anatomical differences. Fossil Wiki +1

Then in 1991, George Olshevsky used the name Giraffatitan brancai — but his usage was in a self-published list and not widely adopted. Wikipedia +1

Finally, in 2009 Taylor published a detailed re-evaluation showing at least 26 osteological differences between the African material (brancai) and the North American type species Brachiosaurus altithorax — justifying full generic separation. BioOne +1

If you like — I can show a short timeline of all taxonomic changes of B. brancai.

--

As an author, you should write things that are tested or at least true. But they did a pretty bad job of testing this and are making assumptions that are not true. Then they're basing their argument/reasoning (restrospectively) on assumptions not gounded in reality.

Post reply on HN