Live data from Hacker News

GPT-5: "How many times does the letter b appear in blueberry?"

bsky.app

291–300 of 339 posts

Re: GPT-5: "How many times does the letter b appear in blueberry?"

#291

Earlier quoted context omitted.

What a terrible analogy. Illusions don't fool our intelligence, they fool our senses, and we use our intelligence to override our senses and see it for what it for it actually is - which is exactly why we find them interesting and have a word for them. Because they create a conflict between our intelligence and our senses. The machine's senses aren't being fooled. The machine doesn't have senses. Nor does it have int…

Analogies are just that, they are meant to put things in perspective. Obviously the LLM doesn't have "senses" in the human way, and it doesn't "see" words, but the point is that the LLM perceives (or whatever other word you want to use here that is less anthropomorphic) the word as a single indivisible thing (a token). In more machine learning terms, it isn't trained to autocomplete answers based on individual letter…

> the LLM perceives [...] the word as a single indivisible thing (a token).

Two actually, "blue" and "berry". https://platform.openai.com/tokenizer

"b l u e b e r r y" is 9 tokens though, and it still failed miserably.

Re: GPT-5: "How many times does the letter b appear in blueberry?"

#292

Let's change this game a bit. Spell "understanding" in your head in reverse order without spending twice more time than forward mode. Can you? I can't. Does that mean we don't really understand even simple spelling? It is a fun activity to dunk on LLMs, but let's have some perspective here.

I can do it if I write the word once and look at it, which is exactly what a transformer based llm is supposed to do.

Re: GPT-5: "How many times does the letter b appear in blueberry?"

#293

The defensive stance of some of the people in this thread is telling. The absolute meltdown that’s going to occur when humanity full internalizes the fact that LLMs are not and will never be intelligent is going to be of epic proportions.

They are still more useful than you

Re: GPT-5: "How many times does the letter b appear in blueberry?"

#294
post #85
post #3

“It’s like talking to a PhD level expert” -Sam Altman https://www.youtube.com/live/0Uu_VJeVVfo?si=PJGU-MomCQP1tyPk

A lot of people confuse access to information with being smart. Because for humans it correlates well - usually the smart people are those that know a lot of facts and can easily manipulate them on demand, and the dumb people are those that can not. LLMs have unique capability of being both very knowledgeable (as in, able to easily access vast quantities of information, way beyond the capabilities of any human, PhD o…

The most reasonable assumption is that the CEO is using dishonest rhetoric to upsell the LLM, instead of taking your approach and assuming the CEO is confused about the LLM's capability.

There are savvy people who know when to say "don't tell me that information" because then it is never a lie, simply "I was not aware"

Re: GPT-5: "How many times does the letter b appear in blueberry?"

#295
post #254

Earlier quoted context omitted.

> a systemic issue It will remain a suggestion of a systemic issue until it will be clear that architecturally all checks are implemented and mandated.

It is clear it is not, given we have examples of models that handles these cases. I don't even know what you mean with "architecturally all checks are implemented and mandated". It suggests you may think these models work very differently to how they actually work.

> given we have examples of models that handles

The suggestions come from the failures, not from the success stories.

> what you mean with "architecturally all checks are implemented and mandated"

That NN-models have an explicit module which works as a conscious mind and does lucid ostensive reasoning ("pointing at things") reliably respected in their conclusion. That module must be stress-tested and proven as reliable. Success stories only result based are not enough.

> you may think these models work very differently to how they actually work

I am interested in how they should work.

Re: GPT-5: "How many times does the letter b appear in blueberry?"

#296

Earlier quoted context omitted.

throwing a pair of dice and getting exactly 2 can also happen on the first try. Doesn't mean the dice are a 1+1 calculating machine

I guess my point is that the parent comment says LLMs get this wrong, but presents no evidence for that, and two anecdotes disagree. The next step is to see some evidence to the contrary.

> LLMs get this wrong

I wrote that of «a dozen models, no one could count». All of those I tried, with reasoning or not.

> presents no evidence

Create an environment to test and look for the failures. System prompt like "count this, this and that in the input"; user prompt some short paragraph. Models, the latest open weights.

> two anecdotes disagree

There is a strong asymmetry between verification and falsification. Said falsification occurred in a full set of selected LLMs - a lot. If two classes are there, the failing class is numerous and the difference between the two must be pointed at clearly. Also since we believe that the failure will be exported beyond the case of counting.

Re: GPT-5: "How many times does the letter b appear in blueberry?"

#297

Earlier quoted context omitted.

It matters to the adult - who is also an user. LLMs do not deliver (they miss important qualities related to intelligence); they are here now; so they must be superseded. There is no excuse: they must be fixed urgently.

LLMs deliver pretty well on their intended functionality: they predict next tokens given a token history and patterns in their training data. If you want to describe that as fully intelligent, that's your call, but I personally wouldn't. And adding functionality that isn't directly related to improving token prediction is just bad practice in an already very complex creation. LLM tools exist for that reason: they're…

> given a token history and patterns in their training data. If you want to describe that as fully intelligent

No, I would call (an easy interpretation of) that an implementation of unintelligence. Following patterns is what an hearsay machine does.

The architecture you describe at the "token prediction" level collides with an architecture in which ideas get related with better justifications than frequent co-occurrance. Given that the outputs will be similar in form, and that "dubious guessers" are now in place, we are now bound to hurry towards the "certified guessers".

Re: GPT-5: "How many times does the letter b appear in blueberry?"

#298

I have done this test extensively days ago, on a dozen models: no one could count - all of them got results wrong, all of them suggested they can't check and will just guess. Until they will be able of procedural thinking they will be radically, structurally unreliable. Structurally delirious. And it is also a good thing that we can check in this easy way - if the producers patched the local fault only, then the abse…

https://claude.ai/share/e7fc2ea5-95a3-4a96-b0fa-c869fa8926e8 It's really not hard to get them to reach the correct answer on this class of problems. Want me to have it spell it backwards and strip out the vowels? I'll be surprised if you can find an example this model can't one shot.

(Can't see it now because of maintenance but of course I trust it - that some get it right is not the issue.)

> if you can find an example this model can't

Then we have a problem of understanding why some work and some do not, and we have a due diligence crucial problem of determining whether the class of issues indicated by the possibility of fault as shown by many models are fully overcome in the architectures of those which work, or whether the boundaries of the problem are just moved but still tainting other classes of results.

Re: GPT-5: "How many times does the letter b appear in blueberry?"

#299

I have done this test extensively days ago, on a dozen models: no one could count - all of them got results wrong, all of them suggested they can't check and will just guess. Until they will be able of procedural thinking they will be radically, structurally unreliable. Structurally delirious. And it is also a good thing that we can check in this easy way - if the producers patched the local fault only, then the abse…

I tested it the other day and Claude with Reasoning got it correct every time

The interesting point is that many fail (100% in the class I had to select), and that raises the question of the difference between the pass-class and fail-class, and the even more important question of the solution inside the pass-class being contextual or definitive.

Re: GPT-5: "How many times does the letter b appear in blueberry?"

#300

Earlier quoted context omitted.

It's an umwelt problem. Bats think we're idiots because we don't hear ultrasonic sound, and thus can't echolocate. And we call the LLMs idiots because they consume tokenized inputs, and don't have access to the raw character stream.

Do bat's know what senses humans have? Or have the concept of what a human is compared to other organisms or moving objects? What is this analogy?

Yeah, I wrote this in a bit too short a hand to meet the critics where they sit...

There's an immense history of humans studying animal intelligence, which has tended pretty uniformly to find that animals are more intelligent than we previously thought at any given point in time. There's a very long history of badly designed experiments which surface 'false negative' results, and are eventually overturned. A common favor in these experiments is that the design assumes that animals have the same prescriptions and/or interests as humans. (For example, trying to do operant conditioning using a color cue with animals who can't perceive the colors. Or tasks that are easy of you happen to have approachable thumbs... That kind of thing.) Experiments eventually come along which better meet the animals where they are, and find true positive results, and our estimation of the intelligence of animals creeps slightly higher.

In other words, humans, in testing intelligence, have a decided bias towards only acknowledging intelligence which is distinctly human, and failing to take into account umwelt.

LLMs have a very different umwelt than we do. If they fail a test which doesn't take that umwelt into account, it doesn't indicate non-intelligence. It is, in fact, very hard to prove non-intelligence, because intelligence is poorly defined. And we have tended consistently to make the definition loftier whenever we're threatened with not being special anymore.

Post reply on HN