Live data from Hacker News

GPT-5: "How many times does the letter b appear in blueberry?"

bsky.app

41–50 of 339 posts

Re: GPT-5: "How many times does the letter b appear in blueberry?"

#41
post #30

The technical explanations to why this happens with strawberry, blueberry and similar is a great way to teach people how LLM works (and not work) https://techcrunch.com/2024/08/27/why-ai-cant-spell-strawber... https://arbisoft.com/blogs/why-ll-ms-can-t-count-the-r-s-in-... https://www.runpod.io/blog/llm-tokenization-limitations

I don’t find the explanation about tokenization to be very compelling.

I don’t see any particular reason the LLM shouldn’t be able to extract the implications about spelling just because its tokens of “straw” and “berry”

Frankly I think that’s probably misleading. Ultimately the problem is that the LLM doesn’t do meta analysis of the text itself. That problem probably still exists in various forms even if its character level tokenization. Best case it manages to go down a reasoning chain of explicit string analysis.

Re: GPT-5: "How many times does the letter b appear in blueberry?"

#42

I think a lot of those trick questions outputting stupid stuff can be explained by simple economics. It's just not sustainable for OpenAI to run GPT at the best of its abilities on every request. Their new router is not trying to give you the most accurate answer, but a balance of speed/accuracy/sustainable cost on their side. (kind of) a similar thing happened when 4o came out, they often tinkered with it and the re…

This is not a demonstration of a trick question.

This is a demonstration of a system that delusionally refuses to accept correction and correct its misunderstanding (which is a thing that is fundamental to their claim of intelligence through reasoning).

Why would anyone believe these things can reason, that they are heading towards AGI, when halfway through a dialogue where you're trying to tell it that it is wrong it doubles down with a dementia-addled explanation about the two bs giving the word that extra bounce?

It's genuinely like the way people with dementia sadly shore up their confabulations with phrases like "I'll never forget", "I'll always remember", etc. (Which is something that... no never mind)

> Even OSS 20b gets it right the first time, I think the author was just mistakenly routed to the dumbest model because it seemed like an easy unimportant question.

Why would you offer up an easy out for them like this? You're not the PR guy for the firm swimming in money paying million dollar bonuses off what increasingly looks, at a fundamental level, like castles in the sand. Why do the labour?

Re: GPT-5: "How many times does the letter b appear in blueberry?"

#43

Earlier quoted context omitted.

This is a great idea. Like, if someone asked me to count the number of B's in your paragraph, I'd yeet it through `grep -o 'B' file.txt | wc -l` or similar, why would I sit there counting it by hand? As a human, if you give me a number on screen like 100000000, I can't be totally sure if that's 100 Million or 1 Billion without getting close and counting carefully. Should ought have my glasses. Mouse pointer helps som…

If you have to build an MCP for every system you aren’t building intelligence in the first place.

Why does it matter? I don't care whether it's intelligent, I just need it to be useful. In order to be useful it needs to start fucking up less, stat. In current form it's borderline useless.

Re: GPT-5: "How many times does the letter b appear in blueberry?"

#44

Earlier quoted context omitted.

This is a great idea. Like, if someone asked me to count the number of B's in your paragraph, I'd yeet it through `grep -o 'B' file.txt | wc -l` or similar, why would I sit there counting it by hand? As a human, if you give me a number on screen like 100000000, I can't be totally sure if that's 100 Million or 1 Billion without getting close and counting carefully. Should ought have my glasses. Mouse pointer helps som…

If you have to build an MCP for every system you aren’t building intelligence in the first place.

Fair criticism, but also this arguably would be preferable. For many use cases it would be strictly better, as you've built some sort of automated drone that can do lots of work but without preferences and personality.

Re: GPT-5: "How many times does the letter b appear in blueberry?"

#45
post #3

“It’s like talking to a PhD level expert” -Sam Altman https://www.youtube.com/live/0Uu_VJeVVfo?si=PJGU-MomCQP1tyPk

There must be smart people at openai who believe in what they're doing and absolutely cringe whenever this clown opens his mouth... like, I hope?

Re: GPT-5: "How many times does the letter b appear in blueberry?"

#46

I think a lot of those trick questions outputting stupid stuff can be explained by simple economics. It's just not sustainable for OpenAI to run GPT at the best of its abilities on every request. Their new router is not trying to give you the most accurate answer, but a balance of speed/accuracy/sustainable cost on their side. (kind of) a similar thing happened when 4o came out, they often tinkered with it and the re…

This is not a demonstration of a trick question. This is a demonstration of a system that delusionally refuses to accept correction and correct its misunderstanding (which is a thing that is fundamental to their claim of intelligence through reasoning). Why would anyone believe these things can reason, that they are heading towards AGI, when halfway through a dialogue where you're trying to tell it that it is wrong i…

the extra bounce was my favorite part!

Re: GPT-5: "How many times does the letter b appear in blueberry?"

#47
post #30

The technical explanations to why this happens with strawberry, blueberry and similar is a great way to teach people how LLM works (and not work) https://techcrunch.com/2024/08/27/why-ai-cant-spell-strawber... https://arbisoft.com/blogs/why-ll-ms-can-t-count-the-r-s-in-... https://www.runpod.io/blog/llm-tokenization-limitations

When Minsky and Papert showed that the perceptron couldn't learn XOR, it contributed to wiping the neural network off the map for decades.

It seems no amount of demonstrating fundamental flaws in this system that should have been solved by all the new improved "reasoning" works anymore. People are willing to call these "trick questions", as if they are disingenuous, when they are discovered in the wild through ordinary interactions.

Does my tiny human brain in, this.

Re: GPT-5: "How many times does the letter b appear in blueberry?"

#48
For what it's worth, it got it right when I tried it.

>simple question should be easy for a genius like you. have many letter b's in the word blueberry? ChatGPT said:

>There are 2 letter b's in blueberry — one at the start and one in the middle.

Re: GPT-5: "How many times does the letter b appear in blueberry?"

#50

Earlier quoted context omitted.

This is not a demonstration of a trick question. This is a demonstration of a system that delusionally refuses to accept correction and correct its misunderstanding (which is a thing that is fundamental to their claim of intelligence through reasoning). Why would anyone believe these things can reason, that they are heading towards AGI, when halfway through a dialogue where you're trying to tell it that it is wrong i…

the extra bounce was my favorite part!

I mean if it was a Black Mirror satire moment it would rapidly become part of meme culture.

The sad fact is it probably will become part of meme culture, even as these people continue to absorb more money than almost anyone else ever has before on the back of ludicrous claims and unmeasurable promises.

Post reply on HN