Live data from Hacker News

Caveman: Why use many token when few token do trick

github.com

81–90 of 396 posts

Re: Caveman: Why use many token when few token do trick

#81
post #5

Oh boy. Someone didn't get the memo that for LLMs, tokens are units of thinking . I.e. whatever feat of computation needs to happen to produce results you seek, it needs to fit in the tokens the LLM produces. Being a finite system, there's only so much computation the LLM internal structure can do per token, so the more you force the model to be concise, the more difficult the task becomes for it - worst case, you ca…

What do you mean? The page explicitly states: > cutting ~75% of tokens while keeping full technical accuracy. I have no clue if this claim holds, but alas, just pretending they did not address the obvious criticism, while they did, is at the very least pretty lazy. An explanation that explains nothing is not very interesting.

The author pretended they addressed the obvious criticism.

You can read the skill. They didn't do anything to mitigate the issue, so the criticism is valid.

Re: Caveman: Why use many token when few token do trick

#85

This is exactly what annoys me most. English is not suitable for computer-human interaction. We should create new programming and query languages for that. We are again in cobol mindset. LLM are not humans and we should stop talking to them as if they are.

Grug says Chinese more suitable, only few runes in word, each take single token. Is great.

Re: Caveman: Why use many token when few token do trick

#86
post #5

Oh boy. Someone didn't get the memo that for LLMs, tokens are units of thinking . I.e. whatever feat of computation needs to happen to produce results you seek, it needs to fit in the tokens the LLM produces. Being a finite system, there's only so much computation the LLM internal structure can do per token, so the more you force the model to be concise, the more difficult the task becomes for it - worst case, you ca…

What do you mean? The page explicitly states: > cutting ~75% of tokens while keeping full technical accuracy. I have no clue if this claim holds, but alas, just pretending they did not address the obvious criticism, while they did, is at the very least pretty lazy. An explanation that explains nothing is not very interesting.

The burden of proof is on the author to provide at least one type of eval for making that claim.

Re: Caveman: Why use many token when few token do trick

#87
post #5

Oh boy. Someone didn't get the memo that for LLMs, tokens are units of thinking . I.e. whatever feat of computation needs to happen to produce results you seek, it needs to fit in the tokens the LLM produces. Being a finite system, there's only so much computation the LLM internal structure can do per token, so the more you force the model to be concise, the more difficult the task becomes for it - worst case, you ca…

Grug says you quite right, token unit thinking, but empty words not real thinking and should avoid. Instead must think problem step by step with good impactful words.

Re: Caveman: Why use many token when few token do trick

#88
post #54
post #3

No articles, no pleasantries, and no hedging. He has combined the best of Slavic and Germanic culture into one :)

Both Slavic languages and German have complex declination systems for nouns, verbs, and adjectives. Which is unlike stereotypical caveman speech.

I speak German, Polish, and English fluently and my take is: German is very precise, almost mathematical, there is little room to be misunderstood. But it also requires the most letters. English is the quickest, get things done kind of language, very compressible , but also risks misunderstanding. Polish is the most fun, with endless possibilities of twisting and bending it's structures, but also lacking the ease of use of English or the precision of German. But it's clearly just my subjective take

Re: Caveman: Why use many token when few token do trick

#89

Earlier quoted context omitted.

>Some languages are far more expressive and specialized in logical conditions, conditionals, recursion and reasoning. Like eskimos have 100 words for snow, but for boolean algebra. This is simply not true.

Really? Because if one accepts that computer languages are languages, then it seems that we could identify one or two that are highly specialized in logical conditions etc. Prolog springs to mind.

Yes, really. The concept GP is alluding to is called the Sapir-Worf hypothesis, which is largely non scientific pop linguistics drivel. Elements of a much weaker version have some scientific merit.

Programming languages are not languages in the human brain nor the culture sense.

Re: Caveman: Why use many token when few token do trick

#90

Earlier quoted context omitted.

> much of the thinking happens in a much higher dimensional space that just happens to be decoded as text. What do you mean by that? It’s literally text prediction, isn’t it?

>It’s literally text prediction, isn’t it? you are discovering that the favorite luddite argument is bullshit

I don't consider these researchers luddites.

https://machinelearning.apple.com/research/illusion-of-think...

https://arxiv.org/abs/2508.01191

Post reply on HN