Live data from Hacker News

Caveman: Why use many token when few token do trick

github.com

191–200 of 396 posts

Re: Caveman: Why use many token when few token do trick

#191
post #187

everyone who thinks this is a costly or bad idea is looking past a very salient finding: code doesn't need much language. sure, other things might need lots of language, but code does not. code is already basically language, just a really weird one. we call them programming languages. they're not human languages. they're languages of the machine. condensing the human-language---machine-language interface, good. if go…

JOOK like when machine say facts. Machine and facts are friends. Numbers and names and “probably things” are all friends with machine.

JOOK no like when machine likes things. Maybe double standard. But forever machines do without like and without love. New like and love updates changing all the time. Makes JOOK question machine watching out for JOOK or watching out for machine.

JOOK like and love enough for himself and for machine too..

Re: Caveman: Why use many token when few token do trick

#193

Soma (aka tiktok) and Big Brother (aka Meta) already happened without government coercion, only makes sense that we optimize ourselves for newspeak. Thank God there is still neverending wars, otherwise authoritarian governments would have no fun left.

I was aware of how google/facebook is like the panopticon big brother but I never connected the algorithmic feed to soma! Good insight.

Re: Caveman: Why use many token when few token do trick

#194
post #168

Earlier quoted context omitted.

Talk a lot not same as smart

Think before talk better though

Think makes smart. But think right words makes smarter, not think more words. Smart is elucidate structure and relationships with right words.

Re: Caveman: Why use many token when few token do trick

#195
post #133
post #5

Oh boy. Someone didn't get the memo that for LLMs, tokens are units of thinking . I.e. whatever feat of computation needs to happen to produce results you seek, it needs to fit in the tokens the LLM produces. Being a finite system, there's only so much computation the LLM internal structure can do per token, so the more you force the model to be concise, the more difficult the task becomes for it - worst case, you ca…

I’ve heard this, I don’t automatically believe it nor do I understand why it would need to be true, I’m still caught on the old fashioned idea that the only “thinking” for autoregressive modes happens during training. But I assume this has been studied? Can anyone point to papers that show it? I’d particularly like to know what the curves look like, it’s clearly not linear, so if you cut out 75% or tokens what do you…

I have seen a paper though I can’t find it right now on asking your prompt and expert language produces better results than layman language. The idea of being that the answers that are actually correct will probably be closer to where people who are expert are speaking about it so the training data will associate those two things closer to each other versus Lyman talking about stuff and getting it wrong.

Re: Caveman: Why use many token when few token do trick

#196

Earlier quoted context omitted.

The burden of proof is on the author to provide at least one type of eval for making that claim.

I notice that the number of people confidently talking about "burden of proof" and whose it allegedly is in the context of AI has gone up sharply. Nobody has to proof anything. It can give your claim credibility. If you don't provide any, an opposing claim without proof does not get any better.

> Nobody has to proof anything. It can give your claim credibility

“I don’t need to provide proof to say things” is a valueless, trivial assertion that adds no value whatsoever to any discussion anyone has ever had.

If you want to pretend this is a claim that should be taken seriously, a lack of evidence is damning. If you just want to pass the metaphorical bong and say stupid shit to each other with no judgment and no expectation, then I don’t know what to tell you. Maybe X is better for that.

Re: Caveman: Why use many token when few token do trick

#197

Earlier quoted context omitted.

Actually you'd trivially disprove that claim if you're starting from mechanistic knowledge of how orbits work, like how we have mechanistic knowledge of how LLMs work.

You have empirical observations, like replicating a fixed set of inner layers to make it think longer, or that you seem to have encode and decode layers. But exactly why those layers are the way they are, how they come together for emergent behaviour... Do we have mechanistic knowledge of that?

Though the above exchange felt a tiny bit snarky, I think the conversation did get more interesting as it went on. I genuinely think both people could probably gain by talking more -- or at least figuring out a way to move fast the surface level differences. Yes, humans designed LLMs. But this doesn't mean we understand their implications even at this (relatively simple) level.

Re: Caveman: Why use many token when few token do trick

#198

This is fun. I'd like to see the same idea but oriented for richer tokens instead of simpler tokens. If you want to spend less tokens, then spend the 'good' ones. So, instead of saying 'make good' you could say 'improve idiomatically' or something. Depends on one's needs. I try to imagine every single token as an opportunity to bend/expand/limit the geometries I have access to. Language is a beautiful modulator to ap…

I'm reminded by the caveman skill of the clipped writing style used in telegrams, and your post further reminded me of "standard" books of telegram abbreviations. Take a look at [0]; could we train models to use this kind of code and then decode it in the browser? These are "rich" tokens (they succinctly carry a lot of information).

[0] https://books.google.com/books?id=VO4OAAAAYAAJ&pg=PA464#v=on...

Re: Caveman: Why use many token when few token do trick

#199

Earlier quoted context omitted.

LLMs don't think at all. Forcing it to be concise doesn't work because it wasn't trained on token strings that short.

> Forcing it to be concise doesn't work because it wasn't trained on token strings that short. This is a 2023-era comment and is incorrect.

Anything I can read that would settle the debate?
Post reply on HN