Live data from Hacker News

Caveman: Why use many token when few token do trick

github.com

321–330 of 396 posts

Re: Caveman: Why use many token when few token do trick

#321
post #239

I've always figured that constraining an LLM to speak in any way other than the default way it wants to speak, reduces its intelligence / reasoning capacity, as at least some of its final layers can be used (on a per-token basis) either to reason about what to say, or about how to say it, but not both at once. (And it's for a similar reason, I think, that deliberative models like rewriting your question in their own…

I find the inverse as well - asking a LLM to be chatty ends up with a much higher output. I've experimented with a few AI personality and telling it to be careful etc matters less than telling it to be talkative.

Re: Caveman: Why use many token when few token do trick

#322
post #228
post #125

Earlier quoted context omitted.

It is text prediction. But to predict text, other things follow that need to be calculated. If you can step back just a minute, i can provide a very simple but adjacent idea that might help to intuit the complexity of “ text prediction “ . I have a list of numbers, 0 to9, and the + , = operators. I will train my model on this dataset, except the model won’t get the list, they will get a bunch of addition problems. A…

>internalize the concepts. This gives the impression that it is doing something more than pattern matching. I think this kind of communication where some human attribute is used to name some concept in the LLM domain is causing a lot of damage, and ends up inadvertently blowing up the hype for the AI marketing...

Except I actually mean to infer the concept of adding things from examples. LLMs are amply capable of applying concepts to data that matches patterns not ever expressed in the training data. It’s called inference for a reason.

Anthropomorphic descriptions are the most expressive because of the fact that LLMs based on human cultural output mimic human behaviours, intrinsically. Other terminology is not nearly as expressive when describing LLM output.

Pattern matching is the same as saying text prediction. While being technically truthy, it fails to convey the external effect. Anthropomorphic terms, while being less truthy overall, do manage to effectively convey the external effect. It does unfortunately imply an internal cause that does not follow, but the externalities are what matter in most non-philosophical contexts.

Re: Caveman: Why use many token when few token do trick

#323

Earlier quoted context omitted.

My theory was that someone should write a specific LLM language, and then spend a whole lot of money to train models using that. A few times other commenters here have pointed out that that would be really difficult . But I think you're onto something, human languages just aren't optimal here. But to actually see this product to conclusion you'd probably need 60 to 100 million. You would have to completely invent a n…

I'm currently downloading Ollama and going to write a simple proof-of-concept with Qwen as local "frontend", talking to OpenAI GPT as "backend". I think the idea is sound, but indeed needs retraining of GPT (hmm like training tiny local LLM in synchronization of a big remote LLM). It might be not bad business venture in the end. I don't think humans should be involved in developing this AI-AI language, just giving so…

If you want to start a Git repo somewhere let me know and I'll do what I can to help.

I imagine it's possible, but just a manner of money.

Re: Caveman: Why use many token when few token do trick

#324
post #266

This is neat but my employer rates my performance based on token consumption; is there one that makes Claude needlessly verbose?

Is this a joke, or are you serious? Do you work for Nvidia?

I’m not poster above but I work at Meta and they are doing this unfortunately. Wish it was a joke.

Re: Caveman: Why use many token when few token do trick

#325

Anyone else worried about the long term consequences of the influence of talking like this all day for the cognitive system of the user ?

I think good, less thinking for you, more thinking you will do

I'm not sure if you're being sarcastic or not, but I did find the caveman examples harder to read than their verbose counterpart.

The verbose ones I could speed read, and consume it at a familiar pace... Almost on autopilot.

Caveman speak no familiar no convention, me no know first time. Need think hard understand. Slower. Good thing?

Re: Caveman: Why use many token when few token do trick

#330
Nothing against this project, it's been the case since forever that you could get better quality responses by simple telling your LLM to be brief and to the point, to ask salient questions rather than reflexively affirm, and eschew cliches and faddish writing styles.
Post reply on HN