I've always figured that constraining an LLM to speak in any way other than the default way it wants to speak, reduces its intelligence / reasoning capacity, as at least some of its final layers can be used (on a per-token basis) either to reason about what to say, or about how to say it, but not both at once. (And it's for a similar reason, I think, that deliberative models like rewriting your question in their own…
Caveman: Why use many token when few token do trick
321–330 of 396 posts
Re: Caveman: Why use many token when few token do trick
#322Earlier quoted context omitted.
It is text prediction. But to predict text, other things follow that need to be calculated. If you can step back just a minute, i can provide a very simple but adjacent idea that might help to intuit the complexity of “ text prediction “ . I have a list of numbers, 0 to9, and the + , = operators. I will train my model on this dataset, except the model won’t get the list, they will get a bunch of addition problems. A…
>internalize the concepts. This gives the impression that it is doing something more than pattern matching. I think this kind of communication where some human attribute is used to name some concept in the LLM domain is causing a lot of damage, and ends up inadvertently blowing up the hype for the AI marketing...
Anthropomorphic descriptions are the most expressive because of the fact that LLMs based on human cultural output mimic human behaviours, intrinsically. Other terminology is not nearly as expressive when describing LLM output.
Pattern matching is the same as saying text prediction. While being technically truthy, it fails to convey the external effect. Anthropomorphic terms, while being less truthy overall, do manage to effectively convey the external effect. It does unfortunately imply an internal cause that does not follow, but the externalities are what matter in most non-philosophical contexts.
Re: Caveman: Why use many token when few token do trick
#323Earlier quoted context omitted.
My theory was that someone should write a specific LLM language, and then spend a whole lot of money to train models using that. A few times other commenters here have pointed out that that would be really difficult . But I think you're onto something, human languages just aren't optimal here. But to actually see this product to conclusion you'd probably need 60 to 100 million. You would have to completely invent a n…
I'm currently downloading Ollama and going to write a simple proof-of-concept with Qwen as local "frontend", talking to OpenAI GPT as "backend". I think the idea is sound, but indeed needs retraining of GPT (hmm like training tiny local LLM in synchronization of a big remote LLM). It might be not bad business venture in the end. I don't think humans should be involved in developing this AI-AI language, just giving so…
I imagine it's possible, but just a manner of money.
Re: Caveman: Why use many token when few token do trick
#324This is neat but my employer rates my performance based on token consumption; is there one that makes Claude needlessly verbose?
Is this a joke, or are you serious? Do you work for Nvidia?
Re: Caveman: Why use many token when few token do trick
#325Anyone else worried about the long term consequences of the influence of talking like this all day for the cognitive system of the user ?
I think good, less thinking for you, more thinking you will do
The verbose ones I could speed read, and consume it at a familiar pace... Almost on autopilot.
Caveman speak no familiar no convention, me no know first time. Need think hard understand. Slower. Good thing?
Re: Caveman: Why use many token when few token do trick
#326Re: Caveman: Why use many token when few token do trick
#327Is what cavemen sound like the same in every culture? Like I know that different cultures have different words for "woof" or "meow"; so it stands to reason maybe also for cavemans speech?
Re: Caveman: Why use many token when few token do trick
#328Re: Caveman: Why use many token when few token do trick
#329There is a reason it is not a common/popular technique.