Earlier quoted context omitted.
Except I actually mean to infer the concept of adding things from examples. LLMs are amply capable of applying concepts to data that matches patterns not ever expressed in the training data. It’s called inference for a reason. Anthropomorphic descriptions are the most expressive because of the fact that LLMs based on human cultural output mimic human behaviours, intrinsically. Other terminology is not nearly as expre…
>do manage to effectively convey the external effect But the problem is that this does not inform about the failure mode. So if I am understanding correctly, you are saying that the behavior of LLM, when it works, is like it has internalized the concepts. But then it does not inform that it can also say stuff that completely contradicts what it said before, there by also contradicting the notion of having "internaliz…
Caveman: Why use many token when few token do trick
381–390 of 396 posts
Re: Caveman: Why use many token when few token do trick
#382Call it Ix
Help caveman save even more tokens
Re: Caveman: Why use many token when few token do trick
#383Re: Caveman: Why use many token when few token do trick
#384Earlier quoted context omitted.
>do manage to effectively convey the external effect But the problem is that this does not inform about the failure mode. So if I am understanding correctly, you are saying that the behavior of LLM, when it works, is like it has internalized the concepts. But then it does not inform that it can also say stuff that completely contradicts what it said before, there by also contradicting the notion of having "internaliz…
If you look at the failure modes, they very closely resemble the failure modes of humans in equivalent situations. I'd say that, in practice, anthropomorphic view is actually the most informative we have about failure modes.
I don't think they do if we are talking about a honest human being.
LLMs will happily hallucinate and even provide "sources" for their wrong responses. That single thing should contradict what you are saying.
Re: Caveman: Why use many token when few token do trick
#385Earlier quoted context omitted.
Ah so obviously making the LLM repeat itself three times for every response it will get smarter
Yes, and observe that people do that too. It gives them more time to notice their own confusion and go "but wait, that's not right" on you.
## More tokens = smarter outputs
When an LLM uses tokens, it is putting more information into its context
## Better context, better results
The more information the LLM has in its context, the more complete and well thought-through the outputs will be
## More complete thinking
When an LLM is able to iterate on itself, results improve
## Better shareholder value
Numbers need to go up in order for us to maintain our shareholder value. This means instead of focusing on results that are qualitative, instead the brand should focus on quantitative, hard results
Re: Caveman: Why use many token when few token do trick
#386Earlier quoted context omitted.
> read much more research about LLMs than any human How long a response is from an LLM is going to be completely individual based on the system prompt and the model itself. You can read all of the "LLM research" in the world and it's not going to give you a correct generalized answer about this topic. It's not like this is some inherent property of LLMs.
FWIW, they also wrote down something that's so obvious you don't have to know much about LLMs to know that it's true. Even the "stochastic parrot" / "glorified Markov chain" / "regurgitation machine" camps people should be on the same page - LLMs are trained on human communication, and in human communications, longer queries, good manners and correct grammar are associated with longer, more correct and quality respon…
I use speech to text with Claude Code and other LLMs and often have terrible grammar and lots of typos and stuff and it never affects the output. But if I go by what you are saying then it would only seem right that the code it outputs is more sloppy? Also the length of a response entirely depends on what I'm using for example ChatGPT always gives me a long response no matter what I ask it and the Claude app always gives short responses unless I specifically ask for something longer. This is because of how they are given instructions and is not inherent to LLMs.
Re: Caveman: Why use many token when few token do trick
#387Idk I try talk like cavemen to claude. Claude seems answer less good. We have more misunderstandings. Feel like sometimes need more words in total to explain previous instructions. Also less context is more damage if typo. Who agrees? Could be just feeling I have. I often ad fluff. Feels like better result from LLM. Me think LLM also get less thinking and less info from own previous replies if talk like caveman.
I once (when ChatGPT first came out) launched into a conversation with ChatGPT using nothing but s-expressions. Didn't bother with a preamble, nor an explanation, just structured my prompt into a tree, forced said tree into an s-expression and hit enter. I was very surprised to see that the response was in s-expressions too. It was incoherent, but the parens balanced at least. Just tried it now and it doesn't seem to…
Re: Caveman: Why use many token when few token do trick
#388Earlier quoted context omitted.
How is your offering different from local ollama?
Its batteries included. No config. We also fine tuned and did RL on our model, developed a custom context engine, trained an embedding model, and modified MLX to improve inference. Everything is built to work with each other. So it’s more like an apple product than Linux. Less config but better optimized for the task.
Re: Caveman: Why use many token when few token do trick
#389Earlier quoted context omitted.
Its batteries included. No config. We also fine tuned and did RL on our model, developed a custom context engine, trained an embedding model, and modified MLX to improve inference. Everything is built to work with each other. So it’s more like an apple product than Linux. Less config but better optimized for the task.
I only understood half of the tech jargon in your answer. If I understood it all I’d probably run it myself. If someone who is less knowing than me is your customer, you need to explain in simpler terms!
All the benefits of claude code, without any of the limitations or rug pulls.
Re: Caveman: Why use many token when few token do trick
#390Earlier quoted context omitted.
Is this a joke, or are you serious? Do you work for Nvidia?
I’m not poster above but I work at Meta and they are doing this unfortunately. Wish it was a joke.