Live data from Hacker News

Overcoming the limits of current LLMs

seanpedersen.github.io

31–40 of 111 posts

Re: Overcoming the limits of current LLMs

#31
We can't develop a universally coherent data set because what we understand as "truth" is so intensely contextual that we can't hope to cover the amount of context needed to make the things work how we want, not to mention the numerous social situations where writing factual statements would be awkward or disastrous.

Here are a few examples of statements that are not "factual" in the sense of being derivable from a universally coherent data set, and that nevertheless we would expect a useful intelligence to be able to generate:

"There is a region called Hobbiton where someone named Frodo Baggins lives."

"We'd like to announce that Mr. Ousted is transitioning from his role as CEO to an advisory position while he looks for a new challenge. We are grateful to Mr. Ousted for his contributions and will be sad to see him go."

"The earth is round."

"Nebraska is flat."

Re: Overcoming the limits of current LLMs

#32
post #28

Earlier quoted context omitted.

"Hallucinations" implies that someone isn't of sound mental state. We can argue forever about what that means for a LLM and whether that's appropriate, but I think it's absolutely the right attitude and approach to be taking toward these things. They simply do not behave like humans of sound minds, and "hallucinations" conveys that in a way that "confabulations" or even "bullshit" does not. (Though "bullshit" isn't b…

I disagree with this take because LLMs are, always, hallucinating. When they get things right it’s because they are lucky. Yes, yes, it’s more complicated than that, but the essence of LLMs is that they are very good at being lucky. So good that they will often give you better results than random search engine clicks, but not good enough to be useful for anything important. I think what calling the times they get thi…

To offer a satirical analogy: "Lastly, I want to reassure investors and members of the press that we take these concerns very seriously: Hindenburg 2 will only contain only normal and unreactive hydrogen gas, and not the rare and unusual explosive kind, which is merely a temporary hurdle in this highly dynamic and growing field."

Edit: It retrospect, perhaps a better analogy would involve gasoline, as its explosive nature is what's being actively being exploited in normal use.

Re: Overcoming the limits of current LLMs

#33

There has been steady improvement since the release of chat gpt into the wild, which is still only less than two years ago (easy to forget). I've been getting a lot of value out of chat gpt 4o, like lots of other people. I find with each model generation my dependence on this stuff for day to day work goes up as the soundness of its answers and reasoning improve. There are still lots of issues and limitations but it'…

This is very close to my use case with Claude 3.5, and I used to only write tests when I was forced to, now it is part of the routine to double check everything while improving the codebase. I also really enjoy the socratic discussions when thinking about new ideas. What it says is mostly generic Wikipedia quality but this is useful when I am exploring domains where I have knowledge gaps.

Re: Overcoming the limits of current LLMs

#34
post #29

Earlier quoted context omitted.

This is almost a year old, thoughts on it today?

LLMs still do not reason or plan. And nothing in their architecture, training, post-training points toward real reasoning as scaling continues. Thinking does not happen one token at a time.

I don't get why some people seem to think the only way to use a LLM is for next token prediction or AGI has to be bult using LLM alone.

You want planning, you can do monte carlo tree search and use LLM to evaluate which node to explore next. You want verifiable reasoning, you can ask it to generate code(an approach used by recent AI olympiad winner and many previous papers).

What is even "planning", finding desirable/optimal solutions to some constrained satisfaction problems? Is the llm based minecraft bot voyager not doing some kind of planning?

LLMs have their limitations. Then augment them with external data sources, code interpreters, give it ways to interact with real world/simulation environment.

Re: Overcoming the limits of current LLMs

#35

Does anyone really believe that having a good corpus will remove hallucinations? Is this article even written by a person? Hard to know; they have a real blog with real article, but stuff like this reads strangely. Maybe it's just not a native english speaker? > Hallucinations are certainly the toughest nut to crack and their negative impact is basically only slightly lessened by good confidence estimates and reliabl…

I've seen such articles more and more recently. In the past, when people had a vague idea, they had to do research before writing. During this process, they often realized some flaws and thoroughly revised the idea or gave up writing. Nowadays, research can be bypassed with the help of eloquent LLMs, allowing any vague idea to turn into a write-up.

Re: Overcoming the limits of current LLMs

#36
post #16

The article suggests a useful line of research. Train an LLM to detect logical fallacies and then see if that can be bootstrapped into something useful because it's pretty clear that all the issues with LLMs is the lack of logical capabilities. If an LLM was capable of logical reasoning then it would be obvious when it was generating made-up nonsense instead of referencing existing sources of consistent information.

I think we should start smaller and make them able to count first.

Yeah, you can train an LLM to recognize the vocabulary and grammatical features of logical fallacies... Except the nature of fallacies is that they look real on that same linguistic level, so those features aren't distinctive for that purpose.

Heck, I think detecting sarcasm would be an easier goal, and still tricky.

Re: Overcoming the limits of current LLMs

#37
post #28

Earlier quoted context omitted.

"Hallucinations" implies that someone isn't of sound mental state. We can argue forever about what that means for a LLM and whether that's appropriate, but I think it's absolutely the right attitude and approach to be taking toward these things. They simply do not behave like humans of sound minds, and "hallucinations" conveys that in a way that "confabulations" or even "bullshit" does not. (Though "bullshit" isn't b…

I disagree with this take because LLMs are, always, hallucinating. When they get things right it’s because they are lucky. Yes, yes, it’s more complicated than that, but the essence of LLMs is that they are very good at being lucky. So good that they will often give you better results than random search engine clicks, but not good enough to be useful for anything important. I think what calling the times they get thi…

But the point is, isn't hallucinating about having malformed, altered or out of touch input rather than producing inaccurate output yourself?

It is the memory pathways leading them astray. It could be thought of a memory system that at certain point any longer can't be fully sure if whatever connections they have are from actually being trained or it or created accidentally.

Re: Overcoming the limits of current LLMs

#38
My biggest problem with them is that I can't quite get it to behave like I want it to. I built myself a "therapy/coaching" telegram bot (I'm healthy, but like to reflect a lot, no worries). I even built a self-reflecting memory component that generates insights (sometimes spot on, sometimes random af). But the more I use it, the more I notice that neither the memory nor the prompt matters much. I just can't get it to behave like a therapist would. So in other words: I can't find the inputs to achieve a desirable prediction from the SOTA LLMs. And I think that's a bigger problem for them not to be a shallow hype.

Re: Overcoming the limits of current LLMs

#39
post #31

We can't develop a universally coherent data set because what we understand as "truth" is so intensely contextual that we can't hope to cover the amount of context needed to make the things work how we want, not to mention the numerous social situations where writing factual statements would be awkward or disastrous. Here are a few examples of statements that are not "factual" in the sense of being derivable from a u…

> We can't develop a universally coherent data set because

Yet every child seems to manage, when raised by a small village, over a period of about 18 years. I guess we just need to give these LLMs a little more love and attention.

Re: Overcoming the limits of current LLMs

#40

Does anyone really believe that having a good corpus will remove hallucinations? Is this article even written by a person? Hard to know; they have a real blog with real article, but stuff like this reads strangely. Maybe it's just not a native english speaker? > Hallucinations are certainly the toughest nut to crack and their negative impact is basically only slightly lessened by good confidence estimates and reliabl…

Thank you. It seems largely ignored that LLMs still sample from a set of tokens based on estimated probability and the given temperature - but not on factuality or the described "confidence estimate" in the article. RAG etc. only move the estimated probabilities into a more factually based direction, but do not change the sampling itself
Post reply on HN