Live data from Hacker News

People are just as bad as my LLMs

wilsoniumite.com

161–170 of 173 posts

Re: People are just as bad as my LLMs

#161

Earlier quoted context omitted.

Does "knowing what today is" count as "Outside STEM"? Coz my interactions with LLMs are certainly way worse than most people. Just tried it: tell me the current date please Today's date is October 3, 2023. Sorry ChatGPT, that's just wrong and your confidence in the answer is not helpful at all. It's also funny how different versions of GPT I've been interacting with always seem to return some date in October 2023, bu…

Me: tell me the current date please Chatgpt.com 4o: Today's date is March 11, 2025. Claud.ai 3.7 sonnet: The current date is Tuesday, March 11, 2025. gemini.google.com 2.0 flash: The current date is Tuesday, March 11, 2025. grok.com: The current date is March 10, 2025. amazon nova pro: The current date, according to the system information available to me, is March 11, 2025. Please keep in mind that my data and knowle…

Oh that's a great list!

Makes a lot of sense, thinking about it. I.e. the models that presumably have been given access to calling out to "live functions" can do stuff like that and/or have been specifically modified to answer such common questions correctly.

I also like it when they just tell you that they're a language model without such capabilities. That's totally fine and OK by me.

What I really don't like is the very confident answer with a specific date that is so obviously wrong. I guess the October 2023 thing is because I've been doing this with models where that's the end of training data and not others / retrained ones.

Re: People are just as bad as my LLMs

#162

Earlier quoted context omitted.

Does "knowing what today is" count as "Outside STEM"? Coz my interactions with LLMs are certainly way worse than most people. Just tried it: tell me the current date please Today's date is October 3, 2023. Sorry ChatGPT, that's just wrong and your confidence in the answer is not helpful at all. It's also funny how different versions of GPT I've been interacting with always seem to return some date in October 2023, bu…

These "LLMs cannot be AGI if they don't have a function to get today's date" remind me of laypeople reviewing phone cameras by seeing which camera's saturation they like more. It's absurd, whether an LLM has access to a function isn't a property of the LLM itself, therefore it's irrelevant, but people use it because LLMs make them feel bad somehow and they'll clutch at any straw.

No straws to clutch here. I've made such and other functions available to LLMs in order to implement some great functionality that would otherwise not have been possible. And they do a relatively good job. One of the issues is that they're not really reliable / deterministic. What the LLM does / is capable of today might not be what it does tomorrow or with just ever so slightly different context added via the prompts used by the user today vs. yesterday.

You are correct in that the date thing by itself, if that was the only thing would not be such a big deal.

But the date thing and confidently telling me the wrong date is a symptom and stand-in example of what LLMs will do in way too many situations and regular people don't understand this. Like I said, not very intelligent / confident people will do the same thing. But with people you generally have a "BS meter" and trust level. If you ask a random stranger on the street what time it is and they confidently tell you that it's exactly 11:20:32 a.m. without looking at their watch/phone, you know it's 99.99% BS. (again, just a stand in example, replace with 'Give me timeline of the most important thing that happened during WWII on a day by day basis' or whatever you can come up with). Yet people trust the output of LLMs with answers to questions where the user has no real way to know where on the BS meter this ranks. And they just believe them.

Happened to me today at work. LLM very confidently made up large swaths of data because it "figured out" that the test env we had was using the Star Trek universe characters and objects for test data. Had no base in reality and it basically had to ignore almost all the data that we actually returned from one of these "Get the current date" type functions we make available to it.

Thanks LLM!

Re: People are just as bad as my LLMs

#163

> ...a lot of the safeguards and policy we have to manage humans own unreliability may serve us well in managing the unreliability of AI systems too. It seems like an incredibly bad outcome if we accept "AI" that's fundamentally flawed in a way similar to if not worse than humans and try to work around it rather than relegating it to unimportant tasks while we work towards a standard of intelligence we'd otherwise ex…

Perhaps that kind of thing could help us finally move on from the "stupid should hurt" mindset to a real safety culture, where we value fault tolerance.

We like to pretend humans can reliably execute basic tasks like telling left from right or counting to ten, or reading a four digit number, and we assume that anyone who fails at these tasks is "not even trying"

But people do make these kinds of mistakes all the time, and some of them lead to patients having the wrong leg amputated.

A lot of people seem to see fault tolerance as cheating or relying on crutches, it's almost like they actively want mistakes to result in major problems.

If we make it so that AI failing to count the Rs doesn't kill anyone, that same attitude might help us build our equipment so that connecting the red wire to R2 instead of R3 results in a self test warning instead of a funeral announcement.

Obviously I'm all for improving the underlying AI tech itself ("Maintain Competence" is a rule in crew resource management), but I'm not a super big fan of unnecessary single points of failure.

Re: People are just as bad as my LLMs

#164

Earlier quoted context omitted.

No thank you. You've just explained "race to the bottom". We've had enough of this race, and it has left us with so many poor services and products.

The race to the bottom happens regardless whether you like it or not. Saying "no thank you" doesn't stop it. If only things in life were that easy.

Races to the bottom are incentive driven as anything else

Re: People are just as bad as my LLMs

#165
post #164

Earlier quoted context omitted.

The race to the bottom happens regardless whether you like it or not. Saying "no thank you" doesn't stop it. If only things in life were that easy.

Races to the bottom are incentive driven as anything else

Sure because the incentive is always quick and easy money.

Apple nearly went bankrupt in the late 90s early 00s by avoiding the race to the bottom of the PC industry till they pivoted to music players. Look at the auto makers today.

Unless you can convince customers why they should pay a premium for your commodity products, you will be wiped out by your competitors who do not refuse the race to the bottom.

Re: People are just as bad as my LLMs

#166
post #136

Earlier quoted context omitted.

Unless you can demonstrate that humans can solve a function that exceeds the Turing computable, it is reasonable to assume we're non more than Turing complete, and all Turing complete systems can compute the same set of functions. As it stands, we don't even know of any functions that exceeds the Turing complete, but are computable.

> As it stands, we don't even know of any functions that exceeds the Turing complete, but are computable. That would require the universe to be discrete, we don't know that. Otherwise most continuous processes compute something that a Turing machine can't, the Turing machine can only approximate it.

You can compute with values that are not discrete just fine by expressing them symbolically. I can't write out 1/3 as a discrete value in base 10, but I can still compute with it just fine.

Re: People are just as bad as my LLMs

#167
post #159

Earlier quoted context omitted.

Thanks for the explanation.. It still does not make sense to me.. A novel solution without deductive reasoning or a novel solution without empirical observation?

To be honest, I don't think their definition of intelligence is very coherent. I was just being pedantic. But if I had to guess, I believe they'd argue that an LLM is basically all a priori knowledge. It is trained on a massive data set and all it can do once trained is reason from those initial axioms (they aren't really axioms, but whatever). While humans, and actually many other animals to a lesser extent, can mak…

Humans derive their ideas from impressions (sensory experiences) and the ideas they form are essentially recombinations or refinements of those impressions. In this sense, human creativity can be viewed as a process of combining, transforming, and reinterpreting past experiences (impressions).

So, if we look at it from this perspective, human thinking is not fundamentally different from LLMs in that both rely on existing material to create new ideas.

The main difference is that LLMs process text statistically, while humans interpret text in context, influenced by emotions, experiences, biases, and goals. LLMs' interpretation is probabilistic, not conceptual.

Additionally, revolutionary thinking often requires rejecting past ideas and forming new conceptual frameworks, but LLMs cannot reject prior data, they are bound by it.

At any rate, the question remains, are LLMs capable of revolutionary ideas just like humans?

Re: People are just as bad as my LLMs

#168

Earlier quoted context omitted.

The idea that LLMs just repeat their training data is just wrong. It’s easy to test them and prove this is not the case. In some situations they may do that, typically when they don’t have much data on some topic. But in many other cases, it’s easy to verify that they are able to synthesize new output that is not simply a repetition of their training data. Software development is a great example, which also illustrat…

If you've implemented a sampler before, the "repeating the training data" is technically the logits array that you do the sampling on. Good samplers and sometimes even the most basic samplers can produce acceptable output but in the end the output is still technically just a bunch of training data predictions averaged together... or something roughly like that. The fact that I don't consider them intelligent doesn't…

First, to be clear, I'm not arguing that you should consider LLMs intelligent. I was responding more narrowly to the claim that an LLM "just repeats its training data."

On a trivial level, it's obviously true that every token in an LLM's output must have existed in the training data. But that's as far as your observation goes.

The point is that LLMs can produce novel sequences of tokens that exhibit the functional equivalent of "understanding" of the input and the expected output. Further, however they achieve it, functionally their output compares well to output that has been produced by a reasoning process.

None of this would be expected if they simply repeated "a bunch of training data predictions averaged together... or something roughly like that." For example, if that were all that was happening, you couldn't reasonably expect them to respond to a prompt with decent, working new code that's fit for purpose. They would produce code that looks plausible, but that doesn't compile, or run, or do what was intended.

One reason your model of the process fails to capture what's happening is because it's not taking into account the effects of latent space embeddings, and the resulting relationships between token representations. This is a major part of what enables an LLM to generalize and produce "correct" output, taking meaning into account, beyond simply repeating its training data.

As for intelligence - again the question comes down to functional equivalence. If we use traditional ways of measuring intelligence, like IQ tests, then LLMs beat the average human. Of course, that's not as significant as might naively be imagined, but it hints at the problem: how do you define intelligence, and on what basis are you claiming an LLM doesn't have it? Ultimately, it's a question of definitions, and I suspect it'd actually be quite difficult to give a rigorous (non-handwavy) definition of intelligence that an LLM can't satisfy. This may partly be an indictment of our understanding of intelligence.

Re: People are just as bad as my LLMs

#169

Earlier quoted context omitted.

To be honest, I don't think their definition of intelligence is very coherent. I was just being pedantic. But if I had to guess, I believe they'd argue that an LLM is basically all a priori knowledge. It is trained on a massive data set and all it can do once trained is reason from those initial axioms (they aren't really axioms, but whatever). While humans, and actually many other animals to a lesser extent, can mak…

Humans derive their ideas from impressions (sensory experiences) and the ideas they form are essentially recombinations or refinements of those impressions. In this sense, human creativity can be viewed as a process of combining, transforming, and reinterpreting past experiences (impressions). So, if we look at it from this perspective, human thinking is not fundamentally different from LLMs in that both rely on exis…

But the major difference between the human perceptual apparatus and data fed to an LLM is that humans are, in a linear temporal fashion, experiencing a physical world that exists outside of our perception. Our observations aren't just large volumes of unstructured data with purely statistical relevance to each other. Instead, we attempt to model the world via objects existing in relative position to each other and events occurring at various point in a timeline. The result is a complex model of cause and effect, actors and things being acted on, etc.

In that way, my dog is far more intelligent than LLM, in that he has a mental model of his world. An LLM is only intelligent relative to a human actor, and so it is no different than any other technology that humans have created to pursue their own ends.

Re: People are just as bad as my LLMs

#170

Earlier quoted context omitted.

If you've implemented a sampler before, the "repeating the training data" is technically the logits array that you do the sampling on. Good samplers and sometimes even the most basic samplers can produce acceptable output but in the end the output is still technically just a bunch of training data predictions averaged together... or something roughly like that. The fact that I don't consider them intelligent doesn't…

First, to be clear, I'm not arguing that you should consider LLMs intelligent. I was responding more narrowly to the claim that an LLM "just repeats its training data." On a trivial level, it's obviously true that every token in an LLM's output must have existed in the training data. But that's as far as your observation goes. The point is that LLMs can produce novel sequences of tokens that exhibit the functional eq…

> how do you define intelligence, and on what basis are you claiming an LLM doesn't have it? Ultimately, it's a question of definitions, and I suspect it'd actually be quite difficult to give a rigorous (non-handwavy) definition of intelligence that an LLM can't satisfy. This may partly be an indictment of our understanding of intelligence.

In my opinion there's nothing wrong with the traditional definition which is "the ability to acquire and apply knowledge and skills". But if you want to reach a minimum of "hand-waviness" then it's additionally required to define 'acquire', 'apply', and 'skills'. My personal definition is that acquiring knowledge requires building some sort of internal semantic model of it, though the occurrence of which there is actually evidence of in LLMs (see "abliteration"), so one out of three so far. But it falls apart at 'apply'. How do we even define applying? Well, I do not define it as what LLMs do, which is to predict the next token of the data.

I, personally, apply my knowledge by recognizing where it may be applicable, bringing it to thought, and then using that in the construction of ideas or strategies that I can act on. There's a degree of separation here between thought and action that doesn't currently seem to exist in LLMs; some creators are trying to simulate it by having an LLM for thoughts and another LLM for actions, or by enabling the thoughts to call tools that perform actions, or by having the LLM think before acting as in DeepSeek R1, but that isn't quite it.

An LLM still doesn't understand, say, spatial reasoning when it is helping me write something like a battle in a story. I have spatial reasoning because I can literally see what is happening while I write. I can see, and feel, and hear, and everything. Maybe that's just my dissociative disorder, but I will continue to await the day where LLMs might be able to do stuff like that. Until they can have that essentially happening "in their head", reason about it, and write using that, I won't believe that LLMs can "apply" much of anything just yet. (Other than machine learning I guess.)

> If we use traditional ways of measuring intelligence, like IQ tests, then LLMs beat the average human.

I think the whole notion of IQ is flawed because of neurodivergence. To put things in vaguely ableist-sounding terms (I don't mean it that way, but it will always sound that way), LLMs right now feel too neurotypical (pattern-based) and I would like to see future LLM developments that allow models leaning closer to autistic (logic-based).

Post reply on HN