Live data from Hacker News

The Unreliability of LLMs and What Lies Ahead

verissimo.substack.com

121–130 of 164 posts

Re: The Unreliability of LLMs and What Lies Ahead

#121
I have been using LLM coding tools to make stuff which I had no chance of making otherwise. They are MVPs, and if anything ever got traction I am very aware that I would need to hire a real dev. For now, I am basically a PM and QA person.

What really concerns me is that the big companies on whose tools we all rely are starting to push a lot of LLM generated code without having increased their QA.

I mean, everybody cut QA teams in recent years. Are they about to make a comeback once big orgs realize that they are pushing out way more bugs?

Am I way off base here?

Re: The Unreliability of LLMs and What Lies Ahead

#122
post #115

Earlier quoted context omitted.

> It's the extraordinary lengths people will go to in order to protect the bullshit they acquired with far less scepticism That's the narrative bias at play. We all are subject to it, and for good reason. People need stories to help maintain a stable mental equilibrium and a sense of identity. Knowledge that contradicts the stories that form the foundation of their understanding of the world can be destabilizing, whi…

I think there is a certain sort of person who, if not "wants" this destabilization, doesn't really experience the alternative. People primarily relating to the world through irony, say. So, characteristically, socrates (, some stand up comedians, and the like) who trade in aporia -- this feeling of destablization.

> I think there is a certain sort of person who, if not "wants" this destabilization, doesn't really experience the alternative.

I agree, but the ability/willingness to engage in that kind of destabilizing irony itself comes from a certain stability, where you can mess with the margins of your own stories' contradictions, without putting the core of your stories under threat.

Re: The Unreliability of LLMs and What Lies Ahead

#123

Earlier quoted context omitted.

> It’s mostly right enough. Honestly this is why your experience is different: your expectations are different (and likely lower). I never find they are "mostly right enough", I find they are "mostly wrong in ways that range from subtle mistakes to extremely incorrect". The more subtly they are wrong, the worse I rate their output actually, because that is what costs me more time when I try to use them I want tools t…

They save me a tremendous amount of time, you just need to be smart about what you try to get them to do. _Busy work_ is what you want to focus on, not anything that takes a ton of domain knowledge and intelligence. Just as an example from today, i had a huge pile of yaml documents that needed to have some transformations done to them -- they were pretty simple and obvious, but I just went into cursor, give it a befo…

Isn't the article saying it's mainly useful for SW?

I'm an electrical engineer and the only cases LLMs useful were developing phyton scripts or translating a text into a foreign language that I'm fluently speaking.

They are absolutely garbage for anything electrical engineering related, even coding RTL.

Re: The Unreliability of LLMs and What Lies Ahead

#124
I think this misses some of the core problems and it suggests there are some more straight forward solutions. We have no solutions to this and the way we're treating this means we aren't going to come up with solutions.

Problem 1: Training

Using any method like RLHF, DPO, or such guarantees that we train our models to be deceptive.

This is because our metric is the Justice Potter metric: I know it when I see it. Well, you're assuming that this accurate. The original case was about defining porn and well... I don't think it is hard to see how people even disagree on this. Go on Reddit and ask if girls in bikinis are safe for work or not. But it gets worse. At times you'll be presented with the choice between two lies. One lie you know is a lie and the other lie you don't know it is. So which do you choose? Obviously the latter! This means we optimize our models to deceive us. This is true too when we come to the choice between truth and a lie we do not know is a lie. They both look like truths.

This will be true even in completely verifiable domains. The problem comes down to truth not having infinite precision. A lot of truth is contextually dependent. Things often have incredible depth, which is why we have experts. As you get more advanced those nuances matter more and more.

Problem 2: Metrics and Alignment

All metrics are proxies. No ifs, ands, or buts. Every single one. You cannot obtain direct measurements which are perfectly aligned with what you intend to measure.

This can be easily observed with even simple forms of measurements like measuring distance. I studied physics and worked as an (aerospace) engineer prior to coming to computing. I did experimental physics, and boy, is there a fuck ton more complexity to measuring things than you'd guess. I have a lot of rules, calipers, micrometers and other stuff at my house. Guess what, none of them actually agree on measurements. They all are pretty close, but they do differ within their marked precision levels. I'm not talking about my ruler with mm hatch marks being off by 1mm. RobertElderSoftware illustrates some of this in this fun video[0]. In engineering, if you send a drawing to a machinist and it doesn't have tolerances, you have actually not provided them measurements.

In physics, you often need to get a hell of a lot more nuanced. If you want to get into that, go find someone that works in an optics lab. Boy does a lot of stuff come up that throws off your measurements. It seems straight forward, you're measuring distances.

This gets less straightforward once we talk about measuring things that aren't concrete. What's a high fidelity image? What is a well written sentence? What is artistic? What is a good science theory? None of these even have answers and are highly subjective. The result of that is your precision is incredibly low. In other words, you have no idea how you align things. It is fucking hard in well defined practical areas, but the stuff we're talking about isn't even close to well defined. I'm sorry, we need more theory. And we need it fast. Ad hoc methods will get you pretty far, but you'll quickly hit a wall if you aren't pushing the theory alongside it. The theory sits invisible in the background, but it is critical to advancements.

We're not even close to figuring this shit out... We don't even know if it is possible! But we should figure out how to put bounds, because even bounding the measurements to certain levels of error provides huge value. These are certainly possible things to accomplish, but we aren't devoting enough time to them. Frankly, it seems many are dismissive. But you can't discuss alignment without understanding these basic things. It only gets more complicated, and very fast.

[0] https://www.youtube.com/watch?v=EstiCb1gA3U

Re: The Unreliability of LLMs and What Lies Ahead

#125

I think I'm settling on a "Gell-mann Amnesia" explanation of why people are so rabidly committed to the "acceptable veracity" of LLM output. When you don't know the facts, you're easily mislead by plausible-sounding analysis, and having been mislead -- a certain default prejudice to existing beliefs takes over. There's a significant asymmetry of effort in belief change vs. acquisition. I think there's also an ego-pro…

You need to be pushing much more data in than you're getting out. 40k tokens of input can result in 400 actual quality tokens of output. Not giving enough input to work off of will result in regressed output.

It's basically like a funnel, which can also be used the other way around if the user is okay with quirky side effects. It feels like a lot of people are using the funnel the wrong way around and complaining that it's not working.

Re: The Unreliability of LLMs and What Lies Ahead

#126
post #120

Earlier quoted context omitted.

Once they start making deals with the relevant organizations, book rooms, handle insurance, replacement hotels, etc, then they'll replace travel agents. These guys don't just Google a bunch of tickets you know.

We're getting into semantics now, but I'm talking about the kind of person who used to sit in a physical store, waiting for someone to walk by and go into the travel agency. In the 80's and 90's, this is how most people booked their holidays. It was labour intensive, people would spend some time talking with a travel agent in a store, who would have a good idea of the packages available, and be able to make recommend…

I imagine a travel agent would have local knowledge and connections, and would know the quality of the hotels they're trying to send to you, a high commission isn't worth it if your customer is unsatisfied and goes to a different agent for their next trip. Of course this is based on the assumption that the customer always wants to use a travel agent (an unrealistic assumption nowadays, because it's so easy to switch to the Internet).

Someone like Rick Steves(1) still goes to the destinations every summer to check out hotels, restaurants and local companies, I imagine someone with more budget would travel with his company rather than try their luck with some booking.com hotel with a high rating...

1: https://www.youtube.com/@RickStevesEuropeOfficial

Re: The Unreliability of LLMs and What Lies Ahead

#127
post #87

This is a good articulation of what is a real concern around the AI bull thesis. If a calculator works great 99% of the time you could not use that calculator to build a bridge. Using AI for more than code generation is still very difficult and requires a human in the loop to verify the results. Sometimes using AI ends up being less productive because you're spending all your time debugging it's outputs. It's great b…

>If a calculator works great 99% of the time you could not use that calculator to build a bridge.

That's happened before with far higher correctness rate than 99%, and it cost Intel $500M. Reliability and accuracy matter. https://en.wikipedia.org/wiki/Pentium_FDIV_bug

Re: The Unreliability of LLMs and What Lies Ahead

#128

I think I'm settling on a "Gell-mann Amnesia" explanation of why people are so rabidly committed to the "acceptable veracity" of LLM output. When you don't know the facts, you're easily mislead by plausible-sounding analysis, and having been mislead -- a certain default prejudice to existing beliefs takes over. There's a significant asymmetry of effort in belief change vs. acquisition. I think there's also an ego-pro…

You need to be pushing much more data in than you're getting out. 40k tokens of input can result in 400 actual quality tokens of output. Not giving enough input to work off of will result in regressed output. It's basically like a funnel, which can also be used the other way around if the user is okay with quirky side effects. It feels like a lot of people are using the funnel the wrong way around and complaining tha…

Sure, if you have a high-quality starting point and need refinement.

The issue is that the vast majority of user-facing LLM use cases are where people don't have these high-quality starting points. They don't have 40k tokens to make 400.

Re: The Unreliability of LLMs and What Lies Ahead

#129
post #103

Earlier quoted context omitted.

LLMs are a legitimate technology with legitimate applications. However in a desperate bid for a new iPhone moment to assure Wall Street that the fantasy of infinite growth in a finite world is possible, they have utterly lost the plot regarding what statistical analysis of words at scale is capable of doing. Useless? Far from it. The basis for a 300 billion company with no meaningful products after almost a decade wo…

You seriously underestimate the appeal of burning cycles on GPUs to get something cool, if barely useful, out. Cryptocurrencies are still very much alive, too.

> Cryptocurrencies are still very much alive, too.

Yeah, like I said, LLMs will be around. Frankly I think they'll be way more around than crypto which as far as the mainstream is concerned might as well be dead.

Re: The Unreliability of LLMs and What Lies Ahead

#130
post #5

> Internally, it uses a sophisticated, multi-path strategy, approximating the sum with one heuristic while precisely determining the final digit with another. Yet, if asked to explain its calculation, the LLM describes the standard 'carry the one' algorithm taught to humans. So, the LLM isn't just wrong, it also lies...

The LLM has no relevant capacities, either to tell the truth or to lie. In generates "appropriate" text, given a history of cases of appropriate textual structures. It is the person who reads this text as-if written by a person who imparts these capacities to the machine, who treats the text as meaningful. But almost no text the LLM generates could be said to be meaningful, if any. In the sense that if a two year old…

One could always argue that the lie is in the ear of the receiver 8-/

I would argue, that if the output of the LLM is to be interpreted as natural speech, and the output makes an authoritative statement, which is factually incorrect, but stated as if it were true, this is a lie.

The problem is that the tech is presented as if it did have the internal state, that you accurately describe it not having.

The lie in this example, is when it is prompted to describe the process by which it reached a result, and that description has no resemblance to the actual process by which it reached the result.

This isn't a misrepresentation of some external facts, but a complete fabrication, that does not represent how it reached that result, at all.

However many users will accept this information, since it only involves internal aspects of the tool itself.

The fact that the LLM doesn't have this introspective information, is part of exactly why LLMs are NOT intelligence, artificial or otherwise.

And yet they are being presented as such, also, a lie...

Post reply on HN