Live data from Hacker News

The Unreliability of LLMs and What Lies Ahead

verissimo.substack.com

141–150 of 164 posts

Re: The Unreliability of LLMs and What Lies Ahead

#141
post #120

Earlier quoted context omitted.

Once they start making deals with the relevant organizations, book rooms, handle insurance, replacement hotels, etc, then they'll replace travel agents. These guys don't just Google a bunch of tickets you know.

We're getting into semantics now, but I'm talking about the kind of person who used to sit in a physical store, waiting for someone to walk by and go into the travel agency. In the 80's and 90's, this is how most people booked their holidays. It was labour intensive, people would spend some time talking with a travel agent in a store, who would have a good idea of the packages available, and be able to make recommend…

Seems like the travel agent has been replace by:

1. the travel blogger who writes about places and why you might/might not want to go there.

2. the tour guide who books everything end to end for everyone on the tour and goes along to show and explain the sites.

Re: The Unreliability of LLMs and What Lies Ahead

#142
post #59

If I take a step back and think back to say a few (or 5) years ago, what LLMs can do is amazing. One has to acknowledge that (or at least, I do). But as a scientist it's been rather interesting to probe the jagged edge and unreliability, including using deep research tools, on any topic I know well. If I read through the reports and summaries it generates, it seems at first glance correct - the jargon is used correct…

You know who else is infamous for making errors due to shallow understanding ? (Non-specialized) journalists !

How do you find they compare?

Re: The Unreliability of LLMs and What Lies Ahead

#143
post #119

hallucinations are essentially the only thing keeping all knowledge workers from being made permanently redundant. if that doesnt make you a little concerned then you are a fool. and the predictions of all the experts in 2010 is that what is currently happening right in front of us could never happen within a hundred years. why are the predictions of experts more reliable now? anyone who dismisses the risks is just a…

I'm a knowledge worker (electrical engineer) but not one bit worried about being replaced by AI in yhe foreseeable future. It does not only neet to be reliable, but also should be able to create, as in create physically working complex systems for me to be worried. I have not seen anything remotely close this yet. I believe AI/ML will eventually get there but definitely not with LLMs or hoarding the whole internet. M…

you are. just change a few words around and you would be reading the confidently incorrect predictions of essentially all scientists and engineers in 2010. you say LLMs wont get us there… and you personally would probably have said word2vec couldnt get us past the turing test… and here we are. citing the existence of a current technology as evidence that another technology, related or not, cannot exist, is lazy and stupid. the simple fact is that there has been an explosion in the progress recently… a corresponding explosion of funding and the specific purpose of every single dollar of research is to create AGI, whether through LLMs or some other framework. to dismiss this situation as totally unconcerning is literally FOOLISH

Re: The Unreliability of LLMs and What Lies Ahead

#144

Earlier quoted context omitted.

Companies are salivating over the idea of cutting staff and replacing them with AI tools, so it's not exactly farfetched to think AI might lead to a lot of unemployment, at least for a while Every type of automation ever invented has led to massive job cuts and yes, some sectors actually did not ever recover

>Every type of automation ever invented has led to massive job cuts It, never has, in fact the opposite is true. Every type of automation has expanded the economic output so much that it created massive amounts of labor demand, which is why cities early absorbed masses of underemployed workers during the industrial revolution. One famous example, there are now more bank tellers than before the invention of the ATM. I…

> It, never has, in fact the opposite is true. Every type of automation has expanded the economic output so much that it created massive amounts of labor demand, …

This seems to be true, but there’s a second issue at work here, which that automation and progress in general can _disrupt_ the labor market. Sure there’s a net gain in labor demand, but there are people involved who are more than just “resources” who can easily be redeployed.

Progress is what built and then killed (injured?) cities and towns like, in the US, Detroit, or Gary, or Pittsburgh.

We want to promote progress and automation while at the same time protecting people who are inadvertently over-exposed to the downside. (Generally less educated people or people with less agency).

Re: The Unreliability of LLMs and What Lies Ahead

#145
Unreliability doesn't matter for some people because their bar was already that low. Unfortunately this is the way of the world and quality has and will continue to suffer. LLMs mostly accelerate this problem... hopefully they get good enough to help solve it.

Re: The Unreliability of LLMs and What Lies Ahead

#146
post #87

This is a good articulation of what is a real concern around the AI bull thesis. If a calculator works great 99% of the time you could not use that calculator to build a bridge. Using AI for more than code generation is still very difficult and requires a human in the loop to verify the results. Sometimes using AI ends up being less productive because you're spending all your time debugging it's outputs. It's great b…

A pedantic but maybe-not-entirely-pedantic point: It depends on what you mean by 99%.

If the calculator has a little gremlin in it that rolls a random 100-sided die, and gives you the wrong answer every time it rolls a 1, then you certainly can use it to build a bridge. You just need to do each calculation say 10 or 20 times and take the majority answer :)

If the gremlin is clever, it might remember the wrong answers it gave you, and then it might give them to you again if you ask about the same numbers. In that case you might need to buy 10 or 20 calculators that all have different gremlins in them, but otherwise the process is the same.

Of course if all your gremlins consistently lie for certain inputs, you might need to do a lot of work to sample all over your input space and see exactly what sorts of numbers they don't like. Then you can breed a new generation of gremlins that...

Re: The Unreliability of LLMs and What Lies Ahead

#147
post #87

This is a good articulation of what is a real concern around the AI bull thesis. If a calculator works great 99% of the time you could not use that calculator to build a bridge. Using AI for more than code generation is still very difficult and requires a human in the loop to verify the results. Sometimes using AI ends up being less productive because you're spending all your time debugging it's outputs. It's great b…

I believe it absolutely will. I think eventually we'll get to a point where people will be measured on now well they can get the AI to behave and how good they are at keeping cost down.

My boss built an AI workflow that cost over $600 that does the same thing I already gave him that cost less than $30. He just wanted to use tools he found and did it his way. Now, this had some value, it got more people in the company exposed to AI and he learned from the experience. It's his prerogative as he's the owner of the company. Though he also isn't concerned about the cost and will continue to pay much more. For now. I think as time goes on this will be more scrutinized.

Re: The Unreliability of LLMs and What Lies Ahead

#148
A few months ago I asked CGPT to create a max operating depth table for scuba diving based on various PPO2 limits and EAN gas profiles, just to test it on something I know (its a trivially easy calculation; and the formula is readily available online). It got it wrong…multiple times…even after correction and supplying the correct formula, the table was still repeatedly wrong (it did finally output a correct table). I just tried it again, with the same result. Obviously not something I would stake my life on anyway, but if it’s getting something so trivial wrong, I’m not inclined to trust it on more complex topics.

Re: The Unreliability of LLMs and What Lies Ahead

#149

Earlier quoted context omitted.

What we are seeing with our customers is that LLM errors are a very manageable problem. End users adapt pretty quickly to the idea that AI systems aren't perfect. In many cases AI products are doing tasks that used to be done by humans and these humans were making mistakes too, so the end user is used to the idea that the task will get accomplished with some non-zero error rate. You just need to build your products i…

> the user has the ability to easily double check the results whenever they like if the user is able to so easily verify that the results are accurate, that means that they are able to generate accurate results through other means, which means they don't need the LLM in the first place

True, but people are so enamored by what they can do that they rarely seem to think about that. We will over spend on AI purely because we think it's cool.

Re: The Unreliability of LLMs and What Lies Ahead

#150

A few months ago I asked CGPT to create a max operating depth table for scuba diving based on various PPO2 limits and EAN gas profiles, just to test it on something I know (its a trivially easy calculation; and the formula is readily available online). It got it wrong…multiple times…even after correction and supplying the correct formula, the table was still repeatedly wrong (it did finally output a correct table). I…

Well it doesn't really do math.
Post reply on HN