Live data from Hacker News

AI isn’t good enough

skventures.substack.com

321–330 of 374 posts

Re: AI isn’t good enough

#321

Earlier quoted context omitted.

> It seems to me that the real question here is what is true human intelligence. IMHO the main weakness with LLMs is they can’t really reason. They can statistically guess their way to an answer - and they do so surprisingly well I will have to admit - but they can’t really “check” themselves to ensure what they are outputting makes any sense like humans do (most of the time) - hence the hallucinations.

As the other poster said, they can check themselves but this requires an iterative process where the output is fed back in as input. Think of LLMs as the output of a human's stream of consciousness: it is intelligent, but has a high chance of being riddled with errors. That's why we iterate on our first thoughts to refine them.

Why do we have to feed the input back? Why doesn’t it do it itself?

Maybe it’s because it can’t tell when it’s wrong and needs to “try again”, and we have to do it for them.

Re: AI isn’t good enough

#322

Earlier quoted context omitted.

Perfectly reasonable, isn't it? But in a sense when YOU learned language YOU also learned a world model. For instance when your teacher explains to you the difference between the tenses (had, have, will have) you realize that time is a thing that you need to think about. Even if you already had some sense of this, you now have it made explicit. Why should we say the LLM hasn't learned a world model when it's done wha…

The reason why the LLM apparent world model should not be considered to be the same as a human's world model is because of the modality of learning. The world model we learn as we learn a language includes the world model embedded in language. But the human world model includes models embedded in flailing about limbs, the permanence of an object, sounds and smells associated with walking through the world. Now, all t…

That makes sense, but isn't this a matter of presenting it with more models? Maybe a physical model discovered via video or something like that? Then it will be similar to what babies are trained with, images and sound. Tactile and olfactory would be similar.

By doing this you'd glue the words to sights, sounds, smells, etc.

But it also seems like this is already someone has thought of and is being explored.

Re: AI isn’t good enough

#323

Earlier quoted context omitted.

A calculator can do arithmetic faster than a human being. How well would a calculator do at proving Fermat’s Last Theorem?

Not well! But that's in no way relevant to the point, which is that we are demonstrably capable of creating machines that can perform tasks of intelligence that we cannot.

Sure, though that really began with the abacus. A skilled abacus user can perform calculations faster than most people can with a calculator. They practice until it's all muscle memory. I think this demonstrates that there's actually very little intelligence involved in arithmetic.

Re: AI isn’t good enough

#324

Earlier quoted context omitted.

> It seems to me that the real question here is what is true human intelligence. IMHO the main weakness with LLMs is they can’t really reason. They can statistically guess their way to an answer - and they do so surprisingly well I will have to admit - but they can’t really “check” themselves to ensure what they are outputting makes any sense like humans do (most of the time) - hence the hallucinations.

Apparently GPT-4 is getting pretty good at knowing when it's wrong: https://thezvi.substack.com/p/ai-26-fine-tuning-time#%C2%A7g... (They asked GPT-3.5 and GPT-4 "are you sure" to see if it would change its answer, both when the original answer was right, and when it was wrong)

Does it do that because it can check it’s own reasoning? Or is it just doing so because OpenAI programmed it to not show alternative answers if the probability of the current answer being right is significantly higher than the alternatives?

Re: AI isn’t good enough

#325

Earlier quoted context omitted.

Humans have repeatedly built things that are beyond their own physical and intellectual capabilities. A calculator can do math problems much more quickly than any human being.

A calculator can do arithmetic faster than a human being. How well would a calculator do at proving Fermat’s Last Theorem?

Sure, but how would the average guy fare at the same task?

Re: AI isn’t good enough

#326

Earlier quoted context omitted.

As the other poster said, they can check themselves but this requires an iterative process where the output is fed back in as input. Think of LLMs as the output of a human's stream of consciousness: it is intelligent, but has a high chance of being riddled with errors. That's why we iterate on our first thoughts to refine them.

Why do we have to feed the input back? Why doesn’t it do it itself? Maybe it’s because it can’t tell when it’s wrong and needs to “try again”, and we have to do it for them.

Because that's how LLM chatbots are designed. Those papers describe systems where the review process is automated for better results.

Re: AI isn’t good enough

#327

Earlier quoted context omitted.

Seems that ~100W is regularly quoted as the average human energy output, so that'd be something like 18 MWh for 20 years (assuming 100% efficiency of food input). I suppose there's also the energy cost of all the training infrastructure (daycare, school teachers, homes, transportation, etc.), and also all the energy consumption of all humans that have come before and built those cities and knowledge infrastructure (a…

I don't think it's fair to include the billion years of human evolution as a cost on the human side of the chart but not include it in the AI side. AIs didn't evolve themselves out of the primordial ooze, they were built by humans and required herculean human effort to develop and improve. They stand on our shoulders, yet they're still in infancy when it comes to capability. I have yet to see any evidence of an LLM b…

>The best use of LLMs that I've seen so far is as a boilerplate-producing autocomplete system. Considering that we have better ways to automate this (better programming languages that can abstract away the boilerplate), this is not very high praise.

I think it's in our nature as software people to look at their ability to work with code, but they're quite good when applied to general language tasks. I've been using them for summarisation and reading comprehension and they're quite effective. I've also been working with a teacher friend on seeing if they can generate well-scoring essays on highschool English essays (as always, the problem is prompting and context).

On code, GPT-3.5 (moreso GPT-4) seems to have a good ability to generate and translate smaller scale code problems (hundreds of lines in low-boilerplate languages) but yes, they're like an eternally junior engineer whose work you're constantly having to oversee for subtle bugs, and I don't know that it actually saves time.

I'm pretty sure people are working on different approaches to applying them to code, with better prompting+context from larger codebases, and multi-step processing (i.e. rather than just a single prompt->response, letting the model iterate through a few steps independently, possibly guided by other adversarial/supervisor agent instances, testcase generators, etc.)

>If you're going to include the whole animal on the human side, you need to include the whole supply chain on the LLM side. The cost of building all the fabs and doing all the R&D to develop and manufacture model training-specific computers (matrix multiplier hardware). Just like with crypto, these resources had to be diverted away from other things (e.g. causing the price of gamers' graphics cards to skyrocket). It's only fair to interrogate the ROI.

That's fair, although it gets complicated to work out numbers because we don't train many LLMs, whereas we're constantly training humans, each of whom cost the planet tons of CO2 emissions every year... and, of course, your point that LLMs just aren't very good yet. I fear that they're good enough (or appear to be to the layperson) that execs will replace customer support staff with them, even if the outcomes overall aren't as good.

Re: AI isn’t good enough

#328
post #232

Earlier quoted context omitted.

Blockchain was a complete bust, and middle managers need buzzwords to sell senior leadership. That’s my personal take on the current wave.

I am a mild LLM skeptic. But I find the response of "oh, it's all just post-crypto scamming" really weird. Crypto was a total scam. There was never a concrete, non-criminal (important caveat) application where crypto was easier than just using PayPal or whatever. LLMs are very imperfect and still have a lot of work to do, but they do actually do some job related tasks today. If I have JS snippet and I wish it were in…

> There was never a concrete, non-criminal (important caveat) application where crypto was easier than just using PayPal or whatever.

That isn't true.

You can use it anywhere that irreversibility matters. Suppose you're going to commit significant resources to the customer's request, so you charge them, commit the resources, deliver the goods, and then discover that they gave you a stolen credit card and you get a chargeback. Cryptocurrency avoids that.

You can use it to accept payments from all over the world. Someone in Asia or Africa may not be able to open a US bank account or get a US credit card, but if they can find a Bitcoin ATM to put their local currency into, they can pay you, or vice versa.

It allows you to pay for something over the internet without giving your name. There are situations where this is important.

The main impediment to using it is, ironically, regulatory. The IRS decided that it's an investment and not money so every time you want to use it for what it's actually supposed to be for, they treat it like a securities transaction where you have to fill out paperwork, even if you're just buying a pack of gum. Which makes it much less convenient for ordinary people to use than cash or credit cards which don't require this -- presumably on purpose in order to destroy its utility in the US.

But it can still be useful for people in countries that don't do this, or in the US if a less explicitly antagonistic regulatory environment could be established.

Re: AI isn’t good enough

#329

Earlier quoted context omitted.

A calculator can do arithmetic faster than a human being. How well would a calculator do at proving Fermat’s Last Theorem?

Sure, but how would the average guy fare at the same task?

Is GPT 4 the "average AI"?

Re: AI isn’t good enough

#330
post #184
post #153

Earlier quoted context omitted.

> LLMs won't get intelligent. That's a fact based on their MO. They are sequence completion engines. A system that could perfectly predict what I would do in response to any particular stimuli, as a continuing sequence, would be exactly as intelligent as me. > They can be fine tuned to specific tasks, but at their core, they remain stochastic parrot Othello GPT was an attempt at answering this exact question, it's a…

A system that could perfectly predict what I would do in response to any particular stimuli, as a continuing sequence, would be exactly as intelligent as me. That's certainly interesting but it's not a depiction of a LLM is it ? LLM's are not deterministic, and (perhaps) so are we so two non-deterministic systems can only occasionally align (or so I assume). Intuition says they may get "close enough", whatever that m…

> LLM's are not deterministic,

They definitely can be, but it doesn't matter.

> but I think you are making a giant assumption to the likes of since we can speed up matter to 1000km/h then IF we sped it up to light speed then ...[something]...

This is an odd comparison.

The point here is that a sequence prediction system can be as intelligent as the system it's predicting unless you invoke woo. That doesn't make llms intelligent but it means the argument that they just predict the next thing isn't enough to say the can't be.

Post reply on HN