Live data from Hacker News

Don't believe the hype: why ChatGPT is not the “holy grail” of AI research

salon.com

131–140 of 164 posts

Re: Don't believe the hype: why ChatGPT is not the “holy grail” of AI research

#131
post #21
post #14

The first steam engines were also written off as being less powerful than a horse. The first electric motors were written off as being less powerful than steam engines. So it goes. I think both of these views can be true at the same time: ChatGPT (or, LLMs really) are revolutionary and they won't revolutionize the world the way technologists/researchers say. Early adopters will use the technology and do amazing thing…

A more modern example is the first iPhone, the first gen was bad even by the standards of the day. If you look passed the novelty of having a lightsaber app on your phone it was terrible.

No, because next word predictors are fundamentally limited in their capabilites and have interest problems. This isn't something you can just iterate on to fix. You need a different architecture.

Re: Don't believe the hype: why ChatGPT is not the “holy grail” of AI research

#132
post #31
post #14

The first steam engines were also written off as being less powerful than a horse. The first electric motors were written off as being less powerful than steam engines. So it goes. I think both of these views can be true at the same time: ChatGPT (or, LLMs really) are revolutionary and they won't revolutionize the world the way technologists/researchers say. Early adopters will use the technology and do amazing thing…

> I think both of these views can be true at the same time: ChatGPT (or, LLMs really) are revolutionary and they won't revolutionize the world the way technologists/researchers say. Yes, because people think that LLMs are almost AGI based on the social media reactions and can't imagine they still have unknown/unsolved problems. But if we take a look at the 14 years of self driving car development, it becomes clear ho…

LLMs are different from self driving cars in that they can be useful even when they make wrong decisions occasionally. Copilots, document drafting (legal, copy, etc) and summarisation are useful services that people and enterprises are currently enjoying.

Re: Don't believe the hype: why ChatGPT is not the “holy grail” of AI research

#133
post #94

Earlier quoted context omitted.

One of the biggest problems in AI for the last 60 years is the grounding problem. The ability of a model to be rooted in objective reality. In other words, for one of these LLMs to understand when they are being accurate vs hallucinating. None of the current crop of LLMs has come close to solving this problem. On the contrary, they make the problem blatantly obvious. No LLMs will achieve AGI until this is solved suff…

A LANGUAGE model cannot solve this because truth and fiction is not a property of LANGUAGE

In causal speech confidence in an answer is communicated; though maybe just as a pause, or even through tone.

Not the sort of data we'll have crop up in CommonCrawl.

Re: Don't believe the hype: why ChatGPT is not the “holy grail” of AI research

#134

Earlier quoted context omitted.

As long as the accuracy of an LLM’s output is unknowable, there’s going to be a pretty hard limit on the kinds of jobs these tools can “replace”. And its not at all clear that this fundamental problem can be fixed at all with the current approach.

The mistake is in believing that LLM's output should be deterministic to be useful. Human output is not deterministic. Fields with text-heavy output are already being upended by this. Being able to summarize long legal briefs, identify contract problems, do classification of discovery documents, or even write first drafts of common legal forms is already upending the legal discipline. Chat-based customer support agen…

It's not about being deterministic.

It's about the LLM itself having any way to determine whether what it says is true.

Re: Don't believe the hype: why ChatGPT is not the “holy grail” of AI research

#135
post #3

This critique asserts that ChatGPT propagates inaccuracies. While there may be instances where this holds true, the given example does not substantiate this claim. The article alleges that ChatGPT stated Russia has launched bears into space, an assertion that is evidently false. However, my own interaction with ChatGPT 4.0 contradicts this. When I posed the same question, the AI unequivocally responded that no nation…

This comment seems to be perpetuating another misunderstanding about how LLMs like ChatGPT work: the idea that they have a "model of the world" that is consistent, but sometimes incorrect.

They do not.

There is no reason why two different people, asking the same question of ChatGPT, would necessarily get the same inaccurate answer. There is also no reason why they would get similar correct answers.

What they will get is answers that are statistically likely based on the text in ChatGPT's training data.

Re: Don't believe the hype: why ChatGPT is not the “holy grail” of AI research

#136
post #14

The first steam engines were also written off as being less powerful than a horse. The first electric motors were written off as being less powerful than steam engines. So it goes. I think both of these views can be true at the same time: ChatGPT (or, LLMs really) are revolutionary and they won't revolutionize the world the way technologists/researchers say. Early adopters will use the technology and do amazing thing…

> The first steam engines were also written off as being less powerful than a horse.

This may not be the best example considering that steam engines were around since at least 20BCE[0] but the first successful application wasn't till almost 1700.

[0] https://en.wikipedia.org/wiki/Aeolipile

Re: Don't believe the hype: why ChatGPT is not the “holy grail” of AI research

#138

Earlier quoted context omitted.

The mistake is in believing that LLM's output should be deterministic to be useful. Human output is not deterministic. Fields with text-heavy output are already being upended by this. Being able to summarize long legal briefs, identify contract problems, do classification of discovery documents, or even write first drafts of common legal forms is already upending the legal discipline. Chat-based customer support agen…

It's not about being deterministic . It's about the LLM itself having any way to determine whether what it says is true .

Not understanding why this is an issue for LLMs but not humans.

This is a simple commercial decision to make governed by three factors.

1. What is the cost of making an error?

2. What is the cost of the human doing the work?

3. What is the likelihood of the human making an error?

It's just evaluating how much more likely AI is to make an error than a human, by the cost of that error, set against the savings by using fewer humans.

Look at the legal profession. Sometimes the cost of an error is high, but usually it is not. There are already tons of little errors in contracts and discovery, and today they're all human. And people are very expensive. There is a giant swath of legal work that looks very attractive to automate at less than 100% accuracy.

Customer service: people offer poor customer service all the time, and usually the cost of that error is low. Human customer service isn't as expensive as legal work, but it's still relatively expensive. Very attractive to automate at less than 100% accuracy.

Re: Don't believe the hype: why ChatGPT is not the “holy grail” of AI research

#139

This is yet another article about "ChatGPT" that doesn't contain the strings "GPT-4" or "GPT4".

I tried two of the "failure mode" examples given: The Russian bears in space, and walking across a river with an average depth of 3 feet. The examples fail in GPT-3.5 while GPT-4 gives the "correct" response. The article is already out of date.

The article was out of date when it was published on March 19. GPT-4 was released on March 14.

Re: Don't believe the hype: why ChatGPT is not the “holy grail” of AI research

#140

Sure, anyone that uses ChatGPT knows it's currently not perfect. But there's a presumption that these tools are going to keep improving over time. Which is presumably why the AI hype is so strong. Whether AI ends up displacing people from their jobs in the long term, well, that's impossible to know. Just because no technological advancement has ever done that in the past doesn't mean it will never happen in the futur…

As long as the accuracy of an LLM’s output is unknowable, there’s going to be a pretty hard limit on the kinds of jobs these tools can “replace”. And its not at all clear that this fundamental problem can be fixed at all with the current approach.

Humans can make mistakes and lie, and we've been able to deal with it by checking their work, giving feedback to help them improve, placing less trust in those who habitually lie, etc..

LLMs making mistakes and "hallucinating" can be dealt with in similar ways, and as this is an open area of research with lots of proposed solutions and probably many more in the years to come, we do/will have plenty of other ways to deal with it too.

Post reply on HN