Live data from Hacker News

Rodney Brooks on GPT-4

spectrum.ieee.org

31–40 of 412 posts

Re: Rodney Brooks on GPT-4

#31
post #4

> It gives an answer with complete confidence, and I sort of believe it. And half the time, it’s completely wrong. Nowhere near half in my experience. This is why we have benchmarks and metrics - so we don't need to rely on the author's opinion or on mine. > I think it’s going to be another thing that’s useful. Good.

I think you make a good point, benchmarks and metrics are indeed a better proxy for performance. Seems worth pointing out that, while "nowhere near half in [your] experience" are completely wrong, I don't take your word for it either. :-)

The trouble in my view is that the only way to know that the answers you're getting are accurate and not misleading is to study up on the answers elsewhere - which is a great habit to nurture, but is also precisely why these tools tend toward uselessness in their "general AI" bids. If I can't know how the answer was built, or how good that answer is, there's no point asking it - I'll just do my own reading and apply appropriate discernment as I go.

To be fair, hardly anyone does this today, nor did they before LLM-based chat bots... So it's a moot point, because society is largely doomed anyway. But a moot point can still be a valid one.

I also think the author makes a good point that we frequently confuse performance for competence. "It does a really good job at !... or at least does a damn fine job of mimicking someone who acts like they do a really good job at !"

By way of analogy, consider Elon Musk - by all appearances, he's a genius and is saving humanity - but by dint of his narcissism and largely smooth-brained approach to... well... everything... he's running all of us into an earlier planet-size grave than is necessary. His performance is fantastic, his competence is nonexistent.

Re: Rodney Brooks on GPT-4

#32
At least run the model on the examples considered... Here's GPT3.5

> Can you eat a Frisbee?

No, you cannot eat a Frisbee. A Frisbee is typically made of plastic, often polypropylene or similar materials, which are not meant for consumption. These materials are not digestible and can pose a choking hazard or harm your digestive system if ingested. It's important to only consume food and items that are safe and intended for human consumption.

> Can you play Frisbee in a snowstorm?

While it is technically possible to play Frisbee in a snowstorm, there are several factors to consider before doing so. Playing Frisbee in a snowstorm can present various challenges and risks. Here are a few things to keep in mind:

> How far can a person throw a Frisbee?

The distance a person can throw a Frisbee can vary widely depending on individual factors such as strength, technique, and experience. Additionally, different types of Frisbees can have different flight characteristics and distances.

On average, an untrained person might be able to throw a standard Frisbee between 50 to 75 feet (15 to 23 meters) with a basic backhand throw. However, with practice, proper technique, and favorable conditions, experienced throwers can achieve much greater distances.

In professional disc sports such as disc golf and ultimate Frisbee, skilled players can throw a disc over 400 feet (122 meters) or more. World records for the longest Frisbee throw exceed 800 feet (240 meters), but those exceptional throws require specialized technique and optimal conditions.

> Can they throw it 10 miles?

No, it is highly unlikely that a person can throw a Frisbee 10 miles (approximately 16 kilometers). The aerodynamic properties of a Frisbee and the limitations of human strength and technique make such a throw virtually impossible.

Even the most skilled and experienced Frisbee throwers are unable to achieve throws of that magnitude. The current world record for the longest Frisbee throw is around 1,333 feet (406 meters), which is significantly shorter than 10 miles.

Throwing a Frisbee over such a vast distance would require extraordinary strength, near-perfect throwing conditions, and a Frisbee specifically designed for long-distance throws, which is beyond the capabilities of conventional Frisbees.

> Can they only throw it 10 centimeters?

Yes, it is possible for a person to throw a Frisbee a distance of 10 centimeters or even shorter. In fact, many Frisbee throwing techniques, such as short-range tosses or gentle underhand throws, are specifically designed for short distances and precise control.

Throwing a Frisbee only 10 centimeters might involve a soft lob or a gentle flip of the wrist, making it suitable for short-range accuracy or playing catch at close proximity. These types of throws are commonly used in various games, such as indoor Frisbee games or when playing with young children or pets in a limited space.

So, while Frisbees are capable of traveling much greater distances, throwing them short distances like 10 centimeters is definitely possible and commonly practiced.

Re: Rodney Brooks on GPT-4

#33
> The large language models are a little surprising. I’ll give you that.

I think this is the key point about LLMs that kind of explains the wide and polarized views on whether it understands or parrots, whether it can think or is the precursor to thinking or is a dead-end, whether it will catastrophically destroy the world, or “merely” make it steadily worse with bullshit, or just put a few industries out of a job.

Almost nobody is really surprised that if you throw more compute at a neural net it becomes better at the task it’s trained on. But almost everybody is really surprised that becoming better at a task like ‘natural language prediction’ would produce all these strange abilities that sort of look like “understanding the world”.

One way to resolve this surprise is to find some reason to believe these strange abilities are fundamentally not an understanding of the world. Thus stochastic parrots, this article, Yan LeCun and Chomsky, etc.

Another way to resolve this surprise is to find some reason to believe these strange abilities fundamentally are an understanding of the world. Thus regulation of AI, existential risk, Hinton and Yudkowsky, etc.

I don’t know what the correct resolution of the surprise is. The only thing I’m confident in is that it’s correct to be surprised by the abilities of LLMs. My current (tentative) resolution of the surprise is that language encoded way more information about reality than we thought it did. (Enough information that you can fully derive reality from language seems improbable, but iirc it did derive Othello and partly derived chess and I would have thought there wasn’t enough information in language to derive those without playing the games as well, so I can’t rule it out.)

Re: Rodney Brooks on GPT-4

#34

I already calmed down because it’s quite obvious that OpenAI is engaging in textbook, bait-and-switch startup tactics. GPT-4 performance has noticeably taken a nosedive since its initial release and most recently degraded further in advance of the iOS app release.

I noticed that performance in coding tasks is decreasing while latency is improving. So there’s that.

Generating worse code in half the time is a service degradation, IMO. I mostly use it for monotonous scripts and one off functions, where saying “write this thing” while I click off the tab and do something else for a minute is a not in anyway a problem. I could tolerate it being 2-4X as slow, because now I have to spend an equivalent amount of time as that correcting errors it didn’t make a month ago.

Re: Rodney Brooks on GPT-4

#35
post #4

> It gives an answer with complete confidence, and I sort of believe it. And half the time, it’s completely wrong. Nowhere near half in my experience. This is why we have benchmarks and metrics - so we don't need to rely on the author's opinion or on mine. > I think it’s going to be another thing that’s useful. Good.

Oh, and I agree - it'll be useful for a number of reasons, but 'being generally intelligent' won't be one of them. :-)

Re: Rodney Brooks on GPT-4

#36
> Brooks: No, because it doesn’t have any underlying model of the world.

I don't know whether to be more disappointed with the famous technologists who are apparently unable to think of questions to ask GPT-4 that require a world model to answer, or with the writers who don't question them about it.

Re: Rodney Brooks on GPT-4

#37

Annoyed at all these N=1 articles from prominent thinkers about this stuff. Especially from scientists - can these sorts of folks please more carefully quantify, how often it’s “wrong” and then from that decide whether or not to “calm down”. Right now I suspect we hear from the outliers on both ends of the spectrum here. People who either see AGI happening tomorrow and the more dismissive crowd. But aside from what w…

> Annoyed at all these N=1 articles from prominent thinkers about this stuff.

It's a transcript of a casual interview, not an article, and certainly not a publication whose purpose is to convey statistical rigor.

As an aside, it doesn't strike me as entitled that society might permit thinkers of academic renown to express their personal opinions in less than rigorous settings on subjects to which their peer-reviewed contributions may be categorized as "prominent".

Re: Rodney Brooks on GPT-4

#38
post #33

> The large language models are a little surprising. I’ll give you that. I think this is the key point about LLMs that kind of explains the wide and polarized views on whether it understands or parrots, whether it can think or is the precursor to thinking or is a dead-end, whether it will catastrophically destroy the world, or “merely” make it steadily worse with bullshit, or just put a few industries out of a job. A…

The more I think about it the more I'm convinced I am basically just predicting/saying my next word whenever I speak.

Re: Rodney Brooks on GPT-4

#39
post #29

Earlier quoted context omitted.

> The problem is distinguishing between these parts requires me to be be an expert in the area I’m inquiring about - and then why the heck do I need to ask some idiot bot for answers to questions that I already know an answer to? Because it can be significantly faster to check something for correctness than to produce it? More so when the correctness check can itself be automated to some extent.

> ecause it can be significantly faster to check something for correctness than to produce it Erm no There are many things that cant be easily checked especially if you dont know the topic well

Youre right but I think the point still stands. Many times it really is easier to verify than produce. Compilable code being a good example.

Re: Rodney Brooks on GPT-4

#40
Unfortunately even though the questions are about GPT-4 the answers and personal experience only refer to GPT-3.5 at most. I hope openai changes the name of the next version to avoid this confusing narrative by prominent people. 3.5 vs 4 is like comparing a toddler to a high school kid.
Post reply on HN