Live data from Hacker News

OpenAI, Google and Anthropic are struggling to build more advanced AI

bloomberg.com

381–390 of 622 posts

Re: OpenAI, Google and Anthropic are struggling to build more advanced AI

#381
post #228
post #115

Question for the group here: do we honestly feel like we've exhausted the options for delivering value on top of the current generation of LLMs? I lead a team exploring cutting edge LLM applications and end-user features. It's my intuition from experience that we have a LONG way to go. GPT-4o / Claude 3.5 are the go-to models for my team. Every combination of technical investment + LLMs yields a new list of potential…

Great question. Im very confident in my answer, even though it’s in the minority here: we’re not even close to exhausting the potential. Imagine that our current capabilities are like the Model-T. There remains many improvements to be made upon this passenger transportation product, with RAG being a great common theme among them. People will use chatbots with much more permissive interfaces instead of clicking throug…

> why are so many experts so worried about AGI?

FYI, I find this line of reasoning to be unconvincing both logically and by counter-example ("why are so many experts so worried about the Y2K bug?")

Personally, I don't find AI foom or AI doom predictions to be probable but I do think there are more convincing arguments for your position than you're making here.

Re: OpenAI, Google and Anthropic are struggling to build more advanced AI

#382
post #142

What a stupid piece. We are making leaps every 6 months still. Tell me this when there are no developments for 3 years.

I'm curious, what was the leap after GPT-4? What about the leaps after that, given a leap every 6 months?

Sora was just one of the many…

Re: OpenAI, Google and Anthropic are struggling to build more advanced AI

#383
post #374
post #258

Earlier quoted context omitted.

> you can have LLMs create reasonable code changes, with automatic review / iteration etc. Nobody who takes code health and sustainability seriously wants to hear this. You absolutely do not want to be in a position where something breaks, but your last 50 commits were all written and reviewed by an LLM. Now you have to go back and review them all with human eyes just to get a handle on how things broke, while custom…

If the last 50 commits were reviewed by an AI and it took that long for an issue to happen I’d immediately mandate all PR’s are reviewed by an AI.

There's a difference between an issue being introduced and being noticed.

Re: OpenAI, Google and Anthropic are struggling to build more advanced AI

#384

Earlier quoted context omitted.

I'm curious, what was the leap after GPT-4? What about the leaps after that, given a leap every 6 months?

Sora was just one of the many…

Your best example is something that doesn't even do the things that GPT-4 does, isn't available to use, and has seemingly only produced a few clips (some of which were edited).

If it were one of many, I think you would name something better.

Re: OpenAI, Google and Anthropic are struggling to build more advanced AI

#385
post #146

Earlier quoted context omitted.

Being aware of its own limitations, for example. Or being aware of how its utterances may come across to its interlocutor. (And by limitations I don’t mean “sorry, I’m not allowed to help you with this dangerous/contentious topic”.)

There is no way of proving awareness in humans let alone machines. We do not even know whether awareness exists or it is just a word that people made up to describe some kind of feeling.

Awareness is exhibited in behavior. It's exactly due to the behavior be observe from LLMs that we don't ascribe them awareness. I agree that it's difficult to define, and it's also not binary, but it's behavior we'd like AI to have and which LLMs are quite lacking.

Re: OpenAI, Google and Anthropic are struggling to build more advanced AI

#386
post #115

Question for the group here: do we honestly feel like we've exhausted the options for delivering value on top of the current generation of LLMs? I lead a team exploring cutting edge LLM applications and end-user features. It's my intuition from experience that we have a LONG way to go. GPT-4o / Claude 3.5 are the go-to models for my team. Every combination of technical investment + LLMs yields a new list of potential…

I think there’s a long way to go also. I think people expected that AI would eventually be like a “point and shoot” where you would tell it to go do some complicated task, or sillier yet, take over someone’s entire job.

More realistically it’s like a really great sidekick for doing very specific mundane but otherwise non deterministic tasks.

I think we’ll start to see AI permeate into nearly every back office job out there, but as a series of tools that help the human work faster. Not as one big brain that replaces the human.

Re: OpenAI, Google and Anthropic are struggling to build more advanced AI

#387
post #248

Earlier quoted context omitted.

> Except our users experience the failed attempts (LLM replies that are wrong, even when backed by RAG) and it's incredibly hard to hide those from them. This has been my team's experience (and frustration) as well, and has led us to look at using LLMs for classifying / structuring, but not entrusting an LLM with making a decision based on things like a database schema or business logic. I think the technology and to…

> "let's only allow the LLM to do things we know it is rock-solid at." Even this is insanely hard in my opinion. The one thing that you would assume LLM to excel at is spelling and grammar checking for the English language, but even the top model (GPT-4o) can be insanely stupid/unpredictable at times. Take the following example from my tool: https://app.gitsense.com/?doc=6c9bada92&model=GPT-4o&samples... 5 models are…

> It keeps complaining that GitHub is spelled like Github, when it isn't

I feel like this is unfair. That's the only thing it got wrong? But we want it to pass all of our evals, even ones the perhaps a dictionary would be better at solving? Or even an LLM augmented with a dictionary.

Re: OpenAI, Google and Anthropic are struggling to build more advanced AI

#388

Earlier quoted context omitted.

> think or actually “understand” anything It doesn't matter if that's happening or not. That's the whole point of the Chinese room - if it can look like it's understanding, it's indistinguishable from actually understanding. This applies to humans too. I'd say most of our regular social communication is done in a habitual intuitive way without understanding what or why we're communicating. Especially the subtle infor…

There’s a distinction in behavior of a human and a Chinese room when things go wrong—when the rule book doesn’t cover the case at hand. I agree that a hypothetical perfectly-functioning Chinese room is, tautologically, impossible to distinguish from a real person who speaks Chinese, but that’s a thought experiment, not something that can actually exist. There’ll remain places where the “behavior” breaks down in ways…

> the “behavior” breaks down in ways that would be surprising from a human who’s actually paying as much attention as they’d need to be to have been interacting the way they had been until things went wrong.

That's an interesting angle. Though of course we're not surprised by human behavior because that's where our expectations of understanding come from. If we were used to dealing with perfectly-correctly-understanding super-intelligences, then normal humans would look like we don't understand much and our deliberate thinking might be no more accurate than the super-intelligence's absent-minded automatic responses. Thus we would conclude that humans are never really thinking or understanding anything.

I agree that default LLM output makes them look like they're thinking like a human more than they really are. I think mistakes are shocking more because our expectation of someone who talks confidently is that they're not constantly revealing themselves to be an obvious liar. But if you take away the social cues and just look at the factual claims they provide, they're not obviously not-understanding vs humans are-understanding.

Re: OpenAI, Google and Anthropic are struggling to build more advanced AI

#389

Go back a few decades and you'd see articles like this about CPU manufacturers struggling to improve processor speeds and questioning if Moore's Law was dead. Obviously those concerns were way overblown. That doesn't mean this article is irrelevant. It's good to know if LLM improvements are going to slow down a bit because the low hanging fruit has seemingly been picked. But in terms of the overall effect of AI and q…

> Go back a few decades and you'd see articles like this about CPU manufacturers struggling to improve processor speeds and questioning if Moore's Law was dead. Obviously those concerns were way overblown.

Am I missing something? I thought general consensus was that Moore's Law in fact did die:

https://cap.csail.mit.edu/death-moores-law-what-it-means-and...

The fact that we've still found ways to speed up computations doesn't obviate that.

We've mostly done that by parallelizing and applying different algorithms. IIUC that's precisely why graphics cards are so good for LLM training - they have highly-parallel architectures well-suited to the problem space.

All that seems to me like an argument that LLMs will hit a point of diminishing returns, and maybe the article gives some evidence we're starting to get there.

Re: OpenAI, Google and Anthropic are struggling to build more advanced AI

#390
post #146

Earlier quoted context omitted.

Being aware of its own limitations, for example. Or being aware of how its utterances may come across to its interlocutor. (And by limitations I don’t mean “sorry, I’m not allowed to help you with this dangerous/contentious topic”.)

Plenty of humans, unfortunately, are incapable of admitting limitations. Many years ago I had a coworker who believed he would never die. At first I thought he was joking, but he was in fact quite serious. Then there are those who are simply narcissistic, and cannot and will not admit fault regardless of the evidence presented them.

Being aware and not admitting are two different things, though. When you confront an LLM with a limitation, it will generally admit having it. That doesn't mean that it exhibits any awareness of having the limitation in contexts where the limitation is glaringly relevant, without first having confronted it with it. This is in itself a limitation of LLMs: In contexts where it should be highly obvious, they don't take their limitations into account without specific prompting.
Post reply on HN