Live data from Hacker News

OpenAI, Google and Anthropic are struggling to build more advanced AI

bloomberg.com

541–550 of 622 posts

Re: OpenAI, Google and Anthropic are struggling to build more advanced AI

#541

Every negative headline I see about AI hitting a wall or being over-hyped makes me think of the early 2000's with that new thing the 'internet' (yes, I know the internet is a lot older than that). There is little doubt in my mind that ten years from now nearly every aspect of life will be deeply connected to AI just like the internet took over everything in the late 90's and early 2000's and is now deeply connected t…

Plus, they're "struggling"? Of course they are! It's cutting edge, and it's hard. If they weren't struggling, it would have been done long ago.

Re: OpenAI, Google and Anthropic are struggling to build more advanced AI

#542
post #115

Question for the group here: do we honestly feel like we've exhausted the options for delivering value on top of the current generation of LLMs? I lead a team exploring cutting edge LLM applications and end-user features. It's my intuition from experience that we have a LONG way to go. GPT-4o / Claude 3.5 are the go-to models for my team. Every combination of technical investment + LLMs yields a new list of potential…

I have tried a few AI coding tools and always found them impressive but I don't really need something to autocomplete obvious code cases. Is there an AI tool that can ingest a codebase and locate code based on abstract questions? Like: "I need to invalidate customers who haven't logged in for a month" and it can locate things like relevant DB tables, controllers, services, etc.

It's ugly, but I've had some success with uploading a few files from a project and a sketch of the schema. Then asking for new functionality.

ChatGPT and Claude seem to be pretty good at maintaining an implicit understanding of the codebase based on a subset of files.

Re: OpenAI, Google and Anthropic are struggling to build more advanced AI

#543
post #248

Earlier quoted context omitted.

> "let's only allow the LLM to do things we know it is rock-solid at." Even this is insanely hard in my opinion. The one thing that you would assume LLM to excel at is spelling and grammar checking for the English language, but even the top model (GPT-4o) can be insanely stupid/unpredictable at times. Take the following example from my tool: https://app.gitsense.com/?doc=6c9bada92&model=GPT-4o&samples... 5 models are…

> "I see LLM as a power tool for domain experts, and you have to assume whatever it spits out may be wrong, and your process should allow for it." this gets to the heart of it for me. I think LLMs are an incredible tool, providing advanced augmentation on our already developed search capabilities. What advanced user doesnt want to have a colleague they can talk about their specific domain capacity with? The problem c…

Welcome to capitalism. The market forces will squeze max value out of them. I imagine that Anthropic and OpenAI will be in the future fully downsized and acquired by their main investors (Microsoft and Amazon) and will simply becoming part of their generic and faceless AI & ML Teams once the current downwards stage of the hype cycle completes it closure in the next 5-8 years.

Re: OpenAI, Google and Anthropic are struggling to build more advanced AI

#544
post #443

Earlier quoted context omitted.

> It keeps complaining that GitHub is spelled like Github, when it isn't I feel like this is unfair. That's the only thing it got wrong? But we want it to pass all of our evals, even ones the perhaps a dictionary would be better at solving? Or even an LLM augmented with a dictionary.

My reason for commenting wasn't to say LLM sucks, but rather we need to get over the honeymoon phase. The fact the GPT-4o (one of the most advanced, if not the most advanced when it comes to non programming tasks) hallucinated "Github" as the input, should give us pause. LLM has its place and it will forever change how we think about UX and other things, but we need to realize you really can't create a public facing…

I believe the honeymoon face has loong been finished. Even in the mainstream, last year of the AI year. 2024 has seen nothing substantially good and the only notesworthy thing is this article finally hitting into the public consciousness that we are past of the AI peak and beyond the plateau and freefalling has already begun.

LLM investors will be reviewing their portfolios and will likely begin declining further investments without clear evidence of profits in the very near future. On the other side, LLM companies will likely try to downplay this and again promise the Moon.

And on and on the market goes

Re: OpenAI, Google and Anthropic are struggling to build more advanced AI

#545

Direct quote from the article: "The companies are facing several challenges. It’s become increasingly difficult to find new, untapped sources of high-quality, human-made training data that can be used to build more advanced AI systems." The irony here is astounding.

That's an interesting limitation. They can't make the LLMs (I still refuse to call them AIs) better, which the current dataset available. So with the sum of all human knowledge, more or less, and mixed in with the dumpster fire that it Internet comments, this is the best we can do with the current models.

I don't know much about LLMs, but that seems to indicate a sort of dead-end. The models are still useful, but limited in their abilities. So now the developers and researchers needs to start looking for new ways to use all this data. That in some sense resets the game. Sucks to be OpenAI, billions of dollars spend on a product that has been match or even outmatched by the competition in a few short years, not nearly enough time to make any of it back.

If there is a take away, it might be that it takes billions, if not trillions of dollars, to develop an AI and the result may still be less than what you hope for, and the investment really hard to recoup.

Re: OpenAI, Google and Anthropic are struggling to build more advanced AI

#546

The next wave won’t be monolithic but network-driven. Orchestration has the potential to integrate diverse AI systems and complementary technologies, such as advanced fact-checking and rule-based output frameworks. This methodological growth could make LLMs more reliable, consistent, and aligned with specific use cases. The skepticism surrounding this vision mirrors early doubts about the early internet fairly concis…

You are definitely on to something here, but the difference is that the fundamental process was proven. It "just" needed to scale. That's hard and complex, but on a different level.

I don't think anyone doubted the nature of the technology. The bits were being sent. It's not like we were unsure of the fundamental possibility of transmitting information. The potential was shown very, very early on (Mother of all demos was in 1968). What we were and to some extent still are unsure of is the practical impact on society.

AI and LLMs in particular are not even at the mother of all demos level yet notwithstanding the grandiose claims and demos. There is no consensus on what these models are even doing. There is (IMO) justified skepticism surrounding the claims of reasoning and ability to abstract. We are in my opinion not yet at the "bits are being sent" stage.

Re: OpenAI, Google and Anthropic are struggling to build more advanced AI

#547
post #195

Earlier quoted context omitted.

> potential applications > if you ... > for example ... Yes there seems to be lots of potential. Yes we can brainstorm things that should work. Yes there is a lot of examples of incredible things in isolation. But it's a little bit like those youtube videos showing amazing basketball shots in 1 try, when in reality lots of failed attempts happened beforehand. Except our users experience the failed attempts (LLM repli…

LLMs are not hype. In education at least, we've actively improved efficiency by ~25% across a large swath of educators (direct time saved) - agentic evaluators, tutors and doubt clarifiers. The wins in this industry are clear. And this is that much more time to spend with students. I also know from 1-1 conversation with my peers in large-finance world, and there too the efficiency improvements on multiple fronts are…

They are partially hype though. That's what people here are arguing. There are benefits but their valuation is largely hype driven. AI is going to transform industries and humanity, yes. But AI does not mean LLM (even if LLM means AI). LLM raw potential was reached last year with GPT-4. From here on, the value will lie on exploiting the potential we already have to generate clever applications. Just like the internet provided a platform for new services, I expect LLMs to be the same but with a much smaller impact

Re: OpenAI, Google and Anthropic are struggling to build more advanced AI

#548
post #482

Earlier quoted context omitted.

The reason people are holding out is that the current generation of models are still pretty poor in many areas. You can have it craft an email, or to review your email, but I wouldn't trust an LLM with anything mission-critical. The accuracy of the generated output is too low be trusted in most practical applications.

Google (even now) wasn't absolutely accurate either. That didn't stop it from becoming many billions worth. > You can have it craft an email, or to review your email, but I wouldn't trust an LLM with anything mission-critical My point is that an entire world lies between these two extremes.

Why don't you give actual concrete testable examples back with evidence where this is the case? Put your skin in the game.

Re: OpenAI, Google and Anthropic are struggling to build more advanced AI

#549

Earlier quoted context omitted.

I don't think we've even started to get the most value out of current gen LLMs. For starters very few people are even looking at sampling which is a major part of the model performance. The theory behind these models so aggressively lags the engineering that I suspect there are many major improvements to be found just by understanding a bit more about what these models are really doing and making re-designs based on…

My big question is what is being done about hallucination? Without a solution it's a giant footgun.

what do you want done about it? Hallucination is an intrinsic part of how LLMs work. What makes a hallucination is the inconsistency between the hallucinated concept and the reality. Reality is not part of how LLMs work. They do amazing things but at the end of the day they are elaborate statistical machines.

Look behind the veil and see LLMs for what they really are and you will maximise their utility, temper your expectations and save you disappointment

Re: OpenAI, Google and Anthropic are struggling to build more advanced AI

#550

Earlier quoted context omitted.

I'm curious, what was the leap after GPT-4? What about the leaps after that, given a leap every 6 months?

O1, new Sonnet, all the music models and video models, the voice models like 4o, etc.

The music/video models are cool, but It's an apples to orange comparison with GPT-4. I don't think there's really any comparison of intelligence or "advanceness" between those models and GPT-4.

I'm surprised to hear someone say that O1 and new Sonnet are "leaps", though. My impression of them is that they're qualitatively similar to GPT-4. Incremental improvements at best. I don't think the gap between GPT-4 and the new Sonnet is anywhere near as large as the gap between GPT-3 and GPT-4, for instance.

Post reply on HN