Every negative headline I see about AI hitting a wall or being over-hyped makes me think of the early 2000's with that new thing the 'internet' (yes, I know the internet is a lot older than that). There is little doubt in my mind that ten years from now nearly every aspect of life will be deeply connected to AI just like the internet took over everything in the late 90's and early 2000's and is now deeply connected t…
OpenAI, Google and Anthropic are struggling to build more advanced AI
541–550 of 622 posts
Re: OpenAI, Google and Anthropic are struggling to build more advanced AI
#542Question for the group here: do we honestly feel like we've exhausted the options for delivering value on top of the current generation of LLMs? I lead a team exploring cutting edge LLM applications and end-user features. It's my intuition from experience that we have a LONG way to go. GPT-4o / Claude 3.5 are the go-to models for my team. Every combination of technical investment + LLMs yields a new list of potential…
I have tried a few AI coding tools and always found them impressive but I don't really need something to autocomplete obvious code cases. Is there an AI tool that can ingest a codebase and locate code based on abstract questions? Like: "I need to invalidate customers who haven't logged in for a month" and it can locate things like relevant DB tables, controllers, services, etc.
ChatGPT and Claude seem to be pretty good at maintaining an implicit understanding of the codebase based on a subset of files.
Re: OpenAI, Google and Anthropic are struggling to build more advanced AI
#543Earlier quoted context omitted.
> "let's only allow the LLM to do things we know it is rock-solid at." Even this is insanely hard in my opinion. The one thing that you would assume LLM to excel at is spelling and grammar checking for the English language, but even the top model (GPT-4o) can be insanely stupid/unpredictable at times. Take the following example from my tool: https://app.gitsense.com/?doc=6c9bada92&model=GPT-4o&samples... 5 models are…
> "I see LLM as a power tool for domain experts, and you have to assume whatever it spits out may be wrong, and your process should allow for it." this gets to the heart of it for me. I think LLMs are an incredible tool, providing advanced augmentation on our already developed search capabilities. What advanced user doesnt want to have a colleague they can talk about their specific domain capacity with? The problem c…
Re: OpenAI, Google and Anthropic are struggling to build more advanced AI
#544Earlier quoted context omitted.
> It keeps complaining that GitHub is spelled like Github, when it isn't I feel like this is unfair. That's the only thing it got wrong? But we want it to pass all of our evals, even ones the perhaps a dictionary would be better at solving? Or even an LLM augmented with a dictionary.
My reason for commenting wasn't to say LLM sucks, but rather we need to get over the honeymoon phase. The fact the GPT-4o (one of the most advanced, if not the most advanced when it comes to non programming tasks) hallucinated "Github" as the input, should give us pause. LLM has its place and it will forever change how we think about UX and other things, but we need to realize you really can't create a public facing…
LLM investors will be reviewing their portfolios and will likely begin declining further investments without clear evidence of profits in the very near future. On the other side, LLM companies will likely try to downplay this and again promise the Moon.
And on and on the market goes
Re: OpenAI, Google and Anthropic are struggling to build more advanced AI
#545Direct quote from the article: "The companies are facing several challenges. It’s become increasingly difficult to find new, untapped sources of high-quality, human-made training data that can be used to build more advanced AI systems." The irony here is astounding.
I don't know much about LLMs, but that seems to indicate a sort of dead-end. The models are still useful, but limited in their abilities. So now the developers and researchers needs to start looking for new ways to use all this data. That in some sense resets the game. Sucks to be OpenAI, billions of dollars spend on a product that has been match or even outmatched by the competition in a few short years, not nearly enough time to make any of it back.
If there is a take away, it might be that it takes billions, if not trillions of dollars, to develop an AI and the result may still be less than what you hope for, and the investment really hard to recoup.
Re: OpenAI, Google and Anthropic are struggling to build more advanced AI
#546The next wave won’t be monolithic but network-driven. Orchestration has the potential to integrate diverse AI systems and complementary technologies, such as advanced fact-checking and rule-based output frameworks. This methodological growth could make LLMs more reliable, consistent, and aligned with specific use cases. The skepticism surrounding this vision mirrors early doubts about the early internet fairly concis…
I don't think anyone doubted the nature of the technology. The bits were being sent. It's not like we were unsure of the fundamental possibility of transmitting information. The potential was shown very, very early on (Mother of all demos was in 1968). What we were and to some extent still are unsure of is the practical impact on society.
AI and LLMs in particular are not even at the mother of all demos level yet notwithstanding the grandiose claims and demos. There is no consensus on what these models are even doing. There is (IMO) justified skepticism surrounding the claims of reasoning and ability to abstract. We are in my opinion not yet at the "bits are being sent" stage.
Re: OpenAI, Google and Anthropic are struggling to build more advanced AI
#547Earlier quoted context omitted.
> potential applications > if you ... > for example ... Yes there seems to be lots of potential. Yes we can brainstorm things that should work. Yes there is a lot of examples of incredible things in isolation. But it's a little bit like those youtube videos showing amazing basketball shots in 1 try, when in reality lots of failed attempts happened beforehand. Except our users experience the failed attempts (LLM repli…
LLMs are not hype. In education at least, we've actively improved efficiency by ~25% across a large swath of educators (direct time saved) - agentic evaluators, tutors and doubt clarifiers. The wins in this industry are clear. And this is that much more time to spend with students. I also know from 1-1 conversation with my peers in large-finance world, and there too the efficiency improvements on multiple fronts are…
Re: OpenAI, Google and Anthropic are struggling to build more advanced AI
#548Earlier quoted context omitted.
The reason people are holding out is that the current generation of models are still pretty poor in many areas. You can have it craft an email, or to review your email, but I wouldn't trust an LLM with anything mission-critical. The accuracy of the generated output is too low be trusted in most practical applications.
Google (even now) wasn't absolutely accurate either. That didn't stop it from becoming many billions worth. > You can have it craft an email, or to review your email, but I wouldn't trust an LLM with anything mission-critical My point is that an entire world lies between these two extremes.
Re: OpenAI, Google and Anthropic are struggling to build more advanced AI
#549Earlier quoted context omitted.
I don't think we've even started to get the most value out of current gen LLMs. For starters very few people are even looking at sampling which is a major part of the model performance. The theory behind these models so aggressively lags the engineering that I suspect there are many major improvements to be found just by understanding a bit more about what these models are really doing and making re-designs based on…
My big question is what is being done about hallucination? Without a solution it's a giant footgun.
Look behind the veil and see LLMs for what they really are and you will maximise their utility, temper your expectations and save you disappointment
Re: OpenAI, Google and Anthropic are struggling to build more advanced AI
#550Earlier quoted context omitted.
I'm curious, what was the leap after GPT-4? What about the leaps after that, given a leap every 6 months?
O1, new Sonnet, all the music models and video models, the voice models like 4o, etc.
I'm surprised to hear someone say that O1 and new Sonnet are "leaps", though. My impression of them is that they're qualitatively similar to GPT-4. Incremental improvements at best. I don't think the gap between GPT-4 and the new Sonnet is anywhere near as large as the gap between GPT-3 and GPT-4, for instance.