Live data from Hacker News

OpenAI, Google and Anthropic are struggling to build more advanced AI

bloomberg.com

451–460 of 622 posts

Re: OpenAI, Google and Anthropic are struggling to build more advanced AI

#451
post #115

Question for the group here: do we honestly feel like we've exhausted the options for delivering value on top of the current generation of LLMs? I lead a team exploring cutting edge LLM applications and end-user features. It's my intuition from experience that we have a LONG way to go. GPT-4o / Claude 3.5 are the go-to models for my team. Every combination of technical investment + LLMs yields a new list of potential…

We might not have exhausted their applications, but everything I’ve witnessed them being used for has been extremely disappointing.

That is, other than me using them to bounce ideas off of and create small snippets of code.

Re: OpenAI, Google and Anthropic are struggling to build more advanced AI

#453
post #195

Earlier quoted context omitted.

> potential applications > if you ... > for example ... Yes there seems to be lots of potential. Yes we can brainstorm things that should work. Yes there is a lot of examples of incredible things in isolation. But it's a little bit like those youtube videos showing amazing basketball shots in 1 try, when in reality lots of failed attempts happened beforehand. Except our users experience the failed attempts (LLM repli…

We have built quite a few highly useful LLM applications in my org that have reduced cost and improved outcomes in several domains - fraud detection, credit analysis, customer support, and a variety of other spaces. By in large they operate as cognitive load reducers but also handle through automation the vast majority of work since in our uses false negatives are not as bad as false positives but the majority of thi…

>We have built quite a few highly useful LLM applications in my org that have reduced cost and improved outcomes in several domains

Apps that use LLMs or apps made with LLMs? In either case can you share them?

>which tells me you aren’t working with very creative and competent people

> In my professional network beyond where I work now I know at least a dozen people who have successful commercial applications of LLMs.

Apps that use LLMs or apps made with LLMs? In either case can you share them?

No one doubts that you can integrate LLMs into an application workflow and get some benefits in certain cases. That has been what the excitement and promise was about all along. They have a demonstrated ability to wrangle, extract, and transform data (mostly correctly) and generate patterns from data and prompts (hit and miss, usually with a lot of human involvement). All of which can be powerful. But outside of textual or visual chatbots or CRUD apps, no one wants to "put up or shut" a solid example that the top management of an existing company would sign off on. Only stories about awesome examples they and their friends are working on ... which often turn out to be CRUD apps or textual or visual chatbots. One notable standout is generative image apps can be quite good in certain circumstances.

So, since you seem to have a real interest and actual examples of this, I am curious to see some that real companies would gamble that company on. And I don't mean some quixotic startup, I mean a company making real money now with customers that is confident on that app to the point they are willing to risk big. Because that last part is what companies do with other (non LLM) apps. I also know that people aren't perfect and wouldn't expect an LLM to be, just want to make sure I am not missing something.

Re: OpenAI, Google and Anthropic are struggling to build more advanced AI

#454

Didn’t Sam Altman just go on some podcast last week and tell the world that he thought “We know exactly what to do to be able to reach AGI now”. What’s going on, is he just posturing?

"We know exactly what we need to do to be able to reach it: figure out how."

Re: OpenAI, Google and Anthropic are struggling to build more advanced AI

#455

Go back a few decades and you'd see articles like this about CPU manufacturers struggling to improve processor speeds and questioning if Moore's Law was dead. Obviously those concerns were way overblown. That doesn't mean this article is irrelevant. It's good to know if LLM improvements are going to slow down a bit because the low hanging fruit has seemingly been picked. But in terms of the overall effect of AI and q…

> Go back a few decades and you'd see articles like this about CPU manufacturers struggling to improve processor speeds and questioning if Moore's Law was dead. Obviously those concerns were way overblown. Am I missing something? I thought general consensus was that Moore's Law in fact did die: https://cap.csail.mit.edu/death-moores-law-what-it-means-and... The fact that we've still found ways to speed up computation…

I wrote "a few decades".

The article you pointed out says the end came in 2016: Eight years ago.

My point is those types of articles have been popping up every few years since the 1990s. Sure, at some point these sort of predictions will be proven correct about LLMs as well. Probably in a few decades.

Re: OpenAI, Google and Anthropic are struggling to build more advanced AI

#456
post #437

Sam Altman might be wrong then? Learning from data is not enough; there is a need for the kind of system-two thinking we humans develop as we grow. It is difficult to see how deep learning and backpropagation alone will help us model that. For tasks where providing enough data is sufficient to cover 95% of cases, deep learning will continue to be useful in the form of 'data-driven knowledge automation.' For other cas…

If Sam Altman concluded that AI is reaching it's limits, it probably wouldn't be a very good strategic decision for him to say it.

I know right ?

Re: OpenAI, Google and Anthropic are struggling to build more advanced AI

#457

Based on recent rumblings about AI scaling hitting a wall, of which this article is perhaps the most visible - and in a high-reach financial publication, I'm considering increasing my estimated probability we might see a major market correction next year (and possibly even a bubble collapse). (example: "CONFIRMED: LLMs have indeed reached a point of diminishing returns" https://garymarcus.substack.com/p/confirmed-llm…

NVIDIA has a strong interest in the financial plausibility of their big orders and the correlation between counterparty risks, and didn't scale up production during the crypto bubble because they understood the dynamics you are describing.

On the other hand, selling to customers who can't pay but who look solvent to public investors sounds like the kind of short-termism nobody should be too surprised to be reading a book about in a few years...

Re: OpenAI, Google and Anthropic are struggling to build more advanced AI

#458

A few important things to remember here: The best engineering minds have been focused on scaling transformer pre and post training for the last three years because they had good reason to believe it would work, and it has up until now. Progress has been measured against benchmarks which are / were largely solvable with scale. There is another emerging paradigm which is still small(er) scale but showing remarkable res…

Tesla are all making rapid progress on functionality which is definitely beyond frontier LLMs because it is distinctly different Sure, that's tautologically true but that doesn't imply that beyondness will lead to significant leaps that offer notable utility like LLMs. Deep Learning overall has been a way around the problem that intelligent behavior is very hard to code and no wants to hire many, many coders needed t…

Are we humans so different? Why do you wear what you wear? People emulate their older siblings, and so learn behavior. LLMs can create new programs, after having initially learned similar examples from others. Likewise for AI media.

Re: OpenAI, Google and Anthropic are struggling to build more advanced AI

#459
post #115

Question for the group here: do we honestly feel like we've exhausted the options for delivering value on top of the current generation of LLMs? I lead a team exploring cutting edge LLM applications and end-user features. It's my intuition from experience that we have a LONG way to go. GPT-4o / Claude 3.5 are the go-to models for my team. Every combination of technical investment + LLMs yields a new list of potential…

> combining a human-moderated knowledge graph with an LLM with RAG allows you to build "expert bots" that understand your business context / your codebase / your specific processes and act almost human-like similar to a coworker in your team

It's been a while though, we've had great models now for a 18 months plus. Why are we still yet to see these type of applications rolling out on a wide scale?

My anecdotal experience is that almost universally, 90-95% type accuracy you get from them is just not good enough. Which is to say, having something be wrong 10% or even 5% of the time is worse than not having at all. At best, you need to implement applications like that in an entirely new paradigm that is designed to extract value without bearing the costs of the risks.

It doesn't mean LLMs can't be useful, but they are kind of stuck with applications that inherently mesh with human oversight (like programming etc). And the thing about those is that they don't really scale, because the human oversight has to scale up with whatever the LLM is doing.

Re: OpenAI, Google and Anthropic are struggling to build more advanced AI

#460

Earlier quoted context omitted.

I disagree that it was over hyped. It has transformed our society so much that I would argue it was vastly under-hyped. Sure, there were a lot of silly companies that sprang up and went away because they weren't sound, but so much of the modern economy is based on the internet that it is hard to say any business isn't somehow internet related today. You would be hard pressed to find any business anywhere that doesn't…

pets.com was valued at $400 million based almost completely on its domain name. That's the classic example. People were throwing buckets of money at any .com that resolved to a site and almost all of it failed. I'm not sure how that doesn't meet the definition of over-hyped. It feels very similar to now. Not even to mention - the web largely doesn't consist of .com sites anymore, it's mostly a few centralized sites a…

Wasn't that mostly from public markets which never invested in tech before?
Post reply on HN