Live data from Hacker News

Will scaling work?

dwarkeshpatel.com

71–80 of 289 posts

Re: Will scaling work?

#71

I think there’s a huge assumption here that more LLM will lead to AGI. Nothing I’ve seen or learned about LLMs leads me to believe that LLMs are in fact a pathway to AGI. LLMs trained on more data with more efficient algorithms will make for more interesting tools built with LLMs, but I don’t see this technology as a foundation for AGI. LLMs don’t “reason” in any sense of the word that I understand and I think the ab…

We do have systems that reason. Prolog comes to mind. It's a niche tool, used in isolated cases by relatively few people. I think that the other candidates are similar: proof assistants, physics simulators, computational chemistry and biology workflows, CAD, etc.

When we get to the point where LLMs are able to invoke these tools for a user, even if that user has no knowledge of them, and are able to translate the results of that reasoning back into the user's context... That'll start to smell like AGI.

The other piece, I think, is going to be improved cataloging of human reasoning. If you can ask a question and get the answer that a specialist who died fifty years ago would've given you because that specialist was a heavy AI user and so their specialty was available for query... That'll also start to smell like AGI.

The foundations have been there for 30 years, LLMs are the paint job, the door handles, and the windows.

Re: Will scaling work?

#72

I think the more interesting question is how long will people cling to the illusion that LLMs will lead us to AGI? Maintaining the illusion is important to keep the money flowing in.

While this is certainly true, I think we can't ignore the intense enthusiasm and faith of a large cohort of our peers (or, you know, HN commenters) who believe this to be The Way, and are not necessarily stakeholders in any meaningful sense. Just look at some of the responses even in this thread. It feels like some people just need this, and respond to balanced skepticism as Alyosha does to his brother Ivan. In part,…

As much as I personally believe that neural networks do bear a lot of resemblance to the human psyche, and that people are just sophisticated biological machines, I don't see how LLMs are capturing all of our thought processes.

What I say is not just regurgitation of my past experiences; there is a logic to it.

Re: Will scaling work?

#73

Earlier quoted context omitted.

The internet did change things pretty dramatically. Productivity at information communication tasks just isn’t the entire economy. I think we are massively more productive. Some of the biggest new companies are ad companies (Google, Facebook), or spend a ton of their time designing devices that can’t be modified by their users (Apple, Microsoft). Even old fashioned companies like tractor and train companies have time…

I feel you are mixing value capture with value generation. If GM produces cars with the same level of margins as Facebook or Google, things will be different. LVMH (Louis Vuitton Group) holds a value equivalent to that of Toyota, Volkswagen, and two-thirds of Ford combined. Louis Vuitton alone was valued more than Red Hat a few months ago. This doesn't mean that Louis Vuitton is more valuable than Red Hat, but rather…

> This doesn't mean that Louis Vuitton is more valuable than Red Hat, but rather that it captures Value more effectively than Red Hat.

What definition of 'valuable' are you using here?

Re: Will scaling work?

#74
post #58

Earlier quoted context omitted.

I strongly disagree. Kids, even infants, show a remarkable degree of sophistication in relation to an LLM. I admit that humans don’t progress much behaviorally, outside of intellect, past our teen years; we’re very instinct driven. But still, I think even very young children have a spark that’s something far beyond rote token generation. I think it’s typical human hubris (and clever marketing) to believe that we can…

Humans are not very smart, individually, and over a single lifetime. We become smart as a species in tens of millennia of gathering experience and sharing it through language. What LLMs learn is exactly the diff between primitive humans and us. It's such a huge jump a human alone can't make it. If we were smarter we should have figured out the germ theory of disease sooner, as we were dying from infections. So don't…

Humans do more than just enhance predictive capabilities. It is also a very strong assumption that we are optimised for survival in many or all aspects (even unclear what that means). Some things could be totally incidental and not optimised. I find appeals to evolutionary optimisation very tricky and often fraught.

Re: Will scaling work?

#75

Earlier quoted context omitted.

The internet did change things pretty dramatically. Productivity at information communication tasks just isn’t the entire economy. I think we are massively more productive. Some of the biggest new companies are ad companies (Google, Facebook), or spend a ton of their time designing devices that can’t be modified by their users (Apple, Microsoft). Even old fashioned companies like tractor and train companies have time…

I feel you are mixing value capture with value generation. If GM produces cars with the same level of margins as Facebook or Google, things will be different. LVMH (Louis Vuitton Group) holds a value equivalent to that of Toyota, Volkswagen, and two-thirds of Ford combined. Louis Vuitton alone was valued more than Red Hat a few months ago. This doesn't mean that Louis Vuitton is more valuable than Red Hat, but rather…

I think I may have just skipped a step or not expressed myself very well.

What I’m saying is, I suspect information technology has made classic production companies vastly more efficient and productive. To the point where we can afford to have massive companies like Facebook that are almost entirely based on value capture.

That’s my speculation at least. Your example puts me in a tough spot, in the sense that Louis Vuitton is pretty old and pretty big. I’d have to know more about the company to quibble, and I don’t feel like researching it. I wonder if the proportion of their value that comes from pointless fashion branding was originally smaller. Or if the whole pointless fashion branding segment was originally just smaller itself. But I’m just spitballing.

In the past we also had mercenary companies and the like to capture value without producing much, so I could just be wrong.

Re: Will scaling work?

#76
post #69

Earlier quoted context omitted.

While this is certainly true, I think we can't ignore the intense enthusiasm and faith of a large cohort of our peers (or, you know, HN commenters) who believe this to be The Way, and are not necessarily stakeholders in any meaningful sense. Just look at some of the responses even in this thread. It feels like some people just need this, and respond to balanced skepticism as Alyosha does to his brother Ivan. In part,…

There is no magic in the brain. There is no magic in LLMs. There is just new experience we gain by interacting with the environment and society. And there is the trove of past experience encoded in our books. We got smart by collecting experience, in other words, from outside. The magic in the brain was not in the brain, but everywhere else. What is experience? We are in state S, and take action A, and observe feedba…

To say there's no magic in the brain drastically *minimizes the complexity of the brain.

Your brain is several orders of magnitude more complex than even the largest LLM.

GPT4 has 1 trillion parameters? Big deal. Your brain has 1 quadrillion synapses, constantly shifting. Beyond that the synapses are analog messages, not binary. Each synapse is approximately like 1000 transistors based on the granularity of messaging it can send and receive.

It is temporally complex as well as structurally complex, well beyond anything we've ever made.

I'm strongly in favor of AGI, for what it's worth, but LLMs aren't even scratching the surface. They're nowhere close to a human. They're a mediocre pastiche and it's equally possible that they're a dead end as it is that they'll ever be AGI.

Re: Will scaling work?

#77
post #49

The best analogy for LLMs (up to and including AGI) is the internet + google search. Imagine explaining the internet/google to someone in 1950. That person might say "Oh my god, everything will change! Instantaneous, cheap communication! The world's information available at light speed! Science will accelerate, productivity will explode!" And yet, 70 years later, things have certainly changed, but we're living in the…

Good point to me the internet was just "other people", what differentiated is not the 4 people you know but literally (almost) and potentially all other people.

With AI, the way I see it, it is just virtual other people. Of course, a bit stranger but more simillar than you think.

Re: Will scaling work?

#78
I'm not sure how one can percentage-wise compare scaling and algorithmic advances - per Dwarkesh's prediction that "70% scaling + 30% algorithmic advance" will get us to AGI ?!

I think a clearer answer is that scaling alone will certainly NOT get us to AGI. There are some things that are just architecturally missing from current LLMs, and no amount of scaling or data cleaning or emergence will make them magically appear.

Some obvious architectural features from top of my list would include:

1) Some sort of planning ahead (cf tree of thought rollouts) which could be implemented in a variety of ways. A simple single-pass feed forward architecture, even a sophisticated one like a transformer, isn't enough. In humans this might be accomplished by some combination of short term memory and the thalamo-cortical feedback loop - iterating on one's perception/reaction to something before "drawing conclusions" (i.e. making predictions) based on it.

2) Online/continual learning so that the model/AGI can learn from it's prediction mistakes via feedback from their consequences, even if that is initially limited to conversational feedback in a ChatGPT setting. To get closer to human-level AGI the model would really need some type of embodiment (either robotic or in a physical simulation virtual word) so that it's actions and feedback go beyond a world of words and let it learn via experimentation how the real world works and responds. You really don't understand the world unless you can touch/poke/feel it, see it, hear it, smell it etc. Reading about it in a book/training set isn't the same.

I think any AGI would also benefit from a real short term memory that can be updated and referred to continuously, although "recalculating" it on each token in a long context window does kind of work. In an LLM-based AGI this could just be an internal context, separate from the input context, but otherwise updated and addressed in the same way via attention.

It depends too on what one means by AGI - is this implicitly human-like (not just human-level) AGI ? If so then it seems there are a host of other missing features too. Can we really call something AGI if it's missing animal capabilities such as emotion and empathy (roughly = predicting other's emotions, based on having learnt how we would feel in similar circumstances)? You can have some type of intelligence without emotion, but that intelligence won't extend to fully understanding humans and animals, and therefore being able to interact with them in a way we'd consider intelligent and natural.

Really we're still a long way from this type of human-like intelligence. What we've got via pre-trained LLMs is more like IBM Watson on steroids - an expert system that would do well on Jeopardy and increasingly well on IQ or SAT tests, and can fool people into thinking it's smarter and more human-like than it really is, just as much simpler systems like Eliza could. The Turing test of "can it fool a human" (in a limited Q&A setting) really doesn't indicate any deeper capability than exactly that ability. It's no indication of intelligence.

Re: Will scaling work?

#79
post #69

Earlier quoted context omitted.

There is no magic in the brain. There is no magic in LLMs. There is just new experience we gain by interacting with the environment and society. And there is the trove of past experience encoded in our books. We got smart by collecting experience, in other words, from outside. The magic in the brain was not in the brain, but everywhere else. What is experience? We are in state S, and take action A, and observe feedba…

To say there's no magic in the brain drastically *minimizes the complexity of the brain. Your brain is several orders of magnitude more complex than even the largest LLM. GPT4 has 1 trillion parameters? Big deal. Your brain has 1 quadrillion synapses, constantly shifting. Beyond that the synapses are analog messages, not binary. Each synapse is approximately like 1000 transistors based on the granularity of messaging…

That kind of explains why humans need to absorb less language to train. It still takes 25 years of focused study to become capable of pushing the frontier of knowledge a tiny bit.
Post reply on HN