Live data from Hacker News

Will scaling work?

dwarkeshpatel.com

141–150 of 289 posts

Re: Will scaling work?

#141
post #49

The best analogy for LLMs (up to and including AGI) is the internet + google search. Imagine explaining the internet/google to someone in 1950. That person might say "Oh my god, everything will change! Instantaneous, cheap communication! The world's information available at light speed! Science will accelerate, productivity will explode!" And yet, 70 years later, things have certainly changed, but we're living in the…

It's important to remember that the internet is still very very new. Like the generation of digital natives are barely in adulthood. Sure, it's existed in some form for about 40 years, but most of the world didn't have access for the longest time. I wouldn't be surprised if we see massive changes in the next 20 years from the people who grew up on the web (specifically people outside the United States and Europe, whe…

"Digital native" are the people who grew up with computers. Many kids born in 1980's and later grew up with computers in their earliest memories.

I'd call the current generation "Social media natives", because that is the biggest difference from the previous generation. 90s kids grew up with games and communication, but they were free from facebook, youtube and instagram.

Re: Will scaling work?

#142

Earlier quoted context omitted.

> We do have systems that reason. Prolog comes to mind. It's a niche tool, used in isolated cases by relatively few people. I think that the other candidates are similar: proof assistants, physics simulators, computational chemistry and biology workflows, CAD, etc. I think OP meant other definition of reason, because by your definition calculator can also reason. These are tools created by humans, that help them to r…

If an expert system is not reasoning, and a statistical apparatus like an LLM is not reasoning, then I think the only definition that remains is the rather antiquated one which defines reason as that capability which makes humans unique and separates us from animals. I don't think it's likely to be a helpful one in this case.

I think he wants "reasoning" to include coming up with rules and not just following rules. Humans can reason by trying to figure out rules for systems and then see if those rules work well, on large scale that is called the scientific method but all humans do that on a small scale, especially as kids.

For a system to be able to solve the same classes of problems human can solve it would need to be able to invent their own rules just like humans can.

Re: Will scaling work?

#143
post #123

>Here’s one of the many astounding finds in Microsoft Research’s Sparks of AGI paper. They found that GPT-4 could write the LaTex code to draw a unicorn. a lot of people have tried to replicate this, I have tried. It's very hard to get GPT-4 to draw a unicorn, also asking it to draw an upside down unicorn is even harder.

The model of GPT-4 those researchers had was not the same that’s available to the public. It’s assumed it was far more capable before alignment training (or whatever it’s called).

Re: Will scaling work?

#144

If the size of the internet is really a bottleneck it seems Google is in quite a strong position. Assuming they have effectively a log of the internet, rather than counting the current state of the internet as usable data we should be thinking about the list of diffs that make up the internet. Maybe this ends up like Millenium Management where a key differentiator is having access to deleted datasets.

I'd guess at most they have 5x more data, but it is probably nowhere near that, and the article says 100,000x more data is needed.

Re: Will scaling work?

#145

I am in the believer camp for simple reasons: 1) we haven’t even scratched the surface of government led investments into AI, 2) AI itself could probably discover better architectures than transformers (willing to bet heavily on this)

> AI itself could probably discover better architectures than transformers

The entire subject of the article is concerned with what it will take and how likely it is than an AI will ever will able to generate improvements like this.

Re: Will scaling work?

#146

Earlier quoted context omitted.

The cognitive 0-day is not the way that humans act like LLMs, it's the way humans anthropomorphize LLMs. It's the blind faith that LLMs do more than they do. The illusion of emergence is fact not fiction. The cognitive biases exposed by stochastic parrots are fact not fiction.

That is no different than saying beauty is only real if it is 100% natural. A woman who wears makeup and colors her hair is just an illusion of beauty. It is a philosophical argument to say that a machine isn't truly intelligent because it isn't using the same type of neural network as a human

Parent is saying that with something as sophisticated as intelligence it's not enough to say that if it behaves like a duck it's a duck (which is what your seem to be saying and which the parent calls a 0-day).

There are some really good bulshitters who have led smart people into deep trouble. These bulshitters behaved really like ducks but they weren't ducks. The duck test just isn't good enough.

The -1 day is where people say that because LLMs behave like humans then humans must be based on the same tech. I just wonder if these people have ever debugged a complex system only to discover that their initial model of how it worked was way off.

Re: Will scaling work?

#147
I was thinking last night about LLMs with respect to Wittgenstein after watching this interesting discussion of his philosophy by John Searle [1].

I think Wittgenstein's ideas are pertinent to the discussion of the relation of language to intelligence (or reasoning in general). I don't meant this in a technical sense (I recall Chomsky mentioning that almost no ideas from Wittgenstein actually have a place in modern linguistics) but from a metaphysical sense (Chomsky also noted that Wittgenstein was one of his formative influences).

The video I linked is a worthy introduction and not too long so I recommend it to anyone interested in how language might be the key to intelligence.

My personal take, when I see skeptics of LLMs approaching AGI, is that they implicitly reject a Wittgenstein view of metaphysics without actually engaging with it. There is an implicit Cartesian aspect to their world view, where there is either some mental aspect not yet captured by machines (a primitive soul) or some physical process missing (some kind of non-language system).

Whenever I read skeptical arguments against LLMs they are not credibly evidence based, nor are they credibly theoretical. They almost always come down to the assumption that language alone isn't sufficient. Wittgenstein was arguing long before LLMs were even a possibility that language wasn't just sufficient, it was inextricably linked to reason.

What excites me about scaling LLMs, is we may actually build evidence that supports (or refutes) his metaphysical ideas.

1. https://www.youtube.com/watch?v=v_hQpvQYhOI&ab_channel=Philo...

Re: Will scaling work?

#148

Earlier quoted context omitted.

The internet did change things pretty dramatically. Productivity at information communication tasks just isn’t the entire economy. I think we are massively more productive. Some of the biggest new companies are ad companies (Google, Facebook), or spend a ton of their time designing devices that can’t be modified by their users (Apple, Microsoft). Even old fashioned companies like tractor and train companies have time…

> The internet did change things pretty dramatically. For sure - I grew up in the mid-late 70s having to walk to the library to research stuff for homework, parents having to use the yellow-pages to find things, etc. Maybe smartphones are more of a game changer than desk-bound internet though - a global communication device in your pocket that'll give you driving directions, etc, etc. BUT ... does the world really FE…

> does the world really FEEL that different now, than pre-internet?

Yes. You said it yourself: you used to have to WALK somewhere to look things up. Added convenience isn't the only side affect; that walk wasn't instantaneous. During the intervening time, you were stimulated in other ways on your trek. You saw, smelled, and heard things and people you wouldn't have otherwise. You may have tried different routes and learned more about your surroundings.

I imagine you, like I, grew up outside, sometimes with friends from a street or two over, that small distance itself requiring some exploration and learning. Running in fresh air, falling down and getting hurt, brushing it off because there was still more woods/quarry/whatever to see, sneaking, imagining what might lie behind the next hill/building; all of that mattered. The minutae people are immersed in today is vastly different in societies where constant internet access is available than it was before, and the people themselves are very different for it. My experience with current teens and very young adults indicates they're plenty bright and capable (30-somethings seem mostly like us older folks, IMO), but many lack the ability or desire to focus long enough to obtain real understanding of context and the details supporting it to really EXPERIENCE things meaningfully.

Admittedly anecdotal example: Explaining to someone why the blue-ish dot that forms in the center of the screen in the final scene of Breaking Bad is meaningful, after watching the series together, is very disheartening. Extrapolation and understanding through collation of subtle details seems to be losing ground to black and white binaries easily digested in minutes without further inquiry as to historical context for those options.

I abhor broad generalizations, and parenting plays a large part in this, but I see a concerning detachment among whatever we're calling post-millenials, and that's a major, real world difference coming after consecutive generations of increasing engagement and activism confronting the real problems we face.

Re: Will scaling work?

#149

Earlier quoted context omitted.

If an expert system is not reasoning, and a statistical apparatus like an LLM is not reasoning, then I think the only definition that remains is the rather antiquated one which defines reason as that capability which makes humans unique and separates us from animals. I don't think it's likely to be a helpful one in this case.

I think he wants "reasoning" to include coming up with rules and not just following rules. Humans can reason by trying to figure out rules for systems and then see if those rules work well, on large scale that is called the scientific method but all humans do that on a small scale, especially as kids. For a system to be able to solve the same classes of problems human can solve it would need to be able to invent thei…

I think that is what I mean by reason. I set the bar for reasoning and AGI pretty high.

Though, I will admit, a system that acts in a way that’s indistinguishable from a human will be awful hard to classify as anything but AGI.

Maybe I’m conflating AGI and consciousness, though given that we don’t understand consciousness and there’s no clear definition of AGI, maybe they ought to be inclusive of each other until we can figure out how to differentiate them.

Still, one interesting outcome, I think, should consciousness be included in the definition of AGI, is that LLMs are deterministic, which, if conscious, would (maybe) eliminate the notion of free will.

I feel like this whole exercise may end up representing a tiny, microscopic scratch on the surface of what it will actually take to build AGI. It feels like we’re extrapolating the capabilities of LLMs far too easily from capable chat bots to full on artificial beings.

We humans are great at imagining the future, but not so good at estimating how long it will take to get there.

Re: Will scaling work?

#150
post #4

>Furthermore, the fact that LLMs seem to need such a stupendous amount of data to get such mediocre reasoning indicates that they simply are not generalizing. If these models can’t get anywhere close to human level performance with the data a human would see in 20,000 years, we should entertain the possibility that 2,000,000,000 years worth of data will be also be insufficient. There’s no amount of jet fuel you can a…

Even the largest LLM has had less "total information" than most humans take in through all of their senses over their lifetime. A single day for a baby is taking in a continuous stream of among other things high quality video and audio and does a large amount of processing on that. Much of that for very young babies is unsupervised learning (clustering), where baby learns that object A and object B are different desp…

You are obviously a believer so you should know I know how to build AGI with a patented and trademarked architecture called "panoptic computronium cathedral"™. Tell all your friends about it. I only need $80B to achieve AGI.
Post reply on HN