Live data from Hacker News

Will scaling work?

dwarkeshpatel.com

161–170 of 289 posts

Re: Will scaling work?

#161

Earlier quoted context omitted.

Demis Hassabis of Deepmind echoes a similar sentiment[0]: > I still think there are missing things with the current systems. […] I regard it a bit like the Industrial Revolution where there was all these amazing new ideas about energy and power and so on, but it was fueled by the fact that there were dead dinosaurs, and coal and oil just lying in the ground. Imagine how much harder the Industrial Revolution would hav…

His view of the Industrial Revolution is completely wrong. Societies pre-IR had multiple periods where energy usage increased significantly, some of them based specifically around coal. No IR. Early IR was largely based around the usage of water power, not coal. IR was pure innovation, people being able to imagine and create the impossible, it was going straight to nuclear already. Ironically, someone who is an innov…

I am very curious on what you mentioned, but not able to comprehend. Can you ELI5? Are you saying fossil fuel based industrial revolution is not as significant as it was or we could have directly jumped to a higher level fuel?

Re: Will scaling work?

#162
post #4

>Furthermore, the fact that LLMs seem to need such a stupendous amount of data to get such mediocre reasoning indicates that they simply are not generalizing. If these models can’t get anywhere close to human level performance with the data a human would see in 20,000 years, we should entertain the possibility that 2,000,000,000 years worth of data will be also be insufficient. There’s no amount of jet fuel you can a…

And yet, we reached the moon, and I would say airplanes were a necessary step on the way, even if only for psychological reasons. For airplanes we had at least an example in nature, birds. But I am not aware of any animal that travelled from earth to the moon on its own, except us.

> But I am not aware of any animal that travelled from earth to the moon on its own, except us.

Tardigrades might :)

Re: Will scaling work?

#163
post #83

Earlier quoted context omitted.

From the article, and relevant here: I’m worried that when people hear ‘5 OOMs off’, how they register it is, “Oh we have 5x less data than we need - we just need a couple of 2x improvements in data efficiency, and we’re golden”. After all, what’s a couple OOMs between friends? No, 5 OOMs off means we have 100,000x less data than we need.

I meant 100,000x. At least for everyone I know, they have 100,000x data in mail/messaging/docs/notes/meeting etc. than their blog or any public site they own. Hell I would even say that if you just have all the meetings of zoom, it will be few order of magnitude higher than the entire public web.

If I have 1MB on my blog, 100,000x would be 100GB. Just, no. OOMs are not to be trifled with.

Re: Will scaling work?

#164

Earlier quoted context omitted.

That is no different than saying beauty is only real if it is 100% natural. A woman who wears makeup and colors her hair is just an illusion of beauty. It is a philosophical argument to say that a machine isn't truly intelligent because it isn't using the same type of neural network as a human

Parent is saying that with something as sophisticated as intelligence it's not enough to say that if it behaves like a duck it's a duck (which is what your seem to be saying and which the parent calls a 0-day). There are some really good bulshitters who have led smart people into deep trouble. These bulshitters behaved really like ducks but they weren't ducks. The duck test just isn't good enough. The -1 day is where…

That is a new definition of intelligence that you are using. You are saying that even when something can outperform humans in the SAT or other tests of intelligence, it isn't actually intelligent due to it not being a carbon based lifeform

Re: Will scaling work?

#165
post #159

Earlier quoted context omitted.

Lol, yes, in fact, I was reacting to the article. The point I was trying to make is that I think better LLMs won’t lead to AGI. The article focused on the mechanics and technology, but I feel that’s missing the point. The point being, AGI is not going to be a direct outcome of LLM development, regardless of the efficiency or volume of data.

I can interpret this in a couple different ways, and I want to make sure I am engaging with what you said, and not with what I thought you said. > I think better LLMs won’t lead to AGI. Does this mean you believe that the Transformer architecture won't be an eventual part of AGI? (possibly true, though I wouldn't bet on it) Does this mean that you see no path for GPT-4 to become an AGI if we just leave it alone sitti…

Yes, I suppose my assertion is that LLMs may be a step toward our understanding of what is required to create AGI. But, the technology (the algorithms) will not be part of the eventual solution.

Having said that, I do agree that LLMs will be transformative technology. As important perhaps as the transistor or the wheel.

I think LLMs will accelerate our ability as a species to solve problems even more than the calculator, computer or internet has.

I think the boost in human capability provided by LLMs will help us more rapidly discover the true nature of reasoning, intelligence and consciousness.

But, like the wheel, transistor, calculator, computer and internet; I feel strongly that LLMs will prove to be just another tool and not a foundational technology for AGI.

Re: Will scaling work?

#166
post #49

The best analogy for LLMs (up to and including AGI) is the internet + google search. Imagine explaining the internet/google to someone in 1950. That person might say "Oh my god, everything will change! Instantaneous, cheap communication! The world's information available at light speed! Science will accelerate, productivity will explode!" And yet, 70 years later, things have certainly changed, but we're living in the…

The internet did change things pretty dramatically. Productivity at information communication tasks just isn’t the entire economy. I think we are massively more productive. Some of the biggest new companies are ad companies (Google, Facebook), or spend a ton of their time designing devices that can’t be modified by their users (Apple, Microsoft). Even old fashioned companies like tractor and train companies have time…

I remember long ago reading an argument that information technology has not actually increased productivity. I really wish I could find a source for this now, but I just can't seem to find it anywhere on the internet. Here it is anyway:

The administration of the Tax Service uses 4% of the total tax revenue it generates. This percentage has stayed relatively fixed over time.

If IT really improved productivity, wouldn't you expect that that number would decrease, since Tax Administration is presumably an area that we should expect to see great gains from computerisation?

We should be able to do the same amount of work more efficiently with IT, thus decreasing the percentage. If instead the efficiency frees up time allowing more work to be done (because there are people dodging taxes and we need to discover that), then you should expect the amount of tax to increase relatively which should also cause the percentage to decrease.

Therefore IT has not increased productivity.

Either it doesn't do so directly, or it does do so directly, but all the efficiency gains are immediately consumed by more useless beurocracy.

Re: Will scaling work?

#167
post #98

I think there’s a huge assumption here that more LLM will lead to AGI. Nothing I’ve seen or learned about LLMs leads me to believe that LLMs are in fact a pathway to AGI. LLMs trained on more data with more efficient algorithms will make for more interesting tools built with LLMs, but I don’t see this technology as a foundation for AGI. LLMs don’t “reason” in any sense of the word that I understand and I think the ab…

> I think there’s a huge assumption here that more LLM will lead to AGI. I'm not sure you realize this, but that is literally what this article was written to explore! I feel like you just autocompleted what you believe about large language models in this thread, rather than engaging with the article. Engagement might look like "I hold the skeptic position because of X, Y, and Z, but I see that the other position has…

I'm not sure you realize this, but that is literally what this article was written to explore!

Yeah but it's "exploration" answers all the reasonable objections by just extrapolating vague "smartness" (EDITED [1]). "LLMs seem smart, more data will make 'em smarter..."

If apparent intelligence were the only measure of where things are going, we could be certain GPT-5 or whatever would reach AGI. But I don't many people think that's the case.

The various critics of LLMs like Gary Marcus make the point that while LLMs increase in ability each iteration, they continue to be weak in particular areas.

My favorite measure is "query intelligence" versus "task accomplishment intelligence". Current "AI" (deep learning/transformers/etc) systems are great at query intelligence but don't seem to scale in their "task accomplishment intelligence" at the same rate. (Notice "baby AGI", ChatGPT+self-talk, fail to produce actual task intelligence).

[1] Edited, original "seemed remarkably unenlightening. Lots of generalities, on-the-one-hand-on-the-other descriptions". Actually, reading more closely the article does raise good objections - but still doesn't answer them well imo.

Re: Will scaling work?

#168

Earlier quoted context omitted.

> The internet did change things pretty dramatically. For sure - I grew up in the mid-late 70s having to walk to the library to research stuff for homework, parents having to use the yellow-pages to find things, etc. Maybe smartphones are more of a game changer than desk-bound internet though - a global communication device in your pocket that'll give you driving directions, etc, etc. BUT ... does the world really FE…

I don’t know really, I was a kid in the 90’s. This is a bit far from the economic aspect, but the world currently seemed to be utterly suffused with a looming sense of dread, I think because we have, or know other people have, news notifications in their pockets telling us all about how bad things are. I don’t remember that feeling from the 90’s, but then, I was a kid. And of course before that there was the constant…

I was a teenager in the 90s in a house that read the Daily Mail every day, and that could deliver a similar sense of dread.

But at least the dread was about things that seemed vaguely tractable and somewhat local, rather than the dizzyingly complex, global and existential threats the news delivers these days.

And of course not everyone read newspapers as intentionally-alarming as the Mail. Whereas now many more people’s information supply is mediated by channels with that brief.

Feels to me like a double-whammy of the alarm-maximising sections of the internet developing at the same time as the climate crisis becomes more imminent, maybe?

Re: Will scaling work?

#169
post #49

The best analogy for LLMs (up to and including AGI) is the internet + google search. Imagine explaining the internet/google to someone in 1950. That person might say "Oh my god, everything will change! Instantaneous, cheap communication! The world's information available at light speed! Science will accelerate, productivity will explode!" And yet, 70 years later, things have certainly changed, but we're living in the…

Imagine explaining to someone from 1950 that we now all have a TV-set on our office desks, with 1000+ channels ...

I bet their reaction would be a facepalm.

Re: Will scaling work?

#170
I think the "self-play" path is where the scary-powerful AI solutions will emerge. This implies persistence of state and logic that lives external to the LLM. The language model is just one tool. AGI/ASI/whatever will be a system of tools, of which the LLM might be the least complicated one to worry about.

In my view, domain modeling, managing state, knowing when to transition between states, techniques for final decision making, consideration for the time domain, and prompt engineering are the real challenges.

Post reply on HN