Live data from Hacker News

Will scaling work?

dwarkeshpatel.com

261–270 of 289 posts

Re: Will scaling work?

#261
post #69

Earlier quoted context omitted.

There is no magic in the brain. There is no magic in LLMs. There is just new experience we gain by interacting with the environment and society. And there is the trove of past experience encoded in our books. We got smart by collecting experience, in other words, from outside. The magic in the brain was not in the brain, but everywhere else. What is experience? We are in state S, and take action A, and observe feedba…

To say there's no magic in the brain drastically *minimizes the complexity of the brain. Your brain is several orders of magnitude more complex than even the largest LLM. GPT4 has 1 trillion parameters? Big deal. Your brain has 1 quadrillion synapses, constantly shifting. Beyond that the synapses are analog messages, not binary. Each synapse is approximately like 1000 transistors based on the granularity of messaging…

Re: synapses being analog messages, isn’t this sort of true of neural networks? In my understanding, the weights, biases and values flowing through the network are floating point numbers so I’d argue closer to analog than binary.

Re: Will scaling work?

#262
post #235

Earlier quoted context omitted.

Such a wild take. Would you want to live in the 50s? I definitely would not.

It doesn't need to be all or nothing. It's possible to acknowledge that the 50s had a better economic outlook for the middle class in developed countries than it does now, despite the incredible advances in computing. And you might still prefer to live now, for social reasons or because you prefer the computing advances, despite it not delivering economically as much. Either way, it does highlight a serious economic…

I’d arguing you cannot meaningfully separate them, they are on a continuum. As a thought experiment, you don’t know where you are going to land socioeconomically, which would you choose? Now or then?

Re: Will scaling work?

#263

Earlier quoted context omitted.

> The administration of the Tax Service uses 4% of the total tax revenue it generates. This percentage has stayed relatively fixed over time. The tax administration is far more efficient than that. The IRS has 79K workers out of a total workforce of 158M, or 1/2000 workers. Federal taxes are about 19% GDP (28% of GDP including state and local taxes.) The IRS costs $14.3B to run and collects 19% of $25.46T = $4,800B o…

The idea of measuring the “efficiency” of the tax system in terms of money in per money out seems a bit odd to me in the first place. Taxes don’t create wealth, the job is to destroy it at the correct rate. Don’t get me wrong, I’m not a “taxes are theft” dummy or anything like that. Taxes are an important knob in shaping the economy. But a better functioning tax collection agency should more effectively implement the…

>Taxes don’t create wealth, the job is to destroy it at the correct rate.

Not exactly, it's more about redistributing or reappropriating wealth because we know there exists flaws in the economic system we have (progressive tax systems). In more general terms, ignoring progressive structures, it's about investing in necessary shared services for everyone and to maintain the government that does such and more.

Looking at taxes as if they destroy wealth is a bit bleak. Governments may not be the most efficient institutions in all possible metrics but they're not out to destroy wealth exactly.

Re: Will scaling work?

#264

Earlier quoted context omitted.

Imagine telling those same people in the 50s that all those changes in productivity would come for the benefit of no one since the work week would be the same and purchasing power would decline

Don't have to imagine, same thing is happening right now with LLMs. I see "AI safety" brought up as a laughable attempt at stopping the progress of LLMs, when in reality the people talking about "AI safety" are the people trying to say that the majority will not benefit from this technology.

There's certainly pockets of the AI Safety movement that is about creating strategic motes to protect their private enterprises, but there are also legitimate concerns out there about these new search techniques can lead rise to concerns that may effect daily life.

I think most people realize there are some gains to be had here, the trick is to do so in a way that doesn't, yet again, massively redistribute wealth and power to a select few and instead share some of those gains. Right now there's not enough legitimate competition to keep things in check and that should be concerning. I'm all for rewarding the early successors who invested and took risk but let's not pretend those investments weren't captured through all sorts of other unequitable approaches to begin with.

Re: Will scaling work?

#265
post #49

The best analogy for LLMs (up to and including AGI) is the internet + google search. Imagine explaining the internet/google to someone in 1950. That person might say "Oh my god, everything will change! Instantaneous, cheap communication! The world's information available at light speed! Science will accelerate, productivity will explode!" And yet, 70 years later, things have certainly changed, but we're living in the…

If we ignore technology for the sake of technology and look at daily life, things we need like food, shelter, healthcare, transportation, socialization, etc. then I'd say technology has definitely improved some of these aspects.

Food distribution has improved as have most logistics in general. These efficiencies have somewhat been shared with the general public but in a lot of cases, those gains were captured by private enterprise.

Healthcare has improved a little bit, iterative progress can be made more quickly, shared, and moved into translational medicine as practice. Drug discovery has improved quite a bit, as have logistics around getting said drugs in the hands of people who need them and doing so affordably. This improved lives and longevity.

Socially we can communicate far easier. It remains to be seen to me if thise is always an improvement. Humans seem to be designed for much smaller social circles and don't seem to be capable of taking much advantage in their daily lives of increases frequency, scale, and reach of socialization.

The list goes on. It's not exactly linearly correlated with technology growth because ultimately it boils down to actionable information. Just because we have more information or more processing capability around information doesn't mean we get direct returns from that or that we don't reach limits where we simply don't have use for the additional gains. Information has to be actionable in some way, otherwise it's just intermediate data products that may or may not benefit us. I know can ready daily news from some small town in Southern Japan if I wanted to. That doesn't improve my life mostly, but it's there.

We have piles and piles of scientific literature we could share and iterate on towards new discoveries for humanity. That doesn't mean in my daily need for survival and balance with recreation I have time to contribute to things I find interesting or necessary, after all I am to some degree a slave of my needs within the economic system I'm entrenched in. I have bills, I have to earn money, and I have to work.

Even if that wasn't the case maybe or maybe not would I be able to contribute more back to society than I do now at my paid profession. Currently I'd say I do pretty well in this department in terms of reach. Without that I might struggle.

Re: Will scaling work?

#266

Earlier quoted context omitted.

The idea of measuring the “efficiency” of the tax system in terms of money in per money out seems a bit odd to me in the first place. Taxes don’t create wealth, the job is to destroy it at the correct rate. Don’t get me wrong, I’m not a “taxes are theft” dummy or anything like that. Taxes are an important knob in shaping the economy. But a better functioning tax collection agency should more effectively implement the…

>Taxes don’t create wealth, the job is to destroy it at the correct rate. Not exactly, it's more about redistributing or reappropriating wealth because we know there exists flaws in the economic system we have (progressive tax systems). In more general terms, ignoring progressive structures, it's about investing in necessary shared services for everyone and to maintain the government that does such and more. Looking…

I don’t think it needs to be seen as bleak; money just exists as a tool, it doesn’t have any other meaning, sometimes destroying it is the best thing to do.

It is fungible anyway, so I think it is really just a matter of semantics or philosophy if the government is collecting and redistributing dollars, or if it is destroying and creating them. My (outsider) understanding is that modern monetary theory leans toward the latter, because it more accurately reflects the latitude the government has, working with a fiat currency and all that.

Re: Will scaling work?

#267

Earlier quoted context omitted.

The idea of measuring the “efficiency” of the tax system in terms of money in per money out seems a bit odd to me in the first place. Taxes don’t create wealth, the job is to destroy it at the correct rate. Don’t get me wrong, I’m not a “taxes are theft” dummy or anything like that. Taxes are an important knob in shaping the economy. But a better functioning tax collection agency should more effectively implement the…

It's still a useful measure to be aware of. Imagine if the IRS costed 30% of collected taxes to run instead of 0.4%. That would be a pretty telling sign that maybe we need to improve the IRS instead of raising taxing.

Only, I think, in the sense that it would be bad if it was very expensive.

It just isn’t an efficiency, in the / sense. The measurement of / is meaningless in the same way that a comparison between the energy required to flip a light switch and the energy that is sent to that socket as a result is basically meaningless.

I mean, sure, if tending a light switch required an appreciable percentage of the amount of the energy the socket could provide, that would be bad, but it is a ridiculous scenario.

Re: Will scaling work?

#268

Earlier quoted context omitted.

There is nothing oversimplified in that description at all. Here is an equivalent description from a neuroscience textbook: "Neuroscience is the study of the nervous system, the collection of nerve cells that interpret all sorts of information which allows the body to coordinate activity in response to the environment." https://openbooks.lib.msu.edu/introneuroscience1/chapter/wha... This is even simpler than the RL a…

Yet again you just pick out a word, "oversimplified" in this case, ignores the rest of the reply - and more importantly, the entire context of all these replies - and goes on to make a nonsensical comparison with the ever-present one-liner that all textbooks of all academic disciplines have. Have you not understood by now that my critic of our field's AI-hubris is a general one - and thus not hinged on exact wordings…

You keep ignoring that your criticism is nonsensical. There is nothing in Visarga's comment that indicates hubris. If there is, you would be able to plainly point it out instead of complaining that people are misinterpreting your complaint.

Re: Will scaling work?

#269

Earlier quoted context omitted.

It's not necessary for the author's purpose of providing more data. We're only training on one kind of input so far, text, from which these models have built some understanding of the world. Humans train on more inputs, and the data to provide those inputs for training a model is readily available, in far larger quantities than individual human brains consume. Data is not the issue.

We're training on text because that's what we're making the model do. It's a fact of neural networks that to train them supervised you need the training data in the expected input for(vector of n thousand preceding tokens for LLMs) with the expected output(the next token for LLMs). "Training them on video" would mean converting the video to a format we can train the llm with, then training the LLM with that info. Thi…

> This would probably be a 1 OOM increase at maximum, if the video transcripts aren't already a part of the training data for gpt.

Human brains aren't trained on video transcripts, which leave out a lot of information from the video that human brains have a shared understanding of due to training. You would train on video embeddings and predict next embeddings, thus learning physics and other properties of the world and the objects within it and how they interact. This is many more than 5 orders of magnitude more data than the text that today's LLMs are trained on.

Re: Will scaling work?

#270

Earlier quoted context omitted.

> He saw that languages are inseparable from the context in which they are used That is one of the things that stood out to me in Searle's summary of his later work because I consider how the transformer architecture works and the way in which the surrounding context plays into the meaning of the words. > Part of learning those language games is experimenting in the real world It is interesting that the article we ar…

Now that I have a bit more time let me try a more substantive and less combative (and more drunk) reply :) > That is one of the things that stood out to me in Searle's summary of his later work because I consider how the transformer architecture works and the way in which the surrounding context plays into the meaning of the words. That's what makes transformer LLMs so interesting! Clearly they have captured a lot of…

> I haven't read Steven Pinkers critiques (could you link them please) so I can't say much about that. What does he mean by goal seeking?

He has spoken about it multiple times on a few different podcasts. Here is a recent discussion he had with physicist David Deutsch [1] where he references this idea (See chapter timestamp for "Does AGI need agency to be 'creative'?").

1. https://www.youtube.com/watch?v=3Ho-vJZsMgk&t=2363s&ab_chann...

Post reply on HN