Live data from Hacker News

Will scaling work?

dwarkeshpatel.com

251–260 of 289 posts

Re: Will scaling work?

#251

Earlier quoted context omitted.

The internet did change things pretty dramatically. Productivity at information communication tasks just isn’t the entire economy. I think we are massively more productive. Some of the biggest new companies are ad companies (Google, Facebook), or spend a ton of their time designing devices that can’t be modified by their users (Apple, Microsoft). Even old fashioned companies like tractor and train companies have time…

Going slightly beyond armchair economics, here are a couple of articles which discuss the lack of evidence for internet-based productivity so far: https://archive.ph/baneA https://archive.ph/TrHYN “Our central theme is that computers and the Internet do not measure up to the Great Inventions of the late nineteenth and early twentieth century, and in this do not merit the label of Industrial Revolution,” — Robert Gord…

“Less than the Industrial Revolution” leaves a pretty good amount of room.

Re: Will scaling work?

#252
post #83

> ‘5 OOMs off’ I think Google, Microsoft and facebook could easily have 5 OOM data than the entire public web combined if we just count text. Majority of people don't have any content on public web except for personal photos. A minority has few public social media posts and it is rare for people to write blog or research paper etc. And almost everyone has some content written in mail or docs or messaging.

From the article, and relevant here: I’m worried that when people hear ‘5 OOMs off’, how they register it is, “Oh we have 5x less data than we need - we just need a couple of 2x improvements in data efficiency, and we’re golden”. After all, what’s a couple OOMs between friends? No, 5 OOMs off means we have 100,000x less data than we need.

> "we just need a couple of 2x improvements in data efficiency, and we’re golden”"

100,000x is only seventeen 2x improvements. This[1] says world volume of data doubles every two years, citing a 2016 McKinsey study as the source. That puts it 34 years away. 2056. This[2] says the estimated compound growth rate of data creation is around 61%, that's a doubling time of 1.32 years and puts it 22 years away. 2045.

It's not tomorrow, but it may not be all that far away.

[1] https://rivery.io/blog/big-data-statistics-how-much-data-is-...

[2] https://theconversation.com/the-worlds-data-explained-how-mu...

Re: Will scaling work?

#253
post #49

The best analogy for LLMs (up to and including AGI) is the internet + google search. Imagine explaining the internet/google to someone in 1950. That person might say "Oh my god, everything will change! Instantaneous, cheap communication! The world's information available at light speed! Science will accelerate, productivity will explode!" And yet, 70 years later, things have certainly changed, but we're living in the…

By that standard, nothing has meaningfully changed since agriculture and domesticated animals. We're still killing each other, forming hierarchical societies, passing down stories, eating, drinking, sleeping, and making families - except now we're killing each other from afar with gunpowder, forming those hierarchies using the guise of democracy or whatever, passing down stories in print rather than speech, can use c…

Yeah in a real sense nothing has changed. I wonder if they finally will when we start modifying our bodies and minds to extreme degrees, that’d be my guess for when the model breaks down.

Re: Will scaling work?

#254

Earlier quoted context omitted.

> The internet did change things pretty dramatically. For sure - I grew up in the mid-late 70s having to walk to the library to research stuff for homework, parents having to use the yellow-pages to find things, etc. Maybe smartphones are more of a game changer than desk-bound internet though - a global communication device in your pocket that'll give you driving directions, etc, etc. BUT ... does the world really FE…

I once asked my mom, who grew up in the 1930s (aside: feels increasingly necessary to specific 19--), what was the biggest technological change she had seen in her lifetime. Her immediate answer was 'indoor plumbing.' But her next answer was the cellphone. She said cars and trains weren't vastly different from when she was a kid, she almost never went on a plane, and that people spent a lot of time watching the TV an…

My grandmother was born at home in 1917. Her father had to hitch up the wagon to go to town to fetch the doctor, and she had been born by the time they arrived. She felt it wasn't any particular innovation that was meaningful so much as the velocity of change. She lived to be nearly 100, so had gone from that horse-driven subsistence farm life to watching people land on the moon and the eventual digitalization of the world. She commented many times that she had a hard time believing that the same rate of change would occur in the next 100 years after she was gone. I've often wondered what that would look like - you'd almost need colonies on Mars to top what the last 100 years have been like in terms of changes. I suspect that the complete reengineering of our world away from fossil fuels may be that level of disruption and change.

Re: Will scaling work?

#255

I was thinking last night about LLMs with respect to Wittgenstein after watching this interesting discussion of his philosophy by John Searle [1]. I think Wittgenstein's ideas are pertinent to the discussion of the relation of language to intelligence (or reasoning in general). I don't meant this in a technical sense (I recall Chomsky mentioning that almost no ideas from Wittgenstein actually have a place in modern l…

> Wittgenstein view of metaphysics

lmao. I'd like to see you elaborate on what you think this means! The fact that you quote Searle, though, tells me all I need to know.

Re: Will scaling work?

#256

Here is an idea. Maybe the most optimized neural network is the brain. Computation to energy consumption ratio. So essentially the way doing this in silicon is just pointless. There must be a reason we can do so much while consuming so little, and then again struggling with other tasks. What is the success if we build a machine that consumes just heaps of energy and then is as bad in maths as us?

There's a couple of false dichotomies here - to say that because we're "more optimised" we must be the most optimised. Our brains are optimised well for certain things, sure, but computers are far more efficient at e.g. crunching numbers than we are - to say that there's no success in a machine that can't currently beat us at math - this year has already proven that false

That google paper really gave a bunch of idiots a whole bunch of ammunition. The four color theorem was proven by a machine long ago and it was worthless, about as worthless as what funsearch did!

Re: Will scaling work?

#257

Earlier quoted context omitted.

Wittgenstein is the perfect lens through which to be skeptical about LLMs and AGI and I think you may be the one not fully engaging with his work. He saw that languages are inseparable from the context in which they are used and that context is much bigger than language itself. Part of learning those language games is experimenting in the real world - interacting with other people, playing language games, and seeing…

> He saw that languages are inseparable from the context in which they are used That is one of the things that stood out to me in Searle's summary of his later work because I consider how the transformer architecture works and the way in which the surrounding context plays into the meaning of the words. > Part of learning those language games is experimenting in the real world It is interesting that the article we ar…

Now that I have a bit more time let me try a more substantive and less combative (and more drunk) reply :)

> That is one of the things that stood out to me in Searle's summary of his later work because I consider how the transformer architecture works and the way in which the surrounding context plays into the meaning of the words.

That's what makes transformer LLMs so interesting! Clearly they have captured a lot of what it means to be intelligent vis a vis language use, but is that enough to capture the kind of innate knowledge that defines intelligence at a human level? Based on my experiments with LLMs, it hasn't (yet). One of the clearest signs IMO is that there is no pedagological foundation to the LLM's answers. It can mimick explanations it learned on the internet but it cannot predict how to best explain a concept by implicitly picking up context from wrong answers or confusing questions. There is no "self reflection" because the algorithm as designed is incapable except for RLHF and finetuning.

> It is interesting that the article we are responding to talks about how we have only just begun to experiment with RL on top of transformers. In the same way that Alpha Go engaged in adversarial play we can envision LLMs being augmented to play language games amongst themselves. That may result in their own language, distinct from human language. But it also may result in the formation of intelligence surpassing human intelligence.

I think they already have their own language - embeddings! That really shines through with the multi-modal LLMs like GPT4V and LLaVa. What's curious is that we stumbled onto the same concept long before our algorithms showed any "intelligence" and it even helped Google move past the PageRank days. That's probably one of the fist steps towards intelligence but far from sufficient.

That brings up the fun question of what is enough to surpass human intelligence? I'm trying to apply LLMs to help make sense of the insane size of the American legal code and I can scale that process up to thousands of GPUs in a matter of seconds (as long as I can afford it). Even if it's at the level of a relatively dumb intern, that's a huge upside when talking about documents that would otherwise take years to read. Is that enough to claim intelligence, even if its not superior to a trained paralegal/lawyer?

>> The fidelity of text is simply too low to communicate the amount of information humans use to build up general intelligence.

> This does not at all follow from anything I've encountered in Wittgenstein. It is an empirical claim that we (as in humanity) are going to test and not something that I would argue either one of us can know simply reasoning from first principles.

Wittgenstein alone is not enough to come to this conclusion because it requires a peek at cognitive neuroscience and information theory which Mr W would have been woefully behind on given his time period. In short, just like the LLM "compresses" its training data to weights, all of human perception has to be compressed into language to communicate, which I think is impossible. We're talking about (age in years) * (365 days/year) * (X hours awake per day) * (500 megapixels per eye) * (2 eyes) + (all the other senses) versus however many bits it takes to represent language. I don't want to do the math on the latter cause I'm several beers in but it's not even close. 10 orders of magnitude wouldn't surprise me. 10 gigabytes of visual and other sensory input per 1 byte of language isn't out of the question.

I'm totally speculating and pulling numbers out of my ass here but the information theoretic part is undeniable: each human has access to more training data than it is possible for ChatGPT to experience. The quality of that training data ranges from "PEEKABOO!" to graduate textbooks, but its volume is incalculable and volume matters a lot to unsupervised algorithms like humans and LLMs.

> What does follow for me is closer to what Steven Pinker has been proposing in his own critiques of LLMs and AGI, which is that there is no necessary correlation between goal seeking (or morality) and intelligence. I also feel this is concordant with Wittgenstein's own work.

I haven't read Steven Pinkers critiques (could you link them please) so I can't say much about that. What does he mean by goal seeking?

IMO the only goal that matters is the will to survive, but lets assume for the sake of this discussion that it's not necessary to intelligence (otherwise we'll have to force our AGI bots to fight in a thunderdome and that's how we probably get Battlestar Gallactica all over again)

> My suspicion is that people don't want this scaling up to work because it would force them to let go of metaphysical commitments they have on both the nature of intelligence as well as the nature of reality. And for this reason they are adamantly disbelieving in even the possibility before the evidence has been gathered.

Let me be clear: I have zero metaphysical commitments and I can't wait until we come up with a richer vocabulary to describe intelligence. LLMs are clearly "intelligent", just not in any human sense quite yet. They don't have the will to survive, or any sense of agency, or even any permanence beyond the hard drive they exist on, but damn if they're not intelligent in some way. We just need better words to describe the levels of intelligence than "human, dog/cat/pig/pet, and everyone else"

However, I have some very strong physical commitments that must be met before I can even consider any algorithm as intelligent:

Neuroplasticity: human brains are incredibly capable of adapting all throughout life. That ranges from the simplest of drug tolerance to neurotransmitter attenuation/potentiation to the growth of new ion channels on cell membranes to very complex rewiring of axons and dendrites. That change is constant. It never stops, and LLMs don't have anything remotely like it. The brain rearchitects itself constantly and it's controlled as much by higher order processes as the neuron itself.

Scale: last time I did the math the minimum number of parameters required to represent the human connectome was over 500 quadrillion. 10+ quintillion is probably more accurate. That's 6-8 orders of magnitude more than we have in SOTA LLMs running on the best of the best hardware and Moore's law isn't going to take us that far. A 2.5D CPU/GPU might not even be theoretically capable of enough elements to simulate a fraction of that.

Quantization: I'm not sure neurons can be fully simulated with the limited precision of FP64, let alone FP32/16 or Q8/7/6/5/4. I've got far less evidence for this point than the others but it's a deeply held suspicion.

Re: Will scaling work?

#258

Earlier quoted context omitted.

The model of GPT-4 those researchers had was not the same that’s available to the public. It’s assumed it was far more capable before alignment training (or whatever it’s called).

That's convenient. Typically one of the markers of good science is reproducability. How can we trust any of the information coming out of these studies if it can't be reproduced?

Also - why not make this clear in all the model access documents? Perhaps call it GPT-4P (for Public?)

Perhaps also provide other researchers with vetted access. There are a lot of groups trying to evaluate these things systematically - for example "Faith and Fate", "Jumbled thoughts","Emergent abilities are a Mirage" were all very good papers published this year which really highlighted hype in LLM evaluation.

Everyone can see that modern LLM's have some great capabilities, the flexibility you can get in an interface by doing intent detection and categorizations using an LLM is great and it is so much easier and quicker than using previous techniques. It's more expensive, but that's improving rapidly. I firmly believe that a new era of great new systems with better interfaces and more functionality will be built on LLM's and other models from this wave of Big Data / Big Model AI, but these are not the precursors of AGI.

The problem with the looky looky AGI bunkum show is that it's pulling money into crappy projects that are going to fail hard and this will then stop a lot of money going into projects that could be successful fast. I am seeing the shape of the dotcom boom/bust in what's happening. Microsoft and Intel used dotcom to build and maintain their monopoly position, I think AWS, MS and Google will do the same this time. I think we will see a wave of new companies like Amazon that will "fail" some of them will really fail and disappear, some will half fail like Sun did, but some will go on to build monopolies anew. In the meantime the technology will evolve not for the greater good but instead to serve purposes like advertising distribution that are trivial compared to the benefit we could have seen. Over all we will not capitalise on the potential of what we have for several decades, ironically because of the failures of capitalism. Children will die, wars will be fought but some of us will have nice sweat pants and fun playing paddleball in the sunshine while it all happens.

When historians write this up in 100 years they won't really see any of this - they will just see a huge surge of innovation. The dead have no voices...

Re: Will scaling work?

#259
post #123

>Here’s one of the many astounding finds in Microsoft Research’s Sparks of AGI paper. They found that GPT-4 could write the LaTex code to draw a unicorn. a lot of people have tried to replicate this, I have tried. It's very hard to get GPT-4 to draw a unicorn, also asking it to draw an upside down unicorn is even harder.

A person commenting on this topic at a different site mentioned that there is a lot of content on the internet around how to draw (animals?) with LaTex and a different tool (can't remember the name), so it's unclear if GPT is just regurgitating or if it's generalizing.

I think it regurgitates badly, an interesting pre-print is the "reasoning or reciting" one from https://arxiv.org/abs/2307.02477

But - it's only a pre-print and I don't think that they have taken it forward to a publication so handle with care. Anyway, the relevant evidence and thinking that I would highlight and somewhat agrees with my findings is seen in figure 7 and section 5.6. I struggle to get the quality of results that they saw, but its believed by some people that ongoing development of GPT-4 and throttling of the reasoning for cost reasons may have limited some of its capabilities by the end of this year so that may be some of my problem.

Re: Will scaling work?

#260

Earlier quoted context omitted.

> The administration of the Tax Service uses 4% of the total tax revenue it generates. This percentage has stayed relatively fixed over time. The tax administration is far more efficient than that. The IRS has 79K workers out of a total workforce of 158M, or 1/2000 workers. Federal taxes are about 19% GDP (28% of GDP including state and local taxes.) The IRS costs $14.3B to run and collects 19% of $25.46T = $4,800B o…

The idea of measuring the “efficiency” of the tax system in terms of money in per money out seems a bit odd to me in the first place. Taxes don’t create wealth, the job is to destroy it at the correct rate. Don’t get me wrong, I’m not a “taxes are theft” dummy or anything like that. Taxes are an important knob in shaping the economy. But a better functioning tax collection agency should more effectively implement the…

It's still a useful measure to be aware of. Imagine if the IRS costed 30% of collected taxes to run instead of 0.4%. That would be a pretty telling sign that maybe we need to improve the IRS instead of raising taxing.
Post reply on HN