Great paper. But wow, I really wish everyone would use more easily searchable names for their projects. In 6 months, there’s a high probability I’ll end up googling/ddging “megabyte model” trying to find this paper again.
Meta AI Unleashes Megabyte, a Scalable Model Architecture
81–90 of 213 posts
Re: Meta AI Unleashes Megabyte, a Scalable Model Architecture
#82Earlier quoted context omitted.
Who said anything about everything? AI could not write code AT ALL until very recently. We could invent a drug that kills 99% of cancers and the next day there would be people bemoaning that it isn't a "true" cure.
Hey you're the one claiming intelligence and all, burden of proving it in squarely on you
Re: Meta AI Unleashes Megabyte, a Scalable Model Architecture
#83Earlier quoted context omitted.
> The gap between an LLM and a "human with below average intelligence" is far more than 10x. In which direction? GPT-4 passes the bar exam with a top 10% score. How do you think a human with below average intelligence (or even with average intelligence) would fare? Copilot generates programming code that solves problems, and in most cases the code is correct. It outperforms many junior professional developers. Do you…
> How do you think a human with below average intelligence (or even with average intelligence) would fare? I don't understand what rote memorization to pass a test has to do with intelligence. For the record Kim Kardashian passed the bar; I imagine anyone given the proper motivation and time to study could do it, it's not a hard test. If AGI is a computer passing a test, then AGI was achieved a long time ago. I don't…
I believe you do have a good point overall, but this here is, IMHO, a rather bad example, because I'd argue GPT-4 is already capable of abstract reasoning, and this seems to be exactly the skill that improves with increasing dimensionality of the latent space.
I agree that LLMs are equivalent of a single component of a human mind. Specifically, I think they're closest equivalent to our inner voice / inner thoughts. But this part is arguably exactly the one that does most of the abstract reasoning (and most reasoning in general), so I think in fact LLMs do have the "hardware" for that specific aspect. What's lacking right now is the equivalent of the higher, "conscious" layer, that guides, filters and censors the stream of thought. That, and long-term recall. Short-term memory might atually correspond to what the context window is in LLMs.
That, and fusing in all the other senses (sight, sound, smell, taste, touch, time, etc.).
Re: Meta AI Unleashes Megabyte, a Scalable Model Architecture
#84Re: Meta AI Unleashes Megabyte, a Scalable Model Architecture
#85In my understanding, this is accomplished through a hierarchical approach of “patches” of tokens.
Re: Meta AI Unleashes Megabyte, a Scalable Model Architecture
#86Earlier quoted context omitted.
Nobody said that nature is optimal. Wheels are trivial, however not present in biology. Nature creates tentacles, not jet engines, nuclear energy etc. Majority of human brain computation is spent on things that are simply not necessary for computer models (how to wiggle limbs, mouth, eyes etc). Current LLM are impressive, but we know they can be much more efficient - we're using very low quality training data, we don…
> Nobody said that nature is optimal. Wheels are trivial, however not present in biology. Nature creates tentacles, not jet engines, nuclear energy etc. Wheels are trivial but useless without bearings. Bearings most certainly aren't trivial. All the things you listed further are dependent on bearings somewhere.
Evolution is a greedy, lazy optimizer, so it promotes things that work a-ok for a given environment.
It's also worth noting that wheels alone are not too useful for transportation, as they're only half of the picture. The other half is roads. That is, because we couldn't (and mostly still can't) figure out all-terrain mobility systems that could navigate diverse environments, we cheated and locally flattened the environment, to reduce the problem to that solvable by a humble wheel. Evolution can't cheat like this.
Re: Meta AI Unleashes Megabyte, a Scalable Model Architecture
#87My main argument against the AI doomsayers has so far been that the current scaling laws simply make runaway singularity style scenarios algorithmically impossible (if for each step of improvement you need 10x parameters and 100x training, you quickly run into a brick wall). This is part of why I’m not worried about the current crop of generative AI. I am however both curious and concerned about what the tsunami of t…
Your premise is wrong. This has nothing to do with 10x-ing parameters. One could argue the current parameter sizes are good enough as we observe "large breadth, shallow depth" behavior from LLM and to some extent, diffusion models. This suggests the problem is the depth of inference, which is single pass "hot takes" for all language models right now, due to cost of inference and our limited understanding of what make…
I'm completely dumbfounded by obviously highly intelligent people consistently not getting this, and dismissing current generation AI systems as not being intelligent because they can't reliably solve massively complex problems in one go. Like anyone would expect a human programmer or researcher to just intuitively come up with a complex program, or the correct answer for a hard problem every time, instantly
Human thinking and problem solving involves a lot of trial and error, iterative thinking, and sharing and discussing the problem with other humans. Processes that AI researchers are just now beginning to explore, with results like increasing reasoning ability by 900% in a recent paper. Every thinking human runs a near constant loop of thought, with no conscious control of which thought will appear next (we're very good at fooling ourselves that we have control though)
We do have super-intelligences already, but they're severely handicapped by lacking a bunch of these - apparently fairly straightforward to implement - abilities, plus a few senses and the ability to directly effect change in the physical world (which really isn't needed if they can get access to human agents who will do their bidding, wittingly or unwittingly), and to self-improve. With regards to self-improvement, the increasing coding skills combined with iterative 'thought' loops should get there in very little time considering the current rate of progress
There's also the idea that a single AI model should be able to do everything our human brains do, when our brains actually contain a number of specialised subunits that handle different aspects of our behavioural repertoire. It reasonable to allow for the same thing with an AI system, where specialised sub-networks handle input, output and other subtasks. AI systems also have the advantage of being able to add any arbitrary number of subunits to increase its capacity to solve various problems
We seem to suffer from a species-wide narcissism with regards to our own intelligence and capabilities, and there's this huge focus on the number of connections in the human brain – most of which deal with things that are by no means necessary to act on the world unless one has a meat body and the need to navigate social situations, make friends and mate. Fact is, we have terrible short-term memory (worse than chimpanzees), slow processing time, lots of cognitive heuristics, many of which cause more harm than good in the modern world. We are emotional and easily fooled. Even the most intelligent people historically have believed in what we now consider fairy tales. We are slow to take in information, bad at storing it, and generally bad at transmitting it. A few of us can generate great ideas – building on accumulated knowledge from our forebears and peers – but most of us are just not that great at coming up with anything original or useful
I've been actively looking for good arguments against AGI being much closer than we should be comfortable with, and reasons why we should not fear systems that surpass us in intelligence. All I've come across so far is some combination of the above, often expressed with a dismissive attitude, disparaging current LLM:s as parrots (that can apparently reason on the level of university level humans, but much more quickly), and pejorative terms like fearmongerers and doomers to describe those of us who really don't think its a good idea to pursue more intelligent systems. My guess is these people will act surprised when the arms race inevitably leads to some very bad unintended consequences. I don't see a way to stop it though, so I'm just strapped in for the ride along with the rest of humankind
Again, if you have good arguments against any of the points above, please do share them with me
Re: Meta AI Unleashes Megabyte, a Scalable Model Architecture
#88Earlier quoted context omitted.
Define grounding things in reality. We only have our 5 senses to go off of. Meta has already put out one multimodal model incorporating multiple data types, openai is undoubtedly working on it too.
Grounding in reality can be something as simple as what openai is experimenting with plugins or something much more integrated. It's not a matter of which senses you have, but about being able to "continuously" use them. The current LLMs are basically unfiltered raw thoughts that must be continuously refined. A similar thing happens in our brains and only a little bit of that is accessible to our consciousness
Exactly. But, AFAIK, it's also the part that does the bulk of actual thinking and decision-making for us. In that sense, LLMs may be closer to AGI than people expect, because they seem to be capturing the actual core of intelligence and reasoning - and the missing bits (like long-term memory and higher-level thought stream filter/censor) may be much easier to bolt on to them.
Re: Meta AI Unleashes Megabyte, a Scalable Model Architecture
#89Earlier quoted context omitted.
> How do you think a human with below average intelligence (or even with average intelligence) would fare? I don't understand what rote memorization to pass a test has to do with intelligence. For the record Kim Kardashian passed the bar; I imagine anyone given the proper motivation and time to study could do it, it's not a hard test. If AGI is a computer passing a test, then AGI was achieved a long time ago. I don't…
> it's not a matter of having a 100x more powerful LLM, I think we all can agree that even the best LLM currently is not AGI. That's not what being disputed here I think. However a 100x more powerful LLM is not just 100x better at recall. A 100x more powerful LLM is not just 100x better at being stupid hallucinatory parrot. A model that is just 100x bigger is not necessarily 100x more powerful if you define power is…
Disagree, for the record. If I’d described the capabilities of contemporary AI to 100 AI scientists 5 years ago, I bet more than half would agree to call that AGI. Further, more than 90% would assume that these capabilities were decades and decades away.
Re: Meta AI Unleashes Megabyte, a Scalable Model Architecture
#90My LLM is totally real! And she’s the best, way better than GPT-4. But you wouldn’t know her, she goes to another school. In Canada.