My main argument against the AI doomsayers has so far been that the current scaling laws simply make runaway singularity style scenarios algorithmically impossible (if for each step of improvement you need 10x parameters and 100x training, you quickly run into a brick wall). This is part of why I’m not worried about the current crop of generative AI. I am however both curious and concerned about what the tsunami of t…
Honestly, I welcome the talent working on AI now. They were working on how to make me spend 5 more seconds on Facebook, or how to click on a Google ad. AI has potentially huge positive productivity potential.
Meta AI Unleashes Megabyte, a Scalable Model Architecture
131–140 of 213 posts
Re: Meta AI Unleashes Megabyte, a Scalable Model Architecture
#132Earlier quoted context omitted.
If that's the case why are we sending wheeled rovers to mars/moon and not something with legs? In any case this discussion is going sideways, the point is that nature doesn't have monopoly on being optimal. This also applies to intelligence/learning/modeling something better than brain.
> If that's the case why are we sending wheeled rovers to mars/moon and not something with legs? Because Mars is a simple and boring environment. Most of its surface, and especially the parts we target with rover missions, are effectively flat sheets peppered with rocks - a decent set of wheels and suspension is close to optimal for navigating such terrain. Now, if we were to send missions to a planet that's mostly f…
Re: Meta AI Unleashes Megabyte, a Scalable Model Architecture
#133Earlier quoted context omitted.
> Considering this, it seems perfectly possible that a model like GPT-4 is just a hair's breadth away from vastly superhuman performance. Except that structurally the brain is clearly has vastly more capacity than the GPT-4 model. So sure one brain doesn't look that much different to the other - and it's in the details of the learning, wiring. But the brain, looks vastly different from a GPT-4 model in terms of capac…
> brain is clearly has vastly more capacity I would not be so sure about "clearly" bit. Brain has 100B neurons x 1000 connections. GPT3 has 175B connections, but to implement operations used in those connections Nvidia H100 uses about 5M transistors, which in human brain would have to be copied to each connection because nature didn't invent software. (assumes one needs one full CUDA core to implement necessary ops,…
Also each neuron can hold much more state than a transistor.
For example they can respond to the timing of incoming events without having to build that capability with recurrent connections etc.
In addition the neural connections themselves have properties.
> because nature didn't invent software.
Eh? How do you explain the fact that brains can learn and aren't a fixed input, output engine?
Neural nets are software - just ones programmed by trial and error. Similarly for the brain.
Re: Meta AI Unleashes Megabyte, a Scalable Model Architecture
#134Earlier quoted context omitted.
> We did primality testing in literally the first sem at uni There's a lot of things you did during university classes. I bet you don't even remember half of them, and of the half you do, you couldn't actually do most of them from memory right now. The advantage LLMs have over programmers is that they've seen much more code than any human ever would, remember pretty much all of it - not necessarily verbatim, but also…
Yeah, no contest, llms can remember things. Big deal. You're comparing memory with intelligence. It's a part of it sure, but there's also cognition, reasoning and i don't even know what more
Yes, and LLMs show all of the specific things you've listed.
Re: Meta AI Unleashes Megabyte, a Scalable Model Architecture
#135Earlier quoted context omitted.
"Humanity will benefit enormously from this huge increase in the availability of intelligence." It's a near certainty that AI will be used to create more effective/destructive weapons (if it hasn't already), and will likely be used by terrorists, scammers, and others who wish to harm humans in some way. As this technology becomes more powerful, easier, and cheaper to use, all sorts of harmful uses of it will be made.…
Yes, great point So many of the people who opine about AI, its trajectory, and its possible effects on society, have latched on to one or two possible effects - like it overtaking jobs, or massively increasing misinformation. These are both very valid concerns, but they're only a tiny part of the big picture The thinker who I perceive as having the best holistic (in the non-wooey sense of the word) understanding of h…
I really hope these doomsayers are wrong, but my suspicion is the risk is real. Unfortunately, I'm not sure what can be done about it, as the profit and power these AI's promise is going to be near impossible for humanity to resist.
Re: Meta AI Unleashes Megabyte, a Scalable Model Architecture
#136Earlier quoted context omitted.
> The gap between an LLM and a "human with below average intelligence" is far more than 10x. In which direction? GPT-4 passes the bar exam with a top 10% score. How do you think a human with below average intelligence (or even with average intelligence) would fare? Copilot generates programming code that solves problems, and in most cases the code is correct. It outperforms many junior professional developers. Do you…
> How do you think a human with below average intelligence (or even with average intelligence) would fare? I don't understand what rote memorization to pass a test has to do with intelligence. For the record Kim Kardashian passed the bar; I imagine anyone given the proper motivation and time to study could do it, it's not a hard test. If AGI is a computer passing a test, then AGI was achieved a long time ago. I don't…
The reason scientific researchers in this area are using the term 'AGI' so much is that it does fit the definition of AGI ... *For some definition of AGI*. And there lies the problem - no one can really come to a good consensus on a good definition of AGI. This is why many scientists in this area are avoiding the question altogether - the question is loaded, and is misinterpreted by the public if statements are made.
So, for example, if I make the statement here that e.g. GPT-4 has intelligence which is general (AGI), it will likely be met with a rabid response from HN. However, the claim may be more dull than you're expecting. People often conflate AGI with things that are not required, such as agency, etc.
This definition from [journal Intelligence Vol 24, No1, 1997] can be that "some definition" of AGI: must be able to 1) think abstractly, 2) comprehend complex ideas, 3) reason, 4) plan, 5) solve problems, 6) learn quickly from experience.
Many of these GPT-4 can do, if some modifiers are allowed - for example, GPT-4 can learn quickly from experience so long as you aren't starting a 'new' GPT-4 system from scratch every time you want to interact with it. This is probably preferable, since much of the experiences it will have are personal to the individual working with it, and it would be highly undesirable to do the opposite here. Planning was difficult for the system early on, but it has appeared to learn that it is helpful to lay out a plan early on in a large task, so that appears to be a capability as well, at least on a basic level qualitatively. There is some good literature on ability to reason (and solve problems in abstract and complex ideas) from Microsoft's group on causal reasoning. In pretty much all tasks of causal reasoning the system can achieve near human performance, and LLMs as a category outperform previous SoTA from more targeted or specific systems made for causal reasoning.
Anyway, I would suggest to anyone that has strong reactions to claims about AGI to realize that they are likely building up the statement to be more than it is. Perhaps similar to 'machine learning' may have been misinterpreted years ago ("A machine can learn?! There is no tomorrow!!"), what is being stated here is often more narrow than you may believe.
Re: Meta AI Unleashes Megabyte, a Scalable Model Architecture
#137Earlier quoted context omitted.
Meta has no horse in the race (i.e. they don't have a search engine). So, they don't mind throwing random things out. Withholding it won't really make much of a difference for them, as they don't have a way to productionize the tech.
While I disagree that having a search engine is the only way to have a "horse in the race", I must agree that at this point Meta does not appear to have a horse in the race. Other companies are providing services that are so useful that it makes us think twice about how secure our jobs are. Then there is Meta, who seems to think that the world at large will forget about the terrible motion sickness that their VR prod…
Their horse seems to be "AI-generated ads". I'm still not sold on the idea though. I can see corporate Marketing departments using AI as a tool to ASSIST with ad copy development, but I'm skeptical that they'll let a third party like Meta generate the ad copy on the fly before pushing to targets. Maybe just tiny parts of it for "personalization".
Re: Meta AI Unleashes Megabyte, a Scalable Model Architecture
#138MEGABYTE is indeed a cool new architecture, but it is still very much just a proof of concept at the moment. The paper shows that the model can compete with (but not decimate) vanilla Transformers on the scale of ~1B parameters in long-sequence prediction tasks. They did not "unleash" anything, and scalability to very large parameter counts and datasets still has not been tested.
I'm certainly very excited to see where this architecture goes as the community gets ahold of it and starts developing it, but to call it "revolutionary" this early on is disingenuous. I personally have a few experiments I want to run with it, but I put the probability of it being a true GPT-killer at <20%. I would love to be wrong, though!
Re: Meta AI Unleashes Megabyte, a Scalable Model Architecture
#139First “Transformer” architecture and now “Megabyte”? I swear they’re trolling us! We’re just a generation or two away from the “computer” model.
Re: Meta AI Unleashes Megabyte, a Scalable Model Architecture
#140Earlier quoted context omitted.
Honestly, I welcome the talent working on AI now. They were working on how to make me spend 5 more seconds on Facebook, or how to click on a Google ad. AI has potentially huge positive productivity potential.
For all you know they might now be working on how to make you spend the rest of your life in a coal mine for AI overlords.