Live data from Hacker News

Meta AI Unleashes Megabyte, a Scalable Model Architecture

artisana.ai

151–160 of 213 posts

Re: Meta AI Unleashes Megabyte, a Scalable Model Architecture

#151

Earlier quoted context omitted.

> If that's the case why are we sending wheeled rovers to mars/moon and not something with legs? Because Mars is a simple and boring environment. Most of its surface, and especially the parts we target with rover missions, are effectively flat sheets peppered with rocks - a decent set of wheels and suspension is close to optimal for navigating such terrain. Now, if we were to send missions to a planet that's mostly f…

Exactly, same with intelligence - it's polluted with emotions and all kind of "nonsense" - but it doesn't have to. We can create emotionless, super-intelligent machines exceeding human capability by far (and use them as hammers). No need to imitate every detail of the brain to extract intelligence.

The counter to this is you can end up with an exceptionally powerful, but unaligned AI, which presents a new series of 'known unknowns' and 'unknown unknowns' that we have to deal with.

Re: Meta AI Unleashes Megabyte, a Scalable Model Architecture

#152
post #138

Ugh. I just spent the past few days reading this exact paper and preparing a detailed presentation on it for my job as an AI researcher, and headlines like this make me roll my eyes very hard. MEGABYTE is indeed a cool new architecture, but it is still very much just a proof of concept at the moment. The paper shows that the model can compete with (but not decimate) vanilla Transformers on the scale of ~1B parameters…

When I started reading this and immediately came across the word "groundbreaking" I paused for a sec and thought to myself:

Let me just pretend this word isn't there, it's probably an exaggeration.

It takes a special kind of adaptation to filter out all the 10%-40% bluff usually found in these kinds of news articles.

Re: Meta AI Unleashes Megabyte, a Scalable Model Architecture

#153
post #138

Ugh. I just spent the past few days reading this exact paper and preparing a detailed presentation on it for my job as an AI researcher, and headlines like this make me roll my eyes very hard. MEGABYTE is indeed a cool new architecture, but it is still very much just a proof of concept at the moment. The paper shows that the model can compete with (but not decimate) vanilla Transformers on the scale of ~1B parameters…

Every 'news' site: "But how are we gonna get clicks if we don't vastly exaggerate and sensationalize everything?"

Thanks for the actual sober analysis. As not-an-AI-researcher, I could have been easily bamboozled by an article like this.

Re: Meta AI Unleashes Megabyte, a Scalable Model Architecture

#155

Earlier quoted context omitted.

Who said anything about everything? AI could not write code AT ALL until very recently. We could invent a drug that kills 99% of cancers and the next day there would be people bemoaning that it isn't a "true" cure.

Hey you're the one claiming intelligence and all, burden of proving it in squarely on you

"intelligence"

This is a very problematic word. For example if you were a civil engineer and went "throw me any old design for a bridge, I have a river I need to cross", you'd have your license removed.

Intelligence is too massively loaded, and too much of a gradient even across humans to try to some up human or AI abilities. It is a multitude of different capabilities that don't necessarily have to be bundled together for something to be 'smart', 'useful', 'capable', and/or 'dangerous'.

Re: Meta AI Unleashes Megabyte, a Scalable Model Architecture

#156
post #15

Earlier quoted context omitted.

> Think of a human with below average intelligence. Then think of a human genius. LLMs are not AGI. A human with below average intelligence is still a league above a chimpanzee. A chimpanzee will never be able to read, not because "it's too dumb", but because a chimp's brain lacks the actual hardware for reading. The LLM is the chimpanzee. The gap between an LLM and a "human with below average intelligence" is far mo…

> The gap between an LLM and a "human with below average intelligence" is far more than 10x. In which direction? GPT-4 passes the bar exam with a top 10% score. How do you think a human with below average intelligence (or even with average intelligence) would fare? Copilot generates programming code that solves problems, and in most cases the code is correct. It outperforms many junior professional developers. Do you…

Sure but I can find examples of this in both directions. GPT-4 fails spectacularly at any task that doesn't have very precisely sanitized inputs presented to it in a certain format. It also fails at any task requiring any sort of interaction with the world. I think criticisms that it isn't a general intelligence are entirely fair even if it seems a bit more like general intelligence than most things that came before.

Re: Meta AI Unleashes Megabyte, a Scalable Model Architecture

#157

Earlier quoted context omitted.

Hey you're the one claiming intelligence and all, burden of proving it in squarely on you

Yes. Let’s define AGI as ability for a single model to pass most human professional tests (no cheating) and to provide genuine human-level flexible cognitive benefit to specialized professionals in diverse fields. Reasonable?

This definition fails badly because it doesn't test anything outside of language. At a bear minimum have the tests involved have pictures and descriptions in them and require the AI to use the same model to synthesize information from both.

Re: Meta AI Unleashes Megabyte, a Scalable Model Architecture

#158
post #15

Earlier quoted context omitted.

> Think of a human with below average intelligence. Then think of a human genius. LLMs are not AGI. A human with below average intelligence is still a league above a chimpanzee. A chimpanzee will never be able to read, not because "it's too dumb", but because a chimp's brain lacks the actual hardware for reading. The LLM is the chimpanzee. The gap between an LLM and a "human with below average intelligence" is far mo…

> The gap between an LLM and a "human with below average intelligence" is far more than 10x. In which direction? GPT-4 passes the bar exam with a top 10% score. How do you think a human with below average intelligence (or even with average intelligence) would fare? Copilot generates programming code that solves problems, and in most cases the code is correct. It outperforms many junior professional developers. Do you…

> GPT-4 passes the bar exam with a top 10% score

There was an article on HN yesterday debunking that. OpenAI is prone to exaggerating

Re: Meta AI Unleashes Megabyte, a Scalable Model Architecture

#159
post #54

Earlier quoted context omitted.

Grounding in reality can be something as simple as what openai is experimenting with plugins or something much more integrated. It's not a matter of which senses you have, but about being able to "continuously" use them. The current LLMs are basically unfiltered raw thoughts that must be continuously refined. A similar thing happens in our brains and only a little bit of that is accessible to our consciousness

> The current LLMs are basically unfiltered raw thoughts that must be continuously refined. A similar thing happens in our brains and only a little bit of that is accessible to our consciousness Exactly. But, AFAIK, it's also the part that does the bulk of actual thinking and decision-making for us. In that sense, LLMs may be closer to AGI than people expect, because they seem to be capturing the actual core of intel…

This is why we typically see better performance out of GPT when plugins are bolted in an chain|tree of though with reflection.

The output of LLMs is kind of like our stream of consciousness, there's a lot of things I think, then discount after internally reflecting on the thought which the often leads to a more correct solution. Having an LLM 'think' like this natively would massively increase the necessary the amount of compute needed, hence the expense, so at least in any public products it's not being done at this time.

Re: Meta AI Unleashes Megabyte, a Scalable Model Architecture

#160
post #91
post #3

My main argument against the AI doomsayers has so far been that the current scaling laws simply make runaway singularity style scenarios algorithmically impossible (if for each step of improvement you need 10x parameters and 100x training, you quickly run into a brick wall). This is part of why I’m not worried about the current crop of generative AI. I am however both curious and concerned about what the tsunami of t…

Honestly, I welcome the talent working on AI now. They were working on how to make me spend 5 more seconds on Facebook, or how to click on a Google ad. AI has potentially huge positive productivity potential.

Why do you think the best AI isn't going to be used to work on how to make you spend 5 more seconds on Facebook?
Post reply on HN