Live data from Hacker News

Meta AI Unleashes Megabyte, a Scalable Model Architecture

artisana.ai

41–50 of 213 posts

Re: Meta AI Unleashes Megabyte, a Scalable Model Architecture

#41
post #4

Earlier quoted context omitted.

I don't buy the idea (with either architecture) that "10x"-type scaling is required for another breakthrough. Think of a human with below average intelligence. Then think of a human genius. Now consider how incredibly similar their brains are, despite the massive performance gap. It's not like one has 10x the number of neurons/synapses/connections etc. of the other. They're both healthy human brains, and you need pow…

> Think of a human with below average intelligence. Then think of a human genius. LLMs are not AGI. A human with below average intelligence is still a league above a chimpanzee. A chimpanzee will never be able to read, not because "it's too dumb", but because a chimp's brain lacks the actual hardware for reading. The LLM is the chimpanzee. The gap between an LLM and a "human with below average intelligence" is far mo…

In your metaphor, maybe what the LLM (chimpanzee) "brain" is missing is some things like an inner thought loop, long term and short term storage and context, etc. Pieces we can understand, build, and surround the LLM with. Given the right arrangement, and the correct "abracadabra," perhaps that gets to where we are.

It seems to me that LLMs are perhaps capable of making up the inner thought loop, and the rest of the obvious systems seem doable. The context seems to be the hard problem; pulling in the proper things, and giving them the correct amount of attention.

Re: Meta AI Unleashes Megabyte, a Scalable Model Architecture

#42
post #10

Great paper. But wow, I really wish everyone would use more easily searchable names for their projects. In 6 months, there’s a high probability I’ll end up googling/ddging “megabyte model” trying to find this paper again.

The trick is to search HN submissions and filter by the date you remember reading about it. That's how I deal with these unsearchable names.

One step more effective is to keep track of things that keep your interest in a notes app.

Re: Meta AI Unleashes Megabyte, a Scalable Model Architecture

#43
post #37

Earlier quoted context omitted.

> How do you think a human with below average intelligence (or even with average intelligence) would fare? I don't understand what rote memorization to pass a test has to do with intelligence. For the record Kim Kardashian passed the bar; I imagine anyone given the proper motivation and time to study could do it, it's not a hard test. If AGI is a computer passing a test, then AGI was achieved a long time ago. I don't…

> If AGI is a computer passing a test, then AGI was achieved a long time ago. AIs couldn't even pass a third-grade reading comprehension exam until about 5 years ago. Computers being able to pass tests designed for humans is a very new thing. > it's clear LLMs are not AGIs And the main argument for that is that "it's clear". They're beating lawyers, doctors, and software engineers, but obviously, that's not real inte…

> Computers being able to pass tests designed for humans is a very new thing.

Yes! Captchas so effective!

Re: Meta AI Unleashes Megabyte, a Scalable Model Architecture

#45
post #13

Earlier quoted context omitted.

By what criteria? In several professional capacities, including coding, LLM's are far superior to an average human already.

No they're not. They're superior at specific tests when prompted. But asking for a complete application nearly always fails, when an average human that knows programming should be able to make one that works.

The LLM knows it doesn't work, too, we just don't run it long enough to let it try it out. It would try, get an error, and know how to handle the error.

https://github.com/drifting-in-space/botsh

Check out some agents that probably can already do more than you knew about.

Re: Meta AI Unleashes Megabyte, a Scalable Model Architecture

#46
post #15

Earlier quoted context omitted.

> The gap between an LLM and a "human with below average intelligence" is far more than 10x. In which direction? GPT-4 passes the bar exam with a top 10% score. How do you think a human with below average intelligence (or even with average intelligence) would fare? Copilot generates programming code that solves problems, and in most cases the code is correct. It outperforms many junior professional developers. Do you…

> How do you think a human with below average intelligence (or even with average intelligence) would fare? I don't understand what rote memorization to pass a test has to do with intelligence. For the record Kim Kardashian passed the bar; I imagine anyone given the proper motivation and time to study could do it, it's not a hard test. If AGI is a computer passing a test, then AGI was achieved a long time ago. I don't…

> it's not a matter of having a 100x more powerful LLM,

I think we all can agree that even the best LLM currently is not AGI. That's not what being disputed here I think.

However a 100x more powerful LLM is not just 100x better at recall. A 100x more powerful LLM is not just 100x better at being stupid hallucinatory parrot. A model that is just 100x bigger is not necessarily 100x more powerful if you define power is the ability to achieve goals.

However pure language models will always lack something else: the ability to ground things in reality.

I recently had a dream where I solved some problems and when I woke up I realized that those solutions were bullshit, but I also realized the whole approach of my dreaming self was very similar to what a LLM would have done.

Re: Meta AI Unleashes Megabyte, a Scalable Model Architecture

#47
post #37

Earlier quoted context omitted.

> How do you think a human with below average intelligence (or even with average intelligence) would fare? I don't understand what rote memorization to pass a test has to do with intelligence. For the record Kim Kardashian passed the bar; I imagine anyone given the proper motivation and time to study could do it, it's not a hard test. If AGI is a computer passing a test, then AGI was achieved a long time ago. I don't…

> If AGI is a computer passing a test, then AGI was achieved a long time ago. AIs couldn't even pass a third-grade reading comprehension exam until about 5 years ago. Computers being able to pass tests designed for humans is a very new thing. > it's clear LLMs are not AGIs And the main argument for that is that "it's clear". They're beating lawyers, doctors, and software engineers, but obviously, that's not real inte…

You're assuming that everything needed to write code or make arguments is purely intelligence based and has nothing to do with patterns, structural repetition and things glorified autocomplete could do, and that's not true.

Re: Meta AI Unleashes Megabyte, a Scalable Model Architecture

#48
post #22

Earlier quoted context omitted.

No they're not. They're superior at specific tests when prompted. But asking for a complete application nearly always fails, when an average human that knows programming should be able to make one that works.

> when an average human that knows programming should be able to make one that works. Very few people who "know programming" are actually capable of creating a complete application that serves a specific purpose. Many professional programmers struggle to implement basic algorithms like prime number testing. Coding AIs absolutely outperform the average programming professional, because the average programming professi…

Who is this average programmer you're talking about? We did primality testing in literally the first sem at uni

Re: Meta AI Unleashes Megabyte, a Scalable Model Architecture

#49
post #26

Earlier quoted context omitted.

Certainty about predictions of the future has always been close to zero. If you take a person from an appropriate time and ask them what the future will look like in 1000, 100 or even 20 years then their predictions will bear little resemblence to what actually occurs. As humans we have a tendancy to make linear future predictions based on past observations. Over the timescales that matter significant effects occur f…

I would say that we are really saying the same thing. I did intend to imply “short term predictions”. However I feel that these LLMs are not like the internet or the release of the iPhone. We’ve gone in a very short time from LLMs in the lab to ChatGPT to passing the bar exam. Considering the rate of research that’s being published and the fact that what we see today is generally at least many months behind the state…

Ah, I did not grasp that reading your comment - then we are saying the same thing.

Yes, I think that progress has become rapid enough that we can't predict six months out. I suspect that we are only 1-2 inventions away from something quite large and transformative. Obviously LLMs have already created a lot of excitement and opened up new areas of content generation already, but I think something larger is coming.

Re: Meta AI Unleashes Megabyte, a Scalable Model Architecture

#50

Earlier quoted context omitted.

But asking for a complete application nearly always fails, when an average human that knows programming should be able to make one that works. "This talking dog is a dumbass. The risotto recipe he gave me sucked, and the C++ code he wrote is full of security holes. I don't see what all the hype is about." Hint: ML will get better. Humans will not. ("Slow moving target," indeed.) What we're learning is just how many a…

This might be true but it's irrelevant to the claim "LLMs are already better than humans".

You are fixated on the current state of the art, when the first couple of time derivatives are what actually matter. The assertion in question is already true in a limited sense, and it's only going to go in one direction from here.
Post reply on HN