Live data from Hacker News

Meta AI Unleashes Megabyte, a Scalable Model Architecture

artisana.ai

161–170 of 213 posts

Re: Meta AI Unleashes Megabyte, a Scalable Model Architecture

#161

Earlier quoted context omitted.

But asking for a complete application nearly always fails, when an average human that knows programming should be able to make one that works. "This talking dog is a dumbass. The risotto recipe he gave me sucked, and the C++ code he wrote is full of security holes. I don't see what all the hype is about." Hint: ML will get better. Humans will not. ("Slow moving target," indeed.) What we're learning is just how many a…

This might be true but it's irrelevant to the claim "LLMs are already better than humans".

Google Translate, which IIRC is also a Transformer model like GPT these days, knows more natural languages than I can remember the names of, to a higher standard than I know my best non-native language, and I moved to Germany 5 years ago.

GPT-3 knows most programming languages better than I do, even though I literally learned to read with the Commodore 64 user manual back in the 80s and haven't stopped being a nerd since; and while GPT code isn't always correct or even compilable, it's not like I don't still make mistakes that cause compilation to fail a few times a day, and there's a reason we all insist on testing code rather than just assuming it will work when a dev stops typing.

Again, I acknowledge GPT-3.5 (and I assume 4 but have not used it) is not perfect: like others I'd call 3.5 a "junior developer" (and it isn't even that good in every subject!); yet, despite that, the only domains where I can regularly beat it at those where my perception of reality is fundamentally different from its perception (how words sound and how numbers are composed, so it's relatively bad at arithmetic and rhyme), or where it has been forced to be bad (ask it to write about conflict, my experience is everyone reconciles and lives happily ever after).

But in most cases it would take me years to get as good as it already is. And there are more of those subjects than I can remember the names of, too.

Re: Meta AI Unleashes Megabyte, a Scalable Model Architecture

#162
post #41

Earlier quoted context omitted.

> Think of a human with below average intelligence. Then think of a human genius. LLMs are not AGI. A human with below average intelligence is still a league above a chimpanzee. A chimpanzee will never be able to read, not because "it's too dumb", but because a chimp's brain lacks the actual hardware for reading. The LLM is the chimpanzee. The gap between an LLM and a "human with below average intelligence" is far mo…

In your metaphor, maybe what the LLM (chimpanzee) "brain" is missing is some things like an inner thought loop, long term and short term storage and context, etc. Pieces we can understand, build, and surround the LLM with. Given the right arrangement, and the correct "abracadabra," perhaps that gets to where we are. It seems to me that LLMs are perhaps capable of making up the inner thought loop, and the rest of the…

>pulling in the proper things, and giving them the correct amount of attention.

This is likely one of the harder problems to solve. One of the strengths humans seem to have is the power of analogy. These analogies quite often can lead us into new paths of thought, or at least in the correct direction.

How do you do this in a machine system without wasting tons of power chasing dead ends, not sure?

Re: Meta AI Unleashes Megabyte, a Scalable Model Architecture

#163

Earlier quoted context omitted.

Yeah, no contest, llms can remember things. Big deal. You're comparing memory with intelligence. It's a part of it sure, but there's also cognition, reasoning and i don't even know what more

> You're comparing memory with intelligence. It's a part of it sure, but there's also cognition, reasoning and i don't even know what more Yes, and LLMs show all of the specific things you've listed.

Ah, here is our disagreement. I see what you mean but I'm not totally convinced by that.. reasoning isn't just being able to verbalize the steps you take, which llms can do but more so the steps themselves, and in my experience llms can fail in that department. Maybe they are somewhere between intelligent and not intelligent? In my opinion that's far more likely, as most things are not usually binaries but spectrums in our world.

Re: Meta AI Unleashes Megabyte, a Scalable Model Architecture

#164
post #128

Earlier quoted context omitted.

Yes, great point So many of the people who opine about AI, its trajectory, and its possible effects on society, have latched on to one or two possible effects - like it overtaking jobs, or massively increasing misinformation. These are both very valid concerns, but they're only a tiny part of the big picture The thinker who I perceive as having the best holistic (in the non-wooey sense of the word) understanding of h…

Also see this interview[1] with Robert Miles. I really hope these doomsayers are wrong, but my suspicion is the risk is real. Unfortunately, I'm not sure what can be done about it, as the profit and power these AI's promise is going to be near impossible for humanity to resist. [1] - https://m.youtube.com/watch?v=kMLKbhY0ji0

Yes, Robert Miles is great at explaining the problems of AI alignment, so I'll second the recommendation!

Re: Meta AI Unleashes Megabyte, a Scalable Model Architecture

#165

Earlier quoted context omitted.

No, because most tests designed for humans test memory and pattern recognition, which computers already can do better than humans so it's not a useful comparison. I'd rather define it to be superhuman AGI when it not only performs better on tests with humans who can use computers during the test but also can perform everyday tasks which are not 'hard' for us humans. That is because we have the hardware in our brains…

Have some examples? What would be tests that, if passed, you’d say “oh yeah, that’s AGI.” For instance, if it could make a peanut butter and jelly sandwich? Most challenging things that are easy for us are in the motor domain. While important, I think “intellectual AGI” is a meaningful milestone and closest to what most people think of when they think AGI.

The problem with defining AGI isn't only in defining intelligence, but also defining general. Also, why do we treat AGI as a yes/no question, when it probably makes sense to think partially... i don't have a definition of either

Re: Meta AI Unleashes Megabyte, a Scalable Model Architecture

#166
post #87

Earlier quoted context omitted.

This. So much this. I'm completely dumbfounded by obviously highly intelligent people consistently not getting this, and dismissing current generation AI systems as not being intelligent because they can't reliably solve massively complex problems in one go. Like anyone would expect a human programmer or researcher to just intuitively come up with a complex program, or the correct answer for a hard problem every time…

I would add to your amazing list that we are really good at denial as a coping mechanism with change. I am not a fan of the concept of AGI though. This means so many different things to people that it seems pointless to debate something when most likely we are not talking about the same thing. François Chollet has said that he believes all intelligence is specialized intelligence. From that perspective, whatever peop…

> Humanity will benefit enormously from this huge increase in the availability of intelligence.

I know corporations will, but Moloch doesn't necessarily represent humanity.

Re: Meta AI Unleashes Megabyte, a Scalable Model Architecture

#167

Earlier quoted context omitted.

You mean the average brain with about 100 billion neurons with about 1000 connections each bringing it to around 100 trillion connections. With an estimated 1000 "AI" neurons required per biologial neuron. I don't think you are givin these "below average" intelligence individuals enough credit. What we consider a genius is the equivalent of a dog show obstacle course. We measure intelligence/genius as whatever is har…

Nobody said that nature is optimal. Wheels are trivial, however not present in biology. Nature creates tentacles, not jet engines, nuclear energy etc. Majority of human brain computation is spent on things that are simply not necessary for computer models (how to wiggle limbs, mouth, eyes etc). Current LLM are impressive, but we know they can be much more efficient - we're using very low quality training data, we don…

Not relevant and not simply picking apart your example but I’ve been nerd sniped:

Wheels are not trivial in a biological sense. Topologically most life is either a tube or a cup depending on digestive systems. Wheels are separated from the body, which would be difficult for base cell division to produce.

Some animals like the pangolin are round shaped and do roll, but it just seems non optimal.

I’d say it’s curious nature hasn’t produced more creates that like making wheels like the dung beetle does, but nature made us and we do like making wheels.

Re: Meta AI Unleashes Megabyte, a Scalable Model Architecture

#168
post #138

Ugh. I just spent the past few days reading this exact paper and preparing a detailed presentation on it for my job as an AI researcher, and headlines like this make me roll my eyes very hard. MEGABYTE is indeed a cool new architecture, but it is still very much just a proof of concept at the moment. The paper shows that the model can compete with (but not decimate) vanilla Transformers on the scale of ~1B parameters…

When I started reading this and immediately came across the word "groundbreaking" I paused for a sec and thought to myself: Let me just pretend this word isn't there, it's probably an exaggeration. It takes a special kind of adaptation to filter out all the 10%-40% bluff usually found in these kinds of news articles.

Now I kind of want to see a browser extension that deletes all the intensifier adjectives from runs of text, and is able to be enabled/disabled per domain.

Re: Meta AI Unleashes Megabyte, a Scalable Model Architecture

#169
post #142
post #138

Ugh. I just spent the past few days reading this exact paper and preparing a detailed presentation on it for my job as an AI researcher, and headlines like this make me roll my eyes very hard. MEGABYTE is indeed a cool new architecture, but it is still very much just a proof of concept at the moment. The paper shows that the model can compete with (but not decimate) vanilla Transformers on the scale of ~1B parameters…

I wish you had a link to your blog or twitter in your bio, you have the kind of nuanced tone I’d like to hear more from

they should consider doing this for a living

Re: Meta AI Unleashes Megabyte, a Scalable Model Architecture

#170
post #142
post #138

Ugh. I just spent the past few days reading this exact paper and preparing a detailed presentation on it for my job as an AI researcher, and headlines like this make me roll my eyes very hard. MEGABYTE is indeed a cool new architecture, but it is still very much just a proof of concept at the moment. The paper shows that the model can compete with (but not decimate) vanilla Transformers on the scale of ~1B parameters…

I wish you had a link to your blog or twitter in your bio, you have the kind of nuanced tone I’d like to hear more from

Oh, thanks! Though to be honest, my blog does have a good bit of wild speculation on it, as I'm a bit of a hypocrite on that front :)

I've added links to both in my bio.

Post reply on HN