Live data from Hacker News

GPU-rich labs have won: What's left for the rest of us is distillation

inference.net

41–50 of 53 posts

Re: GPU-rich labs have won: What's left for the rest of us is distillation

#41
post #26

Earlier quoted context omitted.

Hell, we haven’t even seen actual AI yet. This is all just brute-forcing likely patterns of tokens based on a corpus of existing material, not anything brand new or particularly novel. Who would’ve guessed that giving CompSci and Mathematics researchers billions of dollars in funding and millions of GPUs in parallel without the usual constraints of government research would produce the most expensive brute-force algo…

Yeh. We're still barely beyond the first few pixels that make up the bottom tail of the S-curve for autonomous type AI everyone imagines Energy models and other substrates are going to be key, and it has nothing to do with text at all as human intelligence existed before language. It's Newspeak to run a chat bot on what is obviously a computer and call it an intelligence like a human. 1984 like dystopia crap.

Could you please stop creating accounts for every few comments you post? We ban accounts that do that. This is in the site guidelines: https://news.ycombinator.com/newsguidelines.html.

You needn't use your real name, of course, but for HN to be a community, users need some identity for other users to relate to. Otherwise we may as well have no usernames and no community, and that would be a different kind of forum. https://hn.algolia.com/?sort=byDate&dateRange=all&type=comme...

Re: GPU-rich labs have won: What's left for the rest of us is distillation

#42
post #32

Earlier quoted context omitted.

Hell, we haven’t even seen actual AI yet. This is all just brute-forcing likely patterns of tokens based on a corpus of existing material, not anything brand new or particularly novel. Who would’ve guessed that giving CompSci and Mathematics researchers billions of dollars in funding and millions of GPUs in parallel without the usual constraints of government research would produce the most expensive brute-force algo…

It's a necessary evolution step. Did you know our own ancestors had tails and grills. Do you feel ashamed?

Grills don't sound so bad!

Re: GPU-rich labs have won: What's left for the rest of us is distillation

#43
post #22

Earlier quoted context omitted.

What's your source for any of that? I think the $6 million thing was identified as a lie they felt was necessary because of GPU export laws.

It wasn't a lie, it was a misrepresentation of the total cost. It's not hard to calculate the cost of the training though. It takes 6 * active parameters * tokens flops[1]. To get number of seconds you can divide by Flops/s * MFU, where MFU is around 45% for H100 for large enough models[2]. [1]: https://arxiv.org/abs/2001.08361 [2]: https://github.com/facebookresearch/lingua

That paper's 5 years old at this point, dating back to when Amodei was still an OpenAI employee. Has any newer work superseded it, or are those assumptions still considered solid?

Re: GPU-rich labs have won: What's left for the rest of us is distillation

#45
post #42
post #32

Earlier quoted context omitted.

It's a necessary evolution step. Did you know our own ancestors had tails and grills. Do you feel ashamed?

Grills don't sound so bad!

The world is definitely a better place with grills.

Re: GPU-rich labs have won: What's left for the rest of us is distillation

#46

Earlier quoted context omitted.

It wasn't a lie, it was a misrepresentation of the total cost. It's not hard to calculate the cost of the training though. It takes 6 * active parameters * tokens flops[1]. To get number of seconds you can divide by Flops/s * MFU, where MFU is around 45% for H100 for large enough models[2]. [1]: https://arxiv.org/abs/2001.08361 [2]: https://github.com/facebookresearch/lingua

That paper's 5 years old at this point, dating back to when Amodei was still an OpenAI employee. Has any newer work superseded it, or are those assumptions still considered solid?

Those assumptions are still the same. Although now context length has increased more so the n^2 part is non negligible. See the repo for correct flop calculation[1]

[1]: https://github.com/facebookresearch/lingua/blob/437d680e5218...

Re: GPU-rich labs have won: What's left for the rest of us is distillation

#47
post #4

There is huge pressure to prove and scale radical alternative paradigms like memory-centric compute such as memristors, or SNNs, etc. That's why I am surprised we don't hear a lot about very large speculative investments in these directions to dramatically multiply AI compute efficiency. But one has to imagine that seeing so many huge datacenters go up and not being able to do training runs etc. is motivating a lot o…

> There is huge pressure to prove and scale radical alternative paradigms like memory-centric compute such as memristors, or SNNs, etc. That's why I am surprised we don't hear a lot about very large speculative investments in these directions to dramatically multiply AI compute efficiency.

Because the alternatives lack the breakthroughs that give them an edge against current-state AI and don't generate the hype like transformers or diffusion models. You have stuff like neuromorphic hardware that is hardly accessible and in its infancy, e.g. SpiNNaker. You have disciplines like Computational Neuroscience that try to model the brain and come up with novel models and algorithms for learning, which, however, are computational expensive or just perform worse than conventional deep learning models and may benefit from neuromorphic hardware. But again, access is difficult to such hardware.

Re: GPU-rich labs have won: What's left for the rest of us is distillation

#48

We haven't seen a proper npu and we are in the launch of the first consumer grade unified architectures by Nvidia and AMD. The battle of homebrew AI hasn't even started yet.

Hell, we haven’t even seen actual AI yet. This is all just brute-forcing likely patterns of tokens based on a corpus of existing material, not anything brand new or particularly novel. Who would’ve guessed that giving CompSci and Mathematics researchers billions of dollars in funding and millions of GPUs in parallel without the usual constraints of government research would produce the most expensive brute-force algo…

Even if we somehow create AGI it will only be used to saddle the masses, not free them. There are a lot of people with money who yearn for absolute power over everything. Replacing your capacity to think with a subscription service makes you a serf. Like Spotify, Netflix, Amazon and so on the last 20 years is full of brand new subscriptions that replace the whole concept of ownership so thoroughly that it gets erased. It pumps up the company’s valuation to do that.

Re: GPU-rich labs have won: What's left for the rest of us is distillation

#49
post #24

Earlier quoted context omitted.

Ho boy, should we start listing the 10x number of things that went in the wastebasket too?

If I only have to try 11 things for one of them to be LED lights or electric cars, I'd better get trying. Sure, I might have to empty a wastebasket at some point, but I'll just pay someone for that.

This fundamentally at odds with picking one tech and saying ‘this is the winner’ eh? Which is what the prior comment was about.

Re: GPU-rich labs have won: What's left for the rest of us is distillation

#50
post #49

Earlier quoted context omitted.

If I only have to try 11 things for one of them to be LED lights or electric cars, I'd better get trying. Sure, I might have to empty a wastebasket at some point, but I'll just pay someone for that.

This fundamentally at odds with picking one tech and saying ‘this is the winner’ eh? Which is what the prior comment was about.

Which prior comment? Top level offers two options plus an 'etc.', second comment adds another and says 'could be', and then first reply to you offers other tech that took decades of R&D to suggest we can't rule out memristors. I don't see what you're referring to.
Post reply on HN