Live data from Hacker News

Meta Unveils New AI Supercomputer

wsj.com

171–180 of 199 posts

Re: Meta Unveils New AI Supercomputer

#171
post #76

I can't shake the feelings that a trillion or a quadrillion parameters won't solve the fundamental shortcomings of ML models not being models of artificial intelligence. I guess there's no way of knowing until we reach AGI, but I've never heard a compelling argument for why pure ML would get us there. GPT3 seems more like an argument against that hypothesis (in my view) than for it. Even the best, most expensive mode…

I agree. I read Jeff Hawkins book On Intelligence [0] back when it came out, and it had a profound effect on my thinking. Chasing more data, aka "parameters" doesn't seem to be the right answer. I think more of a Bayes model like spam filtering, but cobbled together with other Bayes models looking at other things until something emerges that we call "intelligent". Heck, I'd consider Google's spam filtering pretty int…

Hawkins way of thinking really maps well for me also. It seems like that more parameters helps until it doesn't, then you need to encapsulate those networks and pin them to some reference frame, they create hierarchies of these networks and a system to generalize and compress those hierarchies (aka patterns), rinse and repeat.

My brother just became a grandpa and I was watching his grandson navigate the world this past weekend. It's unbelievable how quickly the brain can extrapolate a new relationship between objects/actions/etc and then apply it elsewhere. Minimally you see it in the drinking action applied to all sorts of things, this sort of repetitive clenching/releasing of the fingers to find things to grip without looking, etc etc. Watching mom use a fork and very quickly understand how to grasp and manipulate it. The model of just training everything from exogenous data into a flat network seems like it will hit some asymptotic limit.

Re: Meta Unveils New AI Supercomputer

#172
post #94

> Meta’s AI supercomputer houses 6,080 Nvidia graphics-processing units ..... By mid-summer, when the AI Research SuperCluster is fully built, it will house some 16,000 GPUs Honestly ... this is lot of GPUs ... but is it the biggest...? > Model training is done with mixed precision on the NVIDIA DGX SuperPOD-based Selene supercomputer powered by 560 DGX A100 servers networked with HDR InfiniBand in a full fat tree co…

Honestly, this single GPU-based install is child's play compared to Google's multiple TPU exoflop supercomputers with hyper-cube optical interconnects. Google's ML setups allow synchronous weight update on thousand+ TPUs...

For what its worth, for attention based advertising (youtube and display, not search), FB targeting blows Google out of the water. Not sure why but its consistent across brands.

Re: Meta Unveils New AI Supercomputer

#173

I can't shake the feelings that a trillion or a quadrillion parameters won't solve the fundamental shortcomings of ML models not being models of artificial intelligence. I guess there's no way of knowing until we reach AGI, but I've never heard a compelling argument for why pure ML would get us there. GPT3 seems more like an argument against that hypothesis (in my view) than for it. Even the best, most expensive mode…

I've always imagined AGI (perhaps naively) as being achieved by clever usage of ML, plus some utilization of classical/symbolic AI from pre-AI winter days, plus probably some unknown elements. For what it's worth, this is my view as well. And I don't think it's particularly naive. Plenty of people have researched and/or are researching aspects of how to do this. But how to combine something like a neural network, wit…

It seems quite clear to me that human brains are not actually doing much symbolic logic. What symbolic logic we do do has been bolted on using other faculties. I think the problem is that reasoning about our own minds is incredible tough. We want there to be some sort of magic sauce to what makes us, us and so we reject things like ANN's that seem somehow too simple. I think it probably is right that we won't just be able to scale up the number of parameters and get human like performance. There are hints that returns start to level off, but I'm also unsure why people are so sure we can't.

Re: Meta Unveils New AI Supercomputer

#174

I can't shake the feelings that a trillion or a quadrillion parameters won't solve the fundamental shortcomings of ML models not being models of artificial intelligence. I guess there's no way of knowing until we reach AGI, but I've never heard a compelling argument for why pure ML would get us there. GPT3 seems more like an argument against that hypothesis (in my view) than for it. Even the best, most expensive mode…

AGI seems hard because each year more and more problems that were previously considered close to AGI are solved.

Playing Chess at a grandmaster level was considered something only a human could do until the 1990s, and now no human has beat the best computer in 17 years while AGI seems further away than ever.

Mark my words: we'll create an AI that can pass the Turing test this decade, but we'll still be as far away from the badly defined general problem as we ever were.

Re: Meta Unveils New AI Supercomputer

#175

Earlier quoted context omitted.

I'd say that the view against ANN gives humans (especially researchers) more "dignity", in the sense that we still need to figure out some deep stuff and not just add hardware. I wouldn't treat this as an argument either way, just an observation. Heuristically, we came to be by a very dumb process of piling up newer generations. If my pet would communicate with me on the level of GPTx, I would be very impressed. That…

> If my pet would communicate with me on the level of GPTx, I would be very impressed. GPTx is not communicating with anyone. It is generating text that resembles text it had in its training set. The fact that human text is normally a form of communication doesn't make generating quasi-random text communication in itself. GPTx is no more communicating than a printer is when printing out text. A cat or dog leading you…

I wonder how well would a dog+GPT/transformer combo work.

Re: Meta Unveils New AI Supercomputer

#176
post #92

Earlier quoted context omitted.

I feel that for there are three requirements for a NN-based AGI, inspired by biology: a). an internal feedback loop that evaluates a possible output without actuating it, and self-modifies the parameters if the possible output is not what it's needed b). the capability (based on a) to model own behaviours without acting on them, and to model other agents behaviours and incorporate that model into the feedback c). the…

I doubt there will be AGI in our lifetime. Maybe some breakthrough happens but it won't be even close to human intelligence.

My observation with statements like this both for and against some event occurring is that you'd have to be very specific with the definition of "AGI" and "human intelligence", otherwise everyone ends up claiming they predicted the outcome correctly (e.g. ray kurzweil's prediction evaluations seem to me like an exercise in motivated reasoning)

Re: Meta Unveils New AI Supercomputer

#178
post #35

Earlier quoted context omitted.

Zuck has vision, especially for what people will want to use. I am looking forward to what FB comes up with here.

Do you get paid in MetaBucks or a real currency? What are the hours and benefits like? Does Zuck wave out a window at you in lieu of cash bonuses?

I'd rather these absurd empire businesses to at least do something interesting with their empires instead of just make it a bit easier to consolidate money. Zuck is uninspiring and hard to like, but at least this move is somewhat visionary.

Re: Meta Unveils New AI Supercomputer

#179
post #94

Earlier quoted context omitted.

Honestly, this single GPU-based install is child's play compared to Google's multiple TPU exoflop supercomputers with hyper-cube optical interconnects. Google's ML setups allow synchronous weight update on thousand+ TPUs...

Tbh I thought I was being trolled with 'hyper-cube optical interconnects'.

Actually, you are right, I mistyped. Although hypercube interconnects exist, and were used, for example, in AS400, system in question uses hypertorus topology.

Re: Meta Unveils New AI Supercomputer

#180
post #94

> Meta’s AI supercomputer houses 6,080 Nvidia graphics-processing units ..... By mid-summer, when the AI Research SuperCluster is fully built, it will house some 16,000 GPUs Honestly ... this is lot of GPUs ... but is it the biggest...? > Model training is done with mixed precision on the NVIDIA DGX SuperPOD-based Selene supercomputer powered by 560 DGX A100 servers networked with HDR InfiniBand in a full fat tree co…

Honestly, this single GPU-based install is child's play compared to Google's multiple TPU exoflop supercomputers with hyper-cube optical interconnects. Google's ML setups allow synchronous weight update on thousand+ TPUs...

for TPUv3 it's 2D torus, not hyper-cube, right? Not sure if TPUv4 topology is externally published, but IIRC hypercubes are basically never used any more.
Post reply on HN