Live data from Hacker News

Andromeda Cluster: 10 Exaflops* for Startups from Nat Friedman and Daniel Gross

andromedacluster.com

11–20 of 45 posts

Re: Andromeda Cluster: 10 Exaflops* for Startups from Nat Friedman and Daniel Gross

#11
Can the creators explain in more detail: how is this different from (for example) the OpenAI cluster that MSFT built in Azure? Is it hosted in an existing cloud provider, or in a data center? Which data center? Who admins the system, is there an SRE team in case it goes down during training? And can you attempt ot run the same benchmarks that Top500 uses to determine what your double precision flops are and give that number in addition to your "10 exaflops" (which I believe is single precision).

Re: Andromeda Cluster: 10 Exaflops* for Startups from Nat Friedman and Daniel Gross

#13
post #11

Can the creators explain in more detail: how is this different from (for example) the OpenAI cluster that MSFT built in Azure? Is it hosted in an existing cloud provider, or in a data center? Which data center? Who admins the system, is there an SRE team in case it goes down during training? And can you attempt ot run the same benchmarks that Top500 uses to determine what your double precision flops are and give that…

Pretty sure it's FP8, not singles. (Which for the H100 makes a 60x difference.)

Re: Andromeda Cluster: 10 Exaflops* for Startups from Nat Friedman and Daniel Gross

#15
post #13
post #11

Can the creators explain in more detail: how is this different from (for example) the OpenAI cluster that MSFT built in Azure? Is it hosted in an existing cloud provider, or in a data center? Which data center? Who admins the system, is there an SRE team in case it goes down during training? And can you attempt ot run the same benchmarks that Top500 uses to determine what your double precision flops are and give that…

Pretty sure it's FP8, not singles. (Which for the H100 makes a 60x difference.)

as an ex-supercomputer nerd (where the fastest system in teh world finally reached over 1 exaflops of double precision), it seems awfully weird to call FP8 "flops". There's nothing truly wrong with it (since "flops" is a fairly poorly defined term), but it makes it clear that ML supercomputers are very different beasts from classic supercomputers. And also makes me wonder if/when the classic folks will try to make more codes work correctly with smaller precision (for example, in molecular dynamics).

Re: Andromeda Cluster: 10 Exaflops* for Startups from Nat Friedman and Daniel Gross

#19

Same guys behind https://aigrant.org , maybe it's mainly as a way to get dealflow?

AI Grant, back in 2018, offered £2500 and got all sorts of skeptical folks doubting Nat's and Daniel's motives (they will steal IP! there's some gotcha here!): https://news.ycombinator.com/item?id=16760736

They offer 100x that at £250000 per team now, plus this humongous GPU cluster. Way to start small and work your way to this. Amazing execution.

Re: Andromeda Cluster: 10 Exaflops* for Startups from Nat Friedman and Daniel Gross

#20

Forget LinPack and friends. Jack Dongarra is going to need to switch to the new metric for supercomputers—-kilograms of H100 GPUs—- about 3,300 give or take a few grams for this system.

That would be mostly heatsinks, right? If you switch to liquid cooling does your score change?
Post reply on HN