Live data from Hacker News

CUDA Moat Still Alive

semianalysis.com

141–150 of 176 posts

Re: CUDA Moat Still Alive

#141
post #129

Earlier quoted context omitted.

Infiniband is a monopoly from NVidia (Mellanox). Everyone else would much rather use Ethernet which is the actual industry standard.

Others can build infiniband hardware according to the standard. There used to be at least two companies building infiniband hardware until Intel killed QLogic’s infiniband division in a misguided attempt to make its own monopoly. :/

Or they can develop a RoCEv2 that works (which is basically what UltraEthernet is).

Infiniband is not fun - it's a special snowflake of an interconnect that sits parallel to the rest of your datacenter network, and can not really run a standard TCP/IP codebase (yeah IPoIB is a thing but still). Do the Nvidia boxes really need a scale-out IP network as well as an Infiniband network?

Plus the spec is old. Packet spraying and trimming, better ordering guarantees, queue pair scalability... a whole bunch of enhancements have been incorporated into UE all the while being compatible with regular Ethernet.

Qlogic was never really an Infiniband vendor --- their qib driver is still in the Linux codebase and essentially emulates verbs on top of a messaging-based design.

Re: CUDA Moat Still Alive

#142
post #54

> AMD is attempting to vertically integrate next year with their upcoming Pollara 400G NIC, which supports Ultra Ethernet, hopefully making AMD competitive with Nvidia. Infiniband is an industry standard. It is weird to see the industry invent yet another standard to do effectively the same thing just because Nvidia is using it. This “Nvidia does things this way so let’s do it differently” mentality is hurting AMD: *…

> Infiniband is an industry standard Infiniband is not an industry standard lol. Maybe it used to be, but it definitely is not anymore. Most Infiniband vendors are dead. The only product from those days that endures is Cornelis' Omnipath, and even that only emulated the Infiniband API back with its first gen, and then evolved to be its own thing. At this point, Infiniband is as good as a proprietary interconnect only…

Infiniband is very much an industry standard:

https://www.infinibandta.org/member-listing/

As far as I know, everyone is free to sign up with the infiniband trade association and implement the specification.

Furthermore, if you use RDMA over Ethernet, you are using infiniband at a low level. RoCE which enables it was originally called Ethernet over infiniband. It is maintained by the infiniband trade association.

Omnipath was Intel’s failed effort to try to kill an open standard. It purchased QLogic’s infiniband business, killed it in favor of omnipath and sold it when it failed.

Re: CUDA Moat Still Alive

#143
post #135

Earlier quoted context omitted.

Nvidia has a unique problem, wants to move fast and has a shit load of money. No need for Nvidia to go first to an industry standard and neither for AMD. Personally would be great its getting backported but its so far away from an normal use case.

Nvidia is the one who went to an industry standard before their competitors in this space. It was created in 1999 and is called infiniband: https://en.wikipedia.org/wiki/InfiniBand Infiniband is extremely popular in the HPC space, which is why Nvidia adopted it. Everyone else saw Nvidia adopt it and said "Let us make a new network standard to be incompatible". This is mind boggling. Even more mind boggling is that ma…

Sorry to hound you for the third time, but this is wrong:

> Infiniband is extremely popular in the HPC space

Not anymore. There used to be Cray Aries/GNI, psm/psm2, and now there's Slingshot, the new Cornelis stuff etc. There's almost no Infiniband now.

Re: CUDA Moat Still Alive

#144
post #142

Earlier quoted context omitted.

> Infiniband is an industry standard Infiniband is not an industry standard lol. Maybe it used to be, but it definitely is not anymore. Most Infiniband vendors are dead. The only product from those days that endures is Cornelis' Omnipath, and even that only emulated the Infiniband API back with its first gen, and then evolved to be its own thing. At this point, Infiniband is as good as a proprietary interconnect only…

Infiniband is very much an industry standard: https://www.infinibandta.org/member-listing/ As far as I know, everyone is free to sign up with the infiniband trade association and implement the specification. Furthermore, if you use RDMA over Ethernet, you are using infiniband at a low level. RoCE which enables it was originally called Ethernet over infiniband. It is maintained by the infiniband trade association. Omn…

> Furthermore, if you use RDMA over Ethernet, you are using infiniband at a low level. RoCE which enables it was originally called Ethernet over infiniband.

This is wrong.

RDMA over Ethernet is... RDMA over Ethernet. There is no Infiniband involved.

RoCE was motivated by supporting RDMA, which was then an IB-only feature, over regular Ethernet. The user-level APIs are the same (verbs), but the underlying architecture is all different --- it is traditional Ethernet with link-level flow control to make it lossless (pause frames).

Re: CUDA Moat Still Alive

#145
post #13

> Give AMD Engineers more compute and engineering resources to fix and improve the AMD ecosystem, they have very few internal gpu boxes relative to what Nvidia provides to their engineers. This is real. We’ve found ourselves having to give hardware to engineers at AMD because they’re unable to get allocation of it internally.

This is baffling. I’m sure there are many technical reasons I don’t grok that AMD’s job is challenging, but it’s wild that they are dropping the ball on such obvious stuff as this. The prize is trillions of dollars, and they can print hundreds of millions if they can convince the market that they are closing the gap. It’s embarrassing that whoever actually tries to use their product hits these crass bugs (same with g…

In the late 90's, US manufacturers, including high-tech electronics, had 2 mantras:

1) Cash is king

2) Inventory is evil

I think this mindset may still be here, in 2024

Re: CUDA Moat Still Alive

#146
post #142

Earlier quoted context omitted.

Infiniband is very much an industry standard: https://www.infinibandta.org/member-listing/ As far as I know, everyone is free to sign up with the infiniband trade association and implement the specification. Furthermore, if you use RDMA over Ethernet, you are using infiniband at a low level. RoCE which enables it was originally called Ethernet over infiniband. It is maintained by the infiniband trade association. Omn…

> Furthermore, if you use RDMA over Ethernet, you are using infiniband at a low level. RoCE which enables it was originally called Ethernet over infiniband. This is wrong. RDMA over Ethernet is... RDMA over Ethernet. There is no Infiniband involved. RoCE was motivated by supporting RDMA, which was then an IB-only feature, over regular Ethernet. The user-level APIs are the same (verbs), but the underlying architecture…

This is not what I have been told by others.

Re: CUDA Moat Still Alive

#147
post #122

Earlier quoted context omitted.

Steam Decks run on Arch Linux

As means to avoid paying for Windows licenses. All the games that matter are Windows games running via Proton, as Valve has failed to actually build a GNU/Linux native games ecosystem, in spite of UNIX/POSIX underpinnings of Android NDK, PlayStation, the studios hardly bother. The day Microsoft actually decides to challenge Proton, or do a netbooks move on handhelds with XBox OS/Windows, the SteamDeck will lose, just…

As a means of control

Re: CUDA Moat Still Alive

#148
post #95

Earlier quoted context omitted.

You hardly beat someone by copying him. They have way more experience in the field you try to catch up.

AMD doesn't need to beat Nvidia, they just need to match them at a lower price point.

That‘s beating them at the price point.

Re: CUDA Moat Still Alive

#149
post #132

Earlier quoted context omitted.

This was happening before CDNA was even a thing. They didn’t release support even for all GPUs from the same generation and dropped support for GPUs sometime within 6 months of releasing a version that actually “worked”. The entire core architecture behind ROCM is rotten. P.S. NVIDIA usually has multiple CUDA feature levels even within a generation. The difference is that a) they always provide a fallback option, and…

The differences between CUDA feature levels appear minor according to the PTX documentation: https://docs.nvidia.com/cuda/parallel-thread-execution/index... They also appear to be cululative.

It doesn’t matter the point is that they don’t break stuff. You can still compile CUDA today to work on old hardware and your binaries are guaranteed to have forward compatibility.

You don’t get that with ROCm, and this is why it’s garbage unless someone else abstracts all of that from you.

So if Microsoft is happy to maintain an ML as a service solution that just takes prompts and maybe data it’s not your problem.

But if you need to run your own workloads and these can include workloads that are well outside of “AI” and might not be even possible or remotely profitable to have a SAAS wrapper around them it’s all on you.

Re: CUDA Moat Still Alive

#150
post #13

> Give AMD Engineers more compute and engineering resources to fix and improve the AMD ecosystem, they have very few internal gpu boxes relative to what Nvidia provides to their engineers. This is real. We’ve found ourselves having to give hardware to engineers at AMD because they’re unable to get allocation of it internally.

This is baffling. I’m sure there are many technical reasons I don’t grok that AMD’s job is challenging, but it’s wild that they are dropping the ball on such obvious stuff as this. The prize is trillions of dollars, and they can print hundreds of millions if they can convince the market that they are closing the gap. It’s embarrassing that whoever actually tries to use their product hits these crass bugs (same with g…

There's a recent interview with Lisa Su where she basically says she's never been interested in software because hardware is harder, she doesn't believe AMD has any problems in the software department anyway and AMD is doing great in AI. So make of that what you will. Suffice it to say, clearly the AMD board doesn't care either because otherwise they'd replace her.
Post reply on HN