Live data from Hacker News

AMD Alveo V70 AI inference accelerator card

xilinx.com

51–60 of 94 posts

Re: AMD Alveo V70 AI inference accelerator card

#51
post #39

This is inference only. AMD should invest into the full AI stack starting from training. For this they need a product comparable to NVIDIA 4090, so that entry level researchers could use their hardware. Honestly, I don't know why AMD aren't doing that already, they are best positioned to do that in the industry landscape.

The hardware is not their major problem. They have been failing super hard at the software side of machine learning for a solid decade now. It seems like pure management incompetence to me. They need to invest a whole lot more in software, integrating their stuff directly into pytorch/TF/XLA/etc and making sure it works on consumer cards too. The investment would be paid back tenfold. The market is crying out for mor…

Last I checked they see deep learning training as a niche market, their strategy is to try to win big contracts (HPC etc) and then supply software specifically for that. Then "the community" will supply software. Having spent a bunch of time beating my head on this and related walls it's not clear to me that they're entirely wrong from an economic standpoint. Remember that 2/3 public cloud providers have their own chips as well as NVIDIA's so it would be tough to negotiate a good deal. As a user it's super irritating to be stuck on NVIDIA especially when Jensen gets up on stage to say "haha, Moore's law is over, stop expecting our products to get cheaper."

Re: AMD Alveo V70 AI inference accelerator card

#52
post #42
post #39

This is inference only. AMD should invest into the full AI stack starting from training. For this they need a product comparable to NVIDIA 4090, so that entry level researchers could use their hardware. Honestly, I don't know why AMD aren't doing that already, they are best positioned to do that in the industry landscape.

MI200, MI250, and MI300 should work for training. They don't have an exact equivalent to the 4090 but that may be setting the bar too high. Nobody can deliver everything that Nvidia has but better.

It will be harder for a small academic lab to afford MI300, whereas everyone can purchase a few cards costing $1500. And even if I had money to buy MI300, I wouldn't - it is a too risky investment, because we have no idea how suitable they are for common AI research workflows. They need to lower the entry bar, so that people can try out their hardware. Even 80% of the performance of 4090 would be enough at an appropriate price point.

Re: AMD Alveo V70 AI inference accelerator card

#53
post #39

This is inference only. AMD should invest into the full AI stack starting from training. For this they need a product comparable to NVIDIA 4090, so that entry level researchers could use their hardware. Honestly, I don't know why AMD aren't doing that already, they are best positioned to do that in the industry landscape.

Big question is why? Would competing for entry level researchers buy them much?

[deleted]

Re: AMD Alveo V70 AI inference accelerator card

#54
post #47

Douglas Adams said we'd have robots to watch TV for us. That seems to be the designed use case for this. 16gb RAM / 96 video channels ... I haven't done any of that work but it feels like they expect that "96" not to be fully used in practice.

I have models in production that currently monitor ~400 cameras with an addition of 2-3 cameras/month. If it were cheap enough, it would be useful for our use case (Quality Control). We generally pull from cameras roughly 6400 pixels per region of interest, of which one instance may have 4-30 RoIs across N cameras.

Curious where I could learn more about models like this / potentially see some open code outlining tooling / infra required as well?

Re: AMD Alveo V70 AI inference accelerator card

#55
post #39

This is inference only. AMD should invest into the full AI stack starting from training. For this they need a product comparable to NVIDIA 4090, so that entry level researchers could use their hardware. Honestly, I don't know why AMD aren't doing that already, they are best positioned to do that in the industry landscape.

Big question is why? Would competing for entry level researchers buy them much?

Because entry level researchers shape the industry in the long term. I'm in academia, I worked at two universities and I have not seen a lab that uses non-NVIDIA hardware for research. Majority of graduates go to work in the industry.

Re: AMD Alveo V70 AI inference accelerator card

#56
post #39

This is inference only. AMD should invest into the full AI stack starting from training. For this they need a product comparable to NVIDIA 4090, so that entry level researchers could use their hardware. Honestly, I don't know why AMD aren't doing that already, they are best positioned to do that in the industry landscape.

The hardware is not their major problem. They have been failing super hard at the software side of machine learning for a solid decade now. It seems like pure management incompetence to me. They need to invest a whole lot more in software, integrating their stuff directly into pytorch/TF/XLA/etc and making sure it works on consumer cards too. The investment would be paid back tenfold. The market is crying out for mor…

AMD has finite resources, like any company, and they’ve been focusing on CPU/datacenter dominance, which to me is both the safer bet and the more-lucrative bet. It wasn’t that long ago that AMD was on the brink of bankruptcy (~2016), so I appreciate that they’re not trying to divide their attention.

Their attempts at entering the ML space so far have been failures, and they are wise to hold off on really competing with Nvidia until they have the bandwidth to go “all in”. Consciously NOT trying to compete with Nvidia is the reason they didn’t go bankrupt. Their Radeon division minted from 2016-2020 because they focused on a niche Nvidia was neglecting- low-end/eSports (also leveraging their APU expertise to win PS4/Xbox contracts).

I think Nvidia will eventually lose its monopoly on ML/AI stuff as AMD, Apple, Qualcomm, Amazon and Google chip away at their “moat” with their own accelerators/NPUs. As mentioned though, the Nvidia Edge really comes from CUDA and other software, not the hardware. I doubt that Apple, Qualcomm, Amazon or Google will be interested in selling hardware direct to consumers. They want that sweet, sweet cloud money and/or competitive advantages in their phones (like photo processing). I don’t want to be paying AWS $100/mo for a GPU I could pay $600 once for. I do think AMD/RTG will go hard on Nvidia eventually, and it will not matter whether you have an AMD or Nvidia GPU for Tensorflow or spaCy or whatever else.

Re: AMD Alveo V70 AI inference accelerator card

#57
post #50

Earlier quoted context omitted.

> AMD should invest into the full AI stack starting from training. https://www.amd.com/en/graphics/servers-solutions-rocm-ml > For this they need a product comparable to NVIDIA 4090, so that entry level researchers could use their hardware. Why is a high end product a requirement for entry level research?

4090 (or 3090, 1080Ti and so on) is a high-end consumer GPU, but at the same time it is an entry level GPU for AI researchers. Don't forget that workstation cards (RTX 8000) let alone server-grade GPUs such as A100 are an order of magnitude more expensive.

I was doing ML stuff on a GTX 1060 a few years ago. As with everything, it depends on what you’re doing.

Re: AMD Alveo V70 AI inference accelerator card

#58
post #55

Earlier quoted context omitted.

Big question is why? Would competing for entry level researchers buy them much?

Because entry level researchers shape the industry in the long term. I'm in academia, I worked at two universities and I have not seen a lab that uses non-NVIDIA hardware for research. Majority of graduates go to work in the industry.

I used to believe this idea that the tools in academia would carry over to industry. Now I think it's only weakly true, that is there's not really a big barrier to switching to other options like TPU or maybe Trainium (haven't tried it). Supporting independent researchers gives you the Heroku problem, they may like your product but as they get more sophisticated and go to production they'll accept significant pain to save money, scale better, etc. You basically have to re-win that business from scratch and the technical tradeoffs are very different at that point.

Re: AMD Alveo V70 AI inference accelerator card

#60
post #54
post #47

Earlier quoted context omitted.

I have models in production that currently monitor ~400 cameras with an addition of 2-3 cameras/month. If it were cheap enough, it would be useful for our use case (Quality Control). We generally pull from cameras roughly 6400 pixels per region of interest, of which one instance may have 4-30 RoIs across N cameras.

Curious where I could learn more about models like this / potentially see some open code outlining tooling / infra required as well?

Not sure about open information about the models, but from tooling/infra we are running k8s w/ in-house API for image acquisition. Features are defined as x,y coordinates denoting center of a feature, with a pixel count denoting size of rectangle in each direction from center.
Post reply on HN