Live data from Hacker News

Amazon launches Trainium3

techcrunch.com

41–50 of 75 posts

Re: Amazon launches Trainium3

#41

I've had to repeatedly tell our AWS account reps that we're not even a little interested in the Trainium or Inferentia instances unless they have a provably reliable track record of working with the standard libraries we have to use like Transformers and PyTorch. I know they claim they work, but that's only on their happy path with their very specific AMI's and the nightmare that is the neuron SDK. You try to do any…

IMO AWS once you get off the core services is full of beta services. S3, Dynamo, Lambda, ECS, etc are all solid. But there are a lot of services they have that have some big rough patches.

RDS, Route53, and Elasticache are decent, too. But yes, I've also been bitten badly in the distant past by attempting to rely on their higher-level services. I guess some things don't change.

I wonder if the difference is stuff they dogfood versus stuff they don't?

Re: Amazon launches Trainium3

#42

I've had to repeatedly tell our AWS account reps that we're not even a little interested in the Trainium or Inferentia instances unless they have a provably reliable track record of working with the standard libraries we have to use like Transformers and PyTorch. I know they claim they work, but that's only on their happy path with their very specific AMI's and the nightmare that is the neuron SDK. You try to do any…

IMO AWS once you get off the core services is full of beta services. S3, Dynamo, Lambda, ECS, etc are all solid. But there are a lot of services they have that have some big rough patches.

>But there are a lot of services they have that have some big rough patches.

Enlight us...

Re: Amazon launches Trainium3

#43
Heavens to Betsy, I don’t know if you can hear me, But try supporting these things if you actually want them to be successful. About the 3rd day into trying to roll your own LMI container in sagemaker because they haven’t updated the vLLM version in 6 months and you can’t run a regular sagemaker endpoint because of a ridiculous 60s timeout that was determined to be adequate 8 years ago. I can only imagine the hell that awaits the developer that decides to try their custom silicon.

Re: Amazon launches Trainium3

#44

Earlier quoted context omitted.

Can you link to the press releases? The only one I'm aware of by Anthropic says they will use Tranium for future LLMs, not that they are using them.

This is the Anthropic press release from last year saying they will use Trainium: https://www.anthropic.com/news/anthropic-amazon-trainium This is the AWS press release from last month saying Anthropic is using 500k Trainium chips and will use 500k more: https://finance.yahoo.com/news/amazon-says-anthropic-will-us... And this is the Anthropic press release from last month saying they will use more Google TPUs but als…

There is no press release saying that they are using 500k trainium chips. You can search on amazon's site.

Re: Amazon launches Trainium3

#45

Earlier quoted context omitted.

IMO AWS once you get off the core services is full of beta services. S3, Dynamo, Lambda, ECS, etc are all solid. But there are a lot of services they have that have some big rough patches.

RDS, Route53, and Elasticache are decent, too. But yes, I've also been bitten badly in the distant past by attempting to rely on their higher-level services. I guess some things don't change. I wonder if the difference is stuff they dogfood versus stuff they don't?

A big problem for a when three AWS teams launch the same thing. Lowers confidence in dogfooding the “right” one.

Re: Amazon launches Trainium3

#46

I've had to repeatedly tell our AWS account reps that we're not even a little interested in the Trainium or Inferentia instances unless they have a provably reliable track record of working with the standard libraries we have to use like Transformers and PyTorch. I know they claim they work, but that's only on their happy path with their very specific AMI's and the nightmare that is the neuron SDK. You try to do any…

spoiler alert, they don't work without a lot of custom code

Re: Amazon launches Trainium3

#47

Earlier quoted context omitted.

This is the Anthropic press release from last year saying they will use Trainium: https://www.anthropic.com/news/anthropic-amazon-trainium This is the AWS press release from last month saying Anthropic is using 500k Trainium chips and will use 500k more: https://finance.yahoo.com/news/amazon-says-anthropic-will-us... And this is the Anthropic press release from last month saying they will use more Google TPUs but als…

There is no press release saying that they are using 500k trainium chips. You can search on amazon's site.

In the subheading of https://www.aboutamazon.com/news/aws/aws-project-rainier-ai-... Unless there are other Project Ranier customers that it's counting.

Re: Amazon launches Trainium3

#48
post #6

Earlier quoted context omitted.

Training. It's in the name.

Ironically these chips are being targeted at inference as well (the AWS CEO acknowledged the difficulties in naming things during the announcement).

Perhaps they should do their training on their Inferentia chips and see how that works out?

Re: Amazon launches Trainium3

#49

I've had to repeatedly tell our AWS account reps that we're not even a little interested in the Trainium or Inferentia instances unless they have a provably reliable track record of working with the standard libraries we have to use like Transformers and PyTorch. I know they claim they work, but that's only on their happy path with their very specific AMI's and the nightmare that is the neuron SDK. You try to do any…

Agree, Google put a ton of work into making TPUs usable with the ecosystem. Given Amazon’s track record I can’t imagine they would ever do that.
Post reply on HN