Live data from Hacker News

New AMD EPYC-based Compute Engine family, now in beta

cloud.google.com

61–70 of 146 posts

Re: New AMD EPYC-based Compute Engine family, now in beta

#61
post #31

Since people from Google Cloud are likely here, one thing I'd like to ask/talk about: are we getting too many options for compute? One of the great things about Google Cloud was that it was very easy to order. None of this "t2.large" where you'd have to look up how much memory and CPU that it has and potentially how many credits you're going to get per hour and such. I think Google Cloud is still easier, but it's get…

> are we getting too many options for compute

As compared to what, Azure? :)

Re: New AMD EPYC-based Compute Engine family, now in beta

#62

Earlier quoted context omitted.

This article goes into great detail: https://www.servethehome.com/amd-epyc-7002-series-rome-deliv... Another follow-up article: https://www.servethehome.com/amd-epyc-7702p-review-redefinin... AMD is offering incredible performance on every metric: single threaded, multithreaded, total RAM per socket, PCIe 4.0, power consumption, total performance, total price, performance for price, etc. Outside of some very niche ap…

There are literally no server benchmarks anywhere in those articles. Unless you are planning to run a distributed C build node there's nothing in these articles that can inform your choice. Distributed building is an extremely narrow, niche use case. Where are the nginx and mysql and grpc benchmarks?

There's absolutely no way you could've read those articles in the last four minutes. They go into great detail about what makes the Rome processors so important -- it's not just some random amalgamation of benchmarks, but the benchmarks serve to provide hard numbers that back up the textual analysis.

The benchmarks are not just "distributed compilation" either... that's a very misleading characterization. There was one compilation benchmark for the Linux kernel, and that's the only compilation benchmark I remember seeing.

No one benchmarks nginx because nginx can easily saturate the network card on a server without saturating the processor.

Here's a postgres benchmark: https://openbenchmarking.org/embed.php?i=2002066-VE-XEONEPYC...

Or a rocksdb benchmark: https://openbenchmarking.org/embed.php?i=2002066-VE-XEONEPYC...

MariaDB was a rare win for Intel: https://openbenchmarking.org/embed.php?i=2002066-VE-XEONEPYC...

("rare win" is literally the wording used in the Phoronix article: https://www.phoronix.com/scan.php?page=article&item=linux55-...)

ServeTheHome had access to more comprehensive Intel hardware, so I preferred to link to their articles, but Phoronix saw more of the same stuff.

Intel was thoroughly destroyed in every Linux review of Rome vs Intel's latest that I've seen. Intel can eke out some rare wins when applications are heavily optimized for the nuances of their CPUs, but it's not guaranteed even then.

If you can't be bothered to read articles to understand the answer to the question you asked, then this is my last reply.

Re: New AMD EPYC-based Compute Engine family, now in beta

#64
post #31

Since people from Google Cloud are likely here, one thing I'd like to ask/talk about: are we getting too many options for compute? One of the great things about Google Cloud was that it was very easy to order. None of this "t2.large" where you'd have to look up how much memory and CPU that it has and potentially how many credits you're going to get per hour and such. I think Google Cloud is still easier, but it's get…

Customers like having choices. Enterprises typically will "certify" one config and would like to stay on that till they absolutely need to move to something else.

That reflects the lumbering, bureaucratic nature of enterprises.

Re: New AMD EPYC-based Compute Engine family, now in beta

#65
post #31

Since people from Google Cloud are likely here, one thing I'd like to ask/talk about: are we getting too many options for compute? One of the great things about Google Cloud was that it was very easy to order. None of this "t2.large" where you'd have to look up how much memory and CPU that it has and potentially how many credits you're going to get per hour and such. I think Google Cloud is still easier, but it's get…

Different chipsets may have slightly different capabilities. For example, I’ve been using NVIDIA RAPIDS recently. Not all NVIDIA cards support this particular framework’s needs. Sometimes you need to specifically direct customer installations to a specific type of card or chipset.

Re: New AMD EPYC-based Compute Engine family, now in beta

#66
post #48
post #31

Since people from Google Cloud are likely here, one thing I'd like to ask/talk about: are we getting too many options for compute? One of the great things about Google Cloud was that it was very easy to order. None of this "t2.large" where you'd have to look up how much memory and CPU that it has and potentially how many credits you're going to get per hour and such. I think Google Cloud is still easier, but it's get…

I think for enterprise businesses, people just love choices. I don't know about GCP but I do know about highly paid AWS consultants producing detailed comparisons between instance types and make recommendations for companies to "save money." Or maybe some people just like the thrill of using spreadsheets and navigating the puzzle of pricing.

Or they prefer the certainty of being able to test software on specific hardware setups and be able to give customers higher levels of confidence.

Re: New AMD EPYC-based Compute Engine family, now in beta

#67
post #64

Earlier quoted context omitted.

Customers like having choices. Enterprises typically will "certify" one config and would like to stay on that till they absolutely need to move to something else.

That reflects the lumbering, bureaucratic nature of enterprises.

Sure, I am sure they have something to say about how smaller internet companies move fast and break things with no concern for how it affects customers. I am sure you would rather have your bank just work than have some nifty tool with insufficient testing and that flakes out every other day.

Everyone to their own, I am just stating that there is a need.

Re: New AMD EPYC-based Compute Engine family, now in beta

#68

Earlier quoted context omitted.

There are literally no server benchmarks anywhere in those articles. Unless you are planning to run a distributed C build node there's nothing in these articles that can inform your choice. Distributed building is an extremely narrow, niche use case. Where are the nginx and mysql and grpc benchmarks?

There's absolutely no way you could've read those articles in the last four minutes. They go into great detail about what makes the Rome processors so important -- it's not just some random amalgamation of benchmarks, but the benchmarks serve to provide hard numbers that back up the textual analysis. The benchmarks are not just "distributed compilation" either... that's a very misleading characterization. There was o…

Every morning, I beg my God to make morons stop replying to me on HN. Today is the first day anyone has promised to make my dream come true.

I didn't read those articles in the last 4 minutes because I read them when they were published. A massively parallel run of 7zip was a really stupid benchmark in August and it remains stupid today.

These other benchmarks are certainly more relevant but none of them jumps out at me as a killer claim. An EPYC 7402 with 50% more cores, drawing 80% more power, and costing 35% more dollars than a Xeon Silver 4216 delivers 24% more pgsql ops per second. What TCO equation do you plug that into? I would describe these results as mixed.

Re: New AMD EPYC-based Compute Engine family, now in beta

#69
post #55
post #52

Earlier quoted context omitted.

They really are doing something disruptive. I can't quite remember if this is correct (it has been a while since I last studied business), but in business there is a "blue ocean strategy". The basic premise is, if you can provide a product for half the price, with the twice the value, you will destroy the incumbent. What AMD is doing is really insane in my opinion. I'm not sure if they are pricing their processors lo…

Smaller chips have better yield. As AMD's current chips are composed from several smaller ones (I believe two or three), each composite has better yield than one bigger of same real estate size. So yes, they figured out how to produce cheaper solutions.

> As AMD's current chips are composed from several smaller ones (I believe two or three)

For EPYC, AMD is using nine chips: https://images.anandtech.com/doci/13561/amd_rome-678_678x452...

That's 1x I/O chip (kind of like a router), and 8x chips, each of which has 8x cores on it. Total for 64-cores / 128-threads across 8-compute chips, talking together through a central 1x I/O and Memory chip.

The I/O chip is the biggest for reasons: 1. Its made on a cheaper process. 2. It has worse performance than the compute chips. 3. Its required to be big because driving external I/O requires more power.

So the I/O chip can be made on a cheap / inefficient 14nm process, while the CPUs can be made on a more expensive 7nm process (maximizing clock rates, power-efficiency). The big I/O ports are going to eat up a lot of power regardless of 7nm or 14nm process, so might as well save money here.

Re: New AMD EPYC-based Compute Engine family, now in beta

#70
post #31

Since people from Google Cloud are likely here, one thing I'd like to ask/talk about: are we getting too many options for compute? One of the great things about Google Cloud was that it was very easy to order. None of this "t2.large" where you'd have to look up how much memory and CPU that it has and potentially how many credits you're going to get per hour and such. I think Google Cloud is still easier, but it's get…

> For example, the N2D instances are basically the price of the N1 instances or even cheaper with committed-use discounts. Given that they provide 39% more performance, should the N1 instances be considered obsolete once the N2D exits beta? As the name implies, N2 is a newer generation than N1. I don't think Google has announced any official N1 deprecation timeline, but that product line clearly has an expiration dat…

> If Google tried to streamline everything...they'd have another cohort of users screaming that the product doesn't meet their needs

Except that they could simplify it without reducing flexibility.

For example, the difference between E-series and N-series is that E-series instances have the whole balancing thing. Instead of being a different instance type, it could be simplified into an option on available types and it would just give you a discount.

Likewise, some of it is about consistency. How much should sustained-usage give you a discount? 20%? 30%? 0%? There seems to be little difference to Google whether sustained-use an E2, N2, N2D, or N1 in terms of their costs and planning and yet the discount varies a lot.

It's not about fewer choices. It's more that the choices aren't internally consistent. N2 instances are supposed to be simply superior to N1 instance, but N1 instances cost the same as N2 instances for 1-year contract, 3-year contract, and on-demand. They're only more expensive for sustained-use which seems odd. Likewise, E2 instances are meant to give you a discount and they do give you a discount for 3 out of the 4 usage scenarios. The point is that there's no real reason for the pricing not to be consistent across the 4 usage scenarios (1-year, 3-year, on-demand, and sustained-use). That's where the complexity creeps in.

It's really easy to look and say, "ok, I have E2, N2D, and N2 instances in ascending price order and I can choose what I want." Except that the pricing doesn't work out consistently.

> N2 instances are likely faster on a per-core basis

Are they meant to be? Google's announcement makes it seem like they should be equivalent: "N2D instances provide savings of up to 13% over comparable N-series instances".

--

The point I'm trying to make isn't that they shouldn't offer choice. It's that the choice should be consistent to be easily understandable. E2 instances should offer a consistent discount. If N2 machines are the same price as N1 machines across 3 usage scenarios, they should be the same price across all 4. When you browse the pricing page, you can get into situations where you start thinking, "ok, the N1 instances are cheaper so do I need the improvements of the N2?" And then you start looking and you're like, "wait, the N2s are the same price....oh, just the same price most of the time." Then you start thinking, "I can certainly deal with the E2's balancing...oh, but it's the same price...well, it's cheaper except for sustained-use".

There doesn't seem to be a reason why sustained-use on N1s should be cheaper for Google than sustained-use on N2s. There doesn't seem to be a reason why sustained-use on E2s offers no discount - especially given that the 1-year E2 price offers the same 37% discount that the N1s offer.

It would be nice to go to the page and say, "I'm going with E2s." However, I go to the page and it's more like, "I'm going with E2s when I am going to do a 1-year commitment, but I'm going with N2Ds when I'm doing sustained-use without a commitment since those are the same price for better hardware with seemingly no reason and the N1s are just equal or more expensive so why don't they just move them to a 'legacy machine types' page". It's the inconsistency in the pricing for seemingly no reason that makes it tough, not the options. The fact that N2Ds are the same monthly price as E2s for sustained-use, but E2s are significantly cheaper in all other scenarios is the type of complexity that's the annoying bit.

EDIT: As an example, E2 instances are 20.7% cheaper on-demand, 20.7% cheaper with 1-year commitment, and 20.7% cheaper with 3-year commitment compared to N2D instances. That's wonderful consistency. Then we look at sustained use and it's 0.9% cheaper with no real explanation why. It's a weird pricing artifact that means that you aren't choosing, "this is the correct machine for the price/performance balance I'm looking for" but rather you're balancing three things: price, performance, and billing scenario.

Post reply on HN