Earlier quoted context omitted.
> I'm also curious if there is something inherent to ARM that would stop them from making similar optimizations to the AMD chips. Yeah ARM isn’t x86. Apple made big gains by leveraging the difference in instruction length and complexity between x86 and ARM. Something AMD can’t do, and they’ve said as much.
I still don't have a clear understanding. Apple also made big gains by integrating performance sensitive stuff on the package. AMD could do the same. I wonder if they are constrained by some hardware-level interoperability requirements across the motherboard, like with chipsets or DMA controllers or whatnot made by different companies. Apple clearly has a nice advantage here.
The Ampere Altra Review: 2x 80 Cores Arm Server Performance Monster
211–220 of 266 posts
Re: The Ampere Altra Review: 2x 80 Cores Arm Server Performance Monster
#212How close are Intel to releasing anything that's a "next gen" core design and not just incremental design change or node shrink? When will Intel's "Zen2" arrive? Are there any rumors?
Rocket Lake is supposedly a "backport" of Sunny Cove to 14nm for desktop, coming Q1 2020.
Re: The Ampere Altra Review: 2x 80 Cores Arm Server Performance Monster
#213Earlier quoted context omitted.
AMD might already be working on an ARM design[1]. [1] https://www.techspot.com/news/87851-amd-rumored-working-arm-... .
From what I recall, the K12 core was essentially the same micro-architecture as Zen, but with ARM instruction decoders. There were some vague statements that the ISA allowed some optimizations compared to the x86 version, but we never got to see what that really means, since the K12 project was shelved to focus on x86. I wouldn't be surprised if AMD has been sitting on K12 and keeping it just updated enough so that t…
Re: The Ampere Altra Review: 2x 80 Cores Arm Server Performance Monster
#214Earlier quoted context omitted.
There's an interesting rumor that says AMD is building an ARM SoC for OEMs...
If I were them I wouldn't even bother with anything pre-Arm v9. That's going to draw all the hype, especially when Apple announces the Arm v9-based M2 next year. And then I'm sure most sheep-like OEMs will say "Oh, we want THAT, too. Where is it - we want it yesterday!" But AMD won't be able to provide one too soon, because they would've gone all in on Arm v8, and they'd want to squeeze at least a couple of generatio…
ARMv8 (i.e. 64-bit ARM), of course, has been around for a while now.
[1] Keep in mind that anything on Reddit not confirmed elsewhere is very likely either wild speculation or made up fiction.
Re: The Ampere Altra Review: 2x 80 Cores Arm Server Performance Monster
#215Earlier quoted context omitted.
> Xeon Platinum supports 8-socket machines. 6x memory channels x 8-sockets == 48x DIMMs for your massive 6TB RAM servers. Both AMD Epyc Rome and Ampere Altra can support 8TB of RAM in dual socket machines, so that spec isn't particularly impressive anymore. > Platinum is crazy priced for a crazy reason: its a building block to truly massive computers that only IBM matches. Meh. A dual-socket Rome system with 128 core…
But not with 6-channels per socket (48-channels total across an 8-socket computer). That's a heck of a lot of bandwidth. I was doing the specs from memory and conservatively: 128GB sticks across 48-channels. https://www.supermicro.com/en/products/system/7U/7088/SYS-70... This 8-way Supermicro system supports 192x DIMMs. So... 128GB sticks x 192 == 24TBs of RAM or so, maybe 48TBs. > At a certain point, you're better o…
But for that bandwidth to be used efficiently, the processes on each NUMA node need to almost exclusively limit themselves to memory attached to their node -- at which point, well-written software could probably do just as good spread out over several machines that are connected by multiple 100Gb network links, and then you saved two or three bajillion dollars.
If you're heavily using the bandwidth over the NUMA interconnect, then you're not going to be using the memory bandwidth very effectively (and likely not really using the processor cores effectively), and that's when NVDIMMs like Optane Persistent Memory could give you large amounts of bulk memory storage on a smaller system.
Or just use a number of Intel's new PCIe 4.0 Optane SSDs in a single machine in place of the extra memory and memory channels... the latency isn't the same as RAM, but it's much closer than traditional SSDs, and the bandwidth per SSD is like 7GB/s, which is impressive.
It all depends on the application at hand, but there are solutions that cost a lot less than the Platinum machines for virtually every problem, in my opinion.
I don't know... perhaps I'm too cynical of these cost-ineffective systems that just happen to be large, and I should be more impressed.
> NUMA scales better than RDMA / Ethernet / Infiniband. A lot, lot, LOT better.
Disagree.
Re: The Ampere Altra Review: 2x 80 Cores Arm Server Performance Monster
#216Earlier quoted context omitted.
But not with 6-channels per socket (48-channels total across an 8-socket computer). That's a heck of a lot of bandwidth. I was doing the specs from memory and conservatively: 128GB sticks across 48-channels. https://www.supermicro.com/en/products/system/7U/7088/SYS-70... This 8-way Supermicro system supports 192x DIMMs. So... 128GB sticks x 192 == 24TBs of RAM or so, maybe 48TBs. > At a certain point, you're better o…
Sure, I agree it's a lot of bandwidth. But for that bandwidth to be used efficiently, the processes on each NUMA node need to almost exclusively limit themselves to memory attached to their node -- at which point, well-written software could probably do just as good spread out over several machines that are connected by multiple 100Gb network links, and then you saved two or three bajillion dollars. If you're heavily…
Wait, so latency over NUMA is too much, but you're willing to incur a 100Gb network link? Intel's UPI is like 40GBps (Giga-BYTEs, not bits) at sub 500ns latency.
100Gb Ethernet is what? 10GBps in practice? 1/4th the bandwidth with 4x more latency (in the microseconds) or something?
That's a PCIe latency penalty (x2, for the sender + the receiver). That's a penalty for electrical -> optical, and back again.
Any latency, or bandwidth, bound problem is going to prefer a NUMA-link rather than 100Gbps over PCIe.
Re: The Ampere Altra Review: 2x 80 Cores Arm Server Performance Monster
#217Earlier quoted context omitted.
The RISC-V ISA would be the better bet. a. They start investing now and they can have some leverage in the spec. b. The spec is minimal and open to expansion that it sets itself up for more domain specific processors. Which we will be seeing alot more of thanks to the slowing of transistors per dollar growth. c. There is a pretty big theme of moving toward more open and democratized standards in tech. d. They would b…
I'd love to see RISC-V or another open source ISA become the dominant design. But I'm not sure RISC-V is there yet. Companies have been pouring resources into optimizing ARM CPUs for years now, and they're now starting to get to the point where their performance competitive with x86. RISC-V is just getting started. I don't know if there's a RISC-V CPU that's comes anywhere close to the raw performance of higher end A…
I doubt that strategy would work. The problem in the PC space is that there's no Apple-like leader who can boldy drive a change in CPU architecture. Whoever is first out of the gate bears all the risk of the change not catching on. If you lose the bet, you've wasted massive R&D spend to make a lemon. Take a look at the reviews of Microsoft's ARM surface laptops to see what I mean here.
I think there's an opportunity over the next 5 years or so for PC makers to ride the M1's coattails and shift from x86 to ARM. Especially given intel's failure to move off 14nm. But I doubt lightning will strike twice. If windows moves to ARM now, we'll be stuck there for at least another few decades. And the technical argument for RISC-V will be much weaker if ARM rules the roost.
Weirdly the strongest counterargument I can think of is due to electron. Chrome will maintain first class support for any and every popular CPU architecture. So the more that desktop software is written on the web and in electron apps, the easier any architecture transition will be down the road. That and the server space. Linux already supports RISC-V extremely well given how few linux-capable RV chips there are in the wild.
Re: The Ampere Altra Review: 2x 80 Cores Arm Server Performance Monster
#218Earlier quoted context omitted.
AMD should see the writing on the wall and start designing ARM CPUs. I think that would be very exciting because they have the engineering expertise and talent necessary to compete against Apple's M1. I don't know of any other company that can do that (besides Intel, but you know...). Additionally, there is a huge hole open right now for ARM on everything else that is not an Apple product. Apple normalized ARM on con…
> AMD should see the writing on the wall and start designing ARM CPUs. This is a really tough strategy decision. Focus all energy on delivering a knock-out blow to Intel (in x86/64) or try to do other stuff too? For the next several years at least, there is still a whole lot of money to be made selling x86/64 chips. (I don't know if Microsoft has plans to move mainstream Windows users off x86/64, but even if they do,…
- It's in both AMD and Intel's interests for x86 to remain dominant for as long as possible. If they could it would be in their interest to drive Arm new entrants out of the market (I don't believe that will happen though).
- Introducing an Arm product (or even hinting at one) would add credibility to new entrants and undermine existing products.
- They could almost certainly switch to an Arm design reasonably quickly if they needed to.
There is also some uncertainty about the impact that Nvidia will have on Arm when it takes it over. x86 is AMD's jointly owned ISA - they have no control over the Arm ISA.
Re: The Ampere Altra Review: 2x 80 Cores Arm Server Performance Monster
#219Earlier quoted context omitted.
AMD should see the writing on the wall and start designing ARM CPUs. I think that would be very exciting because they have the engineering expertise and talent necessary to compete against Apple's M1. I don't know of any other company that can do that (besides Intel, but you know...). Additionally, there is a huge hole open right now for ARM on everything else that is not an Apple product. Apple normalized ARM on con…
> AMD should see the writing on the wall and start designing ARM CPUs. ARM's brand new 80-core system struggles to match AMD's year old 64-core system, while also being entirely unable to run some of the workloads either at all or in the 2S configuration. And with mostly comparable power draw and a much worse turbo story that ARM is trying to pretend is somehow a good thing. Competition is great, but it's a bit prema…
Re: The Ampere Altra Review: 2x 80 Cores Arm Server Performance Monster
#220Earlier quoted context omitted.
Sure, I agree it's a lot of bandwidth. But for that bandwidth to be used efficiently, the processes on each NUMA node need to almost exclusively limit themselves to memory attached to their node -- at which point, well-written software could probably do just as good spread out over several machines that are connected by multiple 100Gb network links, and then you saved two or three bajillion dollars. If you're heavily…
> But for that bandwidth to be used efficiently, the processes on each NUMA node need to almost exclusively limit themselves to memory attached to their node -- at which point, well-written software could probably do just as good spread out over several machines that are connected by multiple 100Gb network links, and then you saved two or three bajillion dollars. Wait, so latency over NUMA is too much, but you're wil…
The point is not just the latency, but latency and bandwidth.
If the application is relying heavily on the NUMA interconnect to transfer tons of data, it's not going to be making efficient use of the processor cores or the RAM bandwidth. It's a total all around bust. You're just wasting money at that point. ^1
If you aren't relying heavily on the NUMA interconnect, and each node is operating independently with only small bits of information exchanged across the interconnect, then you'd save a metric ton of money by switching to separate machines and using just high speed fiber network links -- such as 100Gb.
I'm not saying that network would be better than the NUMA interconnect. I'm saying that you're not going to be having a happy day if you're relying on the interconnect for large amounts of data transfer.
The only situation where the NUMA set up is better is if you have a need for frequent, low latency communication between NUMA nodes... where very little data is being transferred between nodes, so the entire problem is just latency.
At that point, you're still suffering a lot by the NUMA interconnect, and it would be better to use larger processors... such as Epyc Rome processors.
So you really have to be in a very obscure situation which can't fit onto a dual socket Rome server, but can fit within a machine less than 2x larger. (28 cores * 8 sockets is less than 2x larger than 64 cores * 2 sockets)
You see how complicated this is and how absurdly niche those Platinum 8-socket machines are? They're almost never the right choice.
^1: The exception is if you have an even more niche use case that somehow is built for exactly this scenario in mind, and manages to balance everything perfectly. Such a software is almost certainly ridiculously overcomplicated at this point.