Great practical information. Nice to see people who know what they are talking about putting data out there. I hope eventually these persistent HN memes about M1 memory will die: that it's "on-die" (it's not), that it's the only CPU using LPDDR4X-4267 (it's not), or that it's faster because the memory is 2mm closer to the CPU (not that either). It's faster because it has more microarchitectural resources. It can load…
>that it's "on-die" (it's not) It appears to be mounted on the same chip package. Why did Apple do this if not for speed?
Memory access on the Apple M1 processor
221–230 of 278 posts
Re: Memory access on the Apple M1 processor
#222Earlier quoted context omitted.
Microsoft doesn't need to acquire Intel, they need to do what Apple did and acquire a stellar ARM design house that will build a chip with x86 translation, tailored to accelerate the typical workloads on Windows machines and sell those chips to the likes of Dell and Lenovo and tell developers "ARM Windows is the future, x86 Windows will be sunset in 5 years and no longer supported by us, start porting your apps ASAP…
Microsoft has a pretty good relationship with AMD from the Xbox. AMD already made an Arm Opteron. Windows has been multiplatform since NT 3.1 (Alpha, MIPS) and then in 3.51 adding in PowerPC. You can download Windows for Arm for free and run in on a Raspberry Pi. Microsoft has at least one homegrown processor that it has ported Windows and Linux to with the confusingly named 'Edge'. https://www.theregister.com/2018/0…
Re: Memory access on the Apple M1 processor
#223Earlier quoted context omitted.
As a data scientist, I feel this. Intel and AMD don't own an OS or an app store, and you might be surprised how hard it is to get good data. Data is the new gold. If a company that can corner a piece of the market, they can collect data no one else can, and from that companies are often forced to partner or they can't properly provide services that will keep them competitive.
This makes me think that any sort of data advantage Apple may have has nothing to do with them owning an OS. Intel has a massive computer network, managed by their own IT team, just like any other large corporation. Intel could collect whatever performance data they want from actual users of actual programs just as easily as Apple could.
One early benchmark showed allocating and destroying an NSObject performing drastically better on the M1 vs recent Intel Macs. This wasn’t an accident. It’s probably not representative of performance overall. They have enough vertical integration to make their own first party solutions clear optimization targets.
Re: Memory access on the Apple M1 processor
#224Earlier quoted context omitted.
Interesting point. This would suggest pretty sizable synergies from the oft-rumored Microsoft acquisition of Intel.
Microsoft doesn't need to acquire Intel, they need to do what Apple did and acquire a stellar ARM design house that will build a chip with x86 translation, tailored to accelerate the typical workloads on Windows machines and sell those chips to the likes of Dell and Lenovo and tell developers "ARM Windows is the future, x86 Windows will be sunset in 5 years and no longer supported by us, start porting your apps ASAP…
Re: Memory access on the Apple M1 processor
#225For people who know more about this stuff than me: are these sorts optimizations only possible because Apple controls the whole stack and can make the hardware & OS/software perfectly match up with one another or is this something that Intel can do but doesn't for some reasons (tradeoffs)?
Sort of; Intel and AMD are stuck with the variable width instruction isa that exists due to historical evolution. To do something different you need a new isa. Intel tried this with Itanium a while back and failed because it is difficult to get software developers to target a new isa and provide compilers and compiled code for everything unless you use a translation layer. Apple is one step ahead here because their c…
Without AMD everyone would eventually be dragged into adopting Itanium.
Re: Memory access on the Apple M1 processor
#226I can't find any info about the memory bus of apple m1. Is it 8 channels 16 bit each? That's drastically different from AMDs 2 channels 64 bit each. It looks like apple m1 is much less eager when caching memory rows. Maybe because it doesn't have l3 cache. Edit: This test utilizes the 8x16bit memory bus of apple m1 fully. It's mostly just fetching random locations from memory, which can all be parallelized by the cpu…
I'm a bit of two minds about this: on the one hand, for a long time I've wanted a language for writing allocators which is more explicit about memory, and offers good abstractions for low-level memory operations (maybe Zig is going in this direction). In some sense, it feels like the move towards programmers thinking less about memory management has been a bit of a dead-end, and what we really want is better tools for memory management. Fragmentation in terms of how processors handle memory goes against this goal in some ways.
On the other hand, it's a bit of a "holy grail" to imagine a hardware stack which obviates the need for memory optimization, and really does treat loading from and storing to memory anywhere on the heap as the same. But I imagine that the interesting things which the M1 is doing with memory are helping a lot with the worst case performance, and maybe even average case performance, but they're probably not doing much for the best case.
Re: Memory access on the Apple M1 processor
#227Earlier quoted context omitted.
What other laptop ships with LPDDR4X clocked at 4267? I agree though that being closer to the cpu isn't having any appreciable effect on latency, but being soldered close to the cpu probably does make it easier for them to hit that high clock rate.
As WMF mentions, Tiger Lake laptops like my Razer Book have the same memory. It is not appreciably closer to the CPU in the Apple design. In Intel's Tiger Lake reference designs the memory is also in two chips that are mounted right next to the CPU.
> ./two_or_three
N = 1000000000, 953.7 MB
starting experiments.
two : 30.6 ns
two+ : 39.6 ns
three: 45.1 ns
bogus 1422321000
This is much slower than the times Lemire reported for the M1. `two+` is 62% of the way between `two` and and `three`, vs 88% for the M1.EDIT: Adding `-march=native` didn't really change the results, which makes sense given that it's a memory benchmark.
Re: Memory access on the Apple M1 processor
#228Earlier quoted context omitted.
Thanks for clarifying. But this isn't any fundamental difference IMO. There isn't any functional limitation in an Intel core that means it cannot saturate the memory bandwidth from a single core, unless I am missing something.
One could argue that it's not "fundamental", but it's definitely a functional limitation of the current Intel cores. The memory bandwidth of a single core is hardware limited by the number "Line Fill Buffers". Each buffer keeps track of one outstanding L1 cacheline miss, thus the number of LFB's limits the memory level parallelism (MLP). "Little's Law" gives the relationship between the latency, outstanding requests,…
Re: Memory access on the Apple M1 processor
#229Earlier quoted context omitted.
What's with downvotes? I know some people don't mind, but it's a deal breaker to me, and it something I don't want to support. I don't care how good their hardware is. Moreover, good luck sourcing parts if the device has trouble. Apple will not sell you the parts. Even if you wanted. A walled garden does not make their hardware any better. If anything it makes it worse. I hope for Mac users apply does not clamp down…
> What's with downvotes? Didn't downvote, but I think it's the same as when people read a "letter to the editor" of yore, declaring that some person "cancelled their subscription" because of something in the magazine. A natural response is "Don't let the door hit you on your way out", which on HN might be expressed through a downvote by some. > I don't care how good their hardware is. Moreover, good luck sourcing par…
I never bother to enter apple's tyrannical ecosystem . So there is no door to hit me on the way out.
> Well, they repair all kinds of parts, and have guarantees and guarantee extension programs.
You can not get a lot surface mount chips to repair a mac book or iPhone without having to look on the gray-market. This is even before possible firmware issue if you manage to find parts. Heck even getting full replacement boards is basically impossible, unless they are from donor machines that have other problems.
> "Well, it does in a few ways. Mandating how the software is made, and what software is sold, when it should adapt new libs to continue being sold, etc, means that they can move the platform in different ways faster."
I disagree, allowing people to side-load does not stop apple from having policies in place for it's app stores. That's honestly the only problem. It's the owners hardware, they should not need apples permission to run code on it. Unless the owner can sign software themselves and/or run it without apple's consent this will always be a problem. You can't even install gcc without jailbreaking an iPhone.
At the very least I should be able to install an other OS on the device like GNU/Linux. If apple does not want to open iOS the user should at least have that option for the hardware.
This is even before you get into how apple treats developers. Have you read the entire App store guidelines. Some of it is ridiculous. Some of the insanity prevents Firefox from even porting their own browser engine.