Earlier quoted context omitted.
We'll likely see a transformation in how frontier models are trained as a result of a push towards local inference. While it seems unlikely now, given current pricing for RAM, in 10-15 years it's not unthinkable to assume we could see individual machines with 10-12TB (and well beyond that) of RAM which are accessible to the GPU. Min/max system RAM increased a LOT from 2010-2025 and largely because it was cheap. Once…
The 1080ti is out there for almost 10 years now. It has 11GB of VRAM. A 5090 has 32GB. SOCs with unified memory have shifted this a bit forward, but they're also expensive as shit. 10TB ram in a consumer device is simply not happening in the next 10 years.
Apple to skip high-end M6 Mac chips in favor of AI-focused M7 line
181–190 of 396 posts
Re: Apple to skip high-end M6 Mac chips in favor of AI-focused M7 line
#182The article says base M7 memory bandwidth is targeted at 240GB/s. M1 had 70 GB/s, M1 Pro: 200, M1 Max 400, M1 Ultra 800. Modern RTX 6000: ~1,600 or so. If we get a 1,200-1,500 GB/s bandwidth M7 variant in late 2027 with 512GB of RAM, that will be a very interesting chip. Tracking LLM size and performance improvements, I can imagine that being a sort of inflection point for local inference. I wonder what the power bud…
The article didn't state the M5 Ultra won't be released. It will probably provide 1228GB/s of memory bandwidth this year.
Re: Apple to skip high-end M6 Mac chips in favor of AI-focused M7 line
#183Earlier quoted context omitted.
> Though, I've been saying for a while that the local AI inflectiom point is the death knell for these frontier labs. "Death knell" is a touch hyperbolic. Hardware that can only run quantized models that take up GBs in VRAM falls short of even an A100 (by almost an order of magnitude[0]), which in turn falls short of what an 8xH100 cluster can do (also by another order of magnitude[0]). I'm an avid believer in local…
Is it hyperbolic though? One of the best things about the compute and memory shortage is that people are going to insane lengths to optimize things to run on lower memory / lower compute devices. If we keep this up for a while and then ramp up memory and local compute production, that AI inflection point may actually come. Of course, these are a lot of ifs.
Re: Apple to skip high-end M6 Mac chips in favor of AI-focused M7 line
#184Apple is actually interesting. They are one of the few companies with a chip / PC play with real power AND basically no play I'm the hyperscalar market. That means they're actually incentivized at least short term, to benefit PCs becoming strong enough to do local LLMs. Which makes this play make even more sense. Though, I've been saying for a while that the local AI inflectiom point is the death knell for these fron…
> Though, I've been saying for a while that the local AI inflectiom point is the death knell for these frontier labs. "Death knell" is a touch hyperbolic. Hardware that can only run quantized models that take up GBs in VRAM falls short of even an A100 (by almost an order of magnitude[0]), which in turn falls short of what an 8xH100 cluster can do (also by another order of magnitude[0]). I'm an avid believer in local…
So yeah, commercially it might be a death knell. Yes there's still a market for super computers, but would your rather own Apple or Cray?
Re: Apple to skip high-end M6 Mac chips in favor of AI-focused M7 line
#185Earlier quoted context omitted.
M series chips are system-on-chip with RAM on the same wafer with CPU and GPU, so it impossible to outsource only part of the chip
No it isn't, DRAM is made with a different process and those are chiplets, perfectly possible to outsource, and the only possibility really as TSMC does not make DRAM.
Re: Apple to skip high-end M6 Mac chips in favor of AI-focused M7 line
#186Earlier quoted context omitted.
Doesn’t need to be a winner head to head. If it can do 90% of the tasks the big boys do, at 50% speed, for virtually no extra overhead cost save for the power consumed by a prompt - that’s gonna work for a lot of people. And that’s also basically where we’re at today. Qwen3.6 35b running quantized on 10 year old hardware solves basically all of my uses cases for agents except for coding. The frontier models are faste…
> If it can do 90% of the tasks the big boys do, at 50% speed I want to live in this world too, but these numbers, as of today, are very aspirational and far removed from reality. I'm no tokenmaxxer; I find my modest local setup useful, I also know the limitations, it's slow and it sucks (relatively) at high-level and/or long-context planning, compared to frontier models. Only a minority of my prompts are max-effort…
The real question is, what are 90% of people going to ask llms to do. I’d argue mostly it’s going to be stuff that works-now or almost-works on local models, but that’s just an opinion. It also depends on the frontier models hitting a wall of steeply diminishing returns, since they set the expectations for all of this stuff - my gut says that’s happened already they just won’t admit it for a while - but we’ll see.
Re: Apple to skip high-end M6 Mac chips in favor of AI-focused M7 line
#187Apple is actually interesting. They are one of the few companies with a chip / PC play with real power AND basically no play I'm the hyperscalar market. That means they're actually incentivized at least short term, to benefit PCs becoming strong enough to do local LLMs. Which makes this play make even more sense. Though, I've been saying for a while that the local AI inflectiom point is the death knell for these fron…
> Though, I've been saying for a while that the local AI inflectiom point is the death knell for these frontier labs. "Death knell" is a touch hyperbolic. Hardware that can only run quantized models that take up GBs in VRAM falls short of even an A100 (by almost an order of magnitude[0]), which in turn falls short of what an 8xH100 cluster can do (also by another order of magnitude[0]). I'm an avid believer in local…
Re: Apple to skip high-end M6 Mac chips in favor of AI-focused M7 line
#188Earlier quoted context omitted.
> If it can do 90% of the tasks the big boys do, at 50% speed I want to live in this world too, but these numbers, as of today, are very aspirational and far removed from reality. I'm no tokenmaxxer; I find my modest local setup useful, I also know the limitations, it's slow and it sucks (relatively) at high-level and/or long-context planning, compared to frontier models. Only a minority of my prompts are max-effort…
Consider also that right now LLMs run slowly enough you can watch them think. I've seen a demo of an LLM running at an absurdly high speed and it reminds me of when I moved from a 2400 baud modem to a 14.4 - BBS screens that I could watch draw were all of a sudden nigh-interactive. Faster-than-realtime video generation is also coming, and will also continue to require huge hardware for a long while yet. I love local…
Re: Apple to skip high-end M6 Mac chips in favor of AI-focused M7 line
#189What's their backup plan if the AI world doesn't pan out? What if it turns out people want base compute capability and lots of RAM for filestore cache and programs? Maybe this strategy works, even in that world. Remember when we all thought (were told we thought) the world was heading to 3D views of our 2D lived experience like a solid Cube of GUI we could rotate around and live inside? Well Apple took the simple 2D…
It’s all fairly easy bets to make and correct.
Re: Apple to skip high-end M6 Mac chips in favor of AI-focused M7 line
#190Apple is actually interesting. They are one of the few companies with a chip / PC play with real power AND basically no play I'm the hyperscalar market. That means they're actually incentivized at least short term, to benefit PCs becoming strong enough to do local LLMs. Which makes this play make even more sense. Though, I've been saying for a while that the local AI inflectiom point is the death knell for these fron…
> Though, I've been saying for a while that the local AI inflectiom point is the death knell for these frontier labs. "Death knell" is a touch hyperbolic. Hardware that can only run quantized models that take up GBs in VRAM falls short of even an A100 (by almost an order of magnitude[0]), which in turn falls short of what an 8xH100 cluster can do (also by another order of magnitude[0]). I'm an avid believer in local…
Apple is the only player here where it would play into their natural hardware incentive to get you to pay more for better hardware. It would make sense for them to find a way to run LLM locally (eg, newer architectures that others here have pointed out).
Interesting times.