Why don't they just release a basic GPU with 128GB RAM and eat NVidia's local generative AI lunch? The networking effect of all devs porting their LLMs etc. to that card would instantly put them as a major CUDA threat. But beancounters running the company would never get such an idea...
Judging by the number of 16 GB laptops I see around, 128 GB of RAM would probably cost a bajillion dollars
And Nvidia doesn't want to cannibalize its high end chips but putting more memory into consumer ones.