Earlier quoted context omitted.
Yes, ROCm can be used to run frontier models and is being used by OpenAI, Anthropic, and Meta.
I would prefer direct hardware kernel interface. Like linux DMABUFs with userland hardware command ring buffers (I guess this hardware ring buffer instance would be specific to a VMID and a PASID).
DeepSeek V4 Flash on a Single AMD MI300X
91–100 of 115 posts
Re: DeepSeek V4 Flash on a Single AMD MI300X
#92Re: DeepSeek V4 Flash on a Single AMD MI300X
#93Earlier quoted context omitted.
I thought it was a consumer grade GPU until I saw the 192GB of HBM and 256GB or RAM.
It's a chopped down MI350X (roughly half the performance)
Re: DeepSeek V4 Flash on a Single AMD MI300X
#94Earlier quoted context omitted.
Kimi-K3: 2.8T Qwen3.8-Max: 2.4T DeepSeek V4 Pro: 1.6T DeepSeek V4 Flash: 284B (all are total parameter counts, not active parameters)
Rumors say chatgpt/claude/gemini/etc are in the 100s of teras. True?
Re: DeepSeek V4 Flash on a Single AMD MI300X
#95Earlier quoted context omitted.
Can we see some actual numbers, projections, models instead of vibes?
Debt is at $3 trillion right now: https://fortune.com/2026/07/31/ai-debt-hypescalers-capex-cap... Interest alone, at assumed 5%, amounts to about $150 billions per year. That's probably higher than the combined AI revenue of the top 3 providers.
Now let's build the model out more. What is the projected revenue, backlog, improvements in existing big tech businesses such as AI helping Meta's ad business?
Re: DeepSeek V4 Flash on a Single AMD MI300X
#96Unfortunately, the MI300X is an OAM module. The MI350P is the one you want: It's a PCIe card, but it has less memory: 144GB. Luckily, DeepSeek V4 Flash will run in 144GB too because it's 256 MoE exports are native MXFP4 quantized.
Re: DeepSeek V4 Flash on a Single AMD MI300X
#97Earlier quoted context omitted.
> should not discount that DeepSeek also gets paid in data, which is probably more valuable to them That's agentic feedback loops for training, right? Any more detail on this, such as how they actually tell whether that data is good or not? That seems like a very hard problem, and like the value of that data is low compared to just building their own, controlled RL gyms.
Agents usually start with ingesting the existing code base, and DeepSeek can use those code bases for pretraining. And they will have filters on top of that to throw out garbage. I am not sure how they are using the data for post-training, but there probably are ways to get signal out of it, e.g. sentiment analysis when the user begins cursing at the agent, or checking whether the user continued another session with…
Thank you for elaborating, that's already useful. Anywhere I can learn more about this? I'm very interested in it!
Re: DeepSeek V4 Flash on a Single AMD MI300X
#98Re: DeepSeek V4 Flash on a Single AMD MI300X
#99Earlier quoted context omitted.
Debt is at $3 trillion right now: https://fortune.com/2026/07/31/ai-debt-hypescalers-capex-cap... Interest alone, at assumed 5%, amounts to about $150 billions per year. That's probably higher than the combined AI revenue of the top 3 providers.
Did you read your own article? It's well less than $3 trillion. Now let's build the model out more. What is the projected revenue, backlog, improvements in existing big tech businesses such as AI helping Meta's ad business?
Re: DeepSeek V4 Flash on a Single AMD MI300X
#100Earlier quoted context omitted.
Just put a fan on it. It's just 600W, so nothing super-special is needed. Or add a water cooler.
You're going to need at least a couple loud ass IPPC 3000s to usefully move that kind of heat if you don't want it to throttle. And then another normal sized fan for the doorway of the room it's in. Not exactly super special, but a ~constant 600W+ of heat tends to be a learning experience. It's worse than a high end gaming rig, much closer to a literal space heater. I don't work during the summer because it sucks fig…