Live data from Hacker News

DeepSeek V4 Flash on a Single AMD MI300X

github.com

91–100 of 115 posts

Re: DeepSeek V4 Flash on a Single AMD MI300X

#91
post #84
post #79

Earlier quoted context omitted.

Yes, ROCm can be used to run frontier models and is being used by OpenAI, Anthropic, and Meta.

I would prefer direct hardware kernel interface. Like linux DMABUFs with userland hardware command ring buffers (I guess this hardware ring buffer instance would be specific to a VMID and a PASID).

You can do that (tinygrad style) if you want.

Re: DeepSeek V4 Flash on a Single AMD MI300X

#92
post #82

Earlier quoted context omitted.

Kimi-K3: 2.8T Qwen3.8-Max: 2.4T DeepSeek V4 Pro: 1.6T DeepSeek V4 Flash: 284B (all are total parameter counts, not active parameters)

Rumors say chatgpt/claude/gemini/etc are in the 100s of teras. True?

No, rumors say 5-10T.

Re: DeepSeek V4 Flash on a Single AMD MI300X

#93
post #18

Earlier quoted context omitted.

I thought it was a consumer grade GPU until I saw the 192GB of HBM and 256GB or RAM.

It's a chopped down MI350X (roughly half the performance)

Basically right from Lisa Su's speech: "AMD is essentially taking one of its MI350X accelerators and cutting it in half, resulting in a card with half as many compute resources, half as much memory, and perhaps most importantly, a bit over half of the power consumption"

Re: DeepSeek V4 Flash on a Single AMD MI300X

#94
post #82

Earlier quoted context omitted.

Kimi-K3: 2.8T Qwen3.8-Max: 2.4T DeepSeek V4 Pro: 1.6T DeepSeek V4 Flash: 284B (all are total parameter counts, not active parameters)

Rumors say chatgpt/claude/gemini/etc are in the 100s of teras. True?

Maybe you could estimate frontier model sizes from AWS bedrock pricing?

Re: DeepSeek V4 Flash on a Single AMD MI300X

#95
post #53

Earlier quoted context omitted.

Can we see some actual numbers, projections, models instead of vibes?

Debt is at $3 trillion right now: https://fortune.com/2026/07/31/ai-debt-hypescalers-capex-cap... Interest alone, at assumed 5%, amounts to about $150 billions per year. That's probably higher than the combined AI revenue of the top 3 providers.

Did you read your own article? It's well less than $3 trillion.

Now let's build the model out more. What is the projected revenue, backlog, improvements in existing big tech businesses such as AI helping Meta's ad business?

Re: DeepSeek V4 Flash on a Single AMD MI300X

#96
post #35

Unfortunately, the MI300X is an OAM module. The MI350P is the one you want: It's a PCIe card, but it has less memory: 144GB. Luckily, DeepSeek V4 Flash will run in 144GB too because it's 256 MoE exports are native MXFP4 quantized.

Addendum: I was wrong, you‘ll need two of these cards.

Re: DeepSeek V4 Flash on a Single AMD MI300X

#97
post #57

Earlier quoted context omitted.

> should not discount that DeepSeek also gets paid in data, which is probably more valuable to them That's agentic feedback loops for training, right? Any more detail on this, such as how they actually tell whether that data is good or not? That seems like a very hard problem, and like the value of that data is low compared to just building their own, controlled RL gyms.

Agents usually start with ingesting the existing code base, and DeepSeek can use those code bases for pretraining. And they will have filters on top of that to throw out garbage. I am not sure how they are using the data for post-training, but there probably are ways to get signal out of it, e.g. sentiment analysis when the user begins cursing at the agent, or checking whether the user continued another session with…

> e.g. sentiment analysis when the user begins cursing at the agent, or checking whether the user continued another session with the generated code, or started a new session with the same starting point as before, i.e. they git-stashed.

Thank you for elaborating, that's already useful. Anywhere I can learn more about this? I'm very interested in it!

Re: DeepSeek V4 Flash on a Single AMD MI300X

#98

Earlier quoted context omitted.

Give it an AI-bubble pop and these will be flooding the market.

no they won't , the bubble is a financial thing. the demand is real and not going away.

Demand is money.

No bubble, where money?

Re: DeepSeek V4 Flash on a Single AMD MI300X

#99
post #53

Earlier quoted context omitted.

Debt is at $3 trillion right now: https://fortune.com/2026/07/31/ai-debt-hypescalers-capex-cap... Interest alone, at assumed 5%, amounts to about $150 billions per year. That's probably higher than the combined AI revenue of the top 3 providers.

Did you read your own article? It's well less than $3 trillion. Now let's build the model out more. What is the projected revenue, backlog, improvements in existing big tech businesses such as AI helping Meta's ad business?

Did you read beyond the headline? It's 1.65 trillion hidden and 1.35 trillion overt - which sums up to 3 trillion.

Re: DeepSeek V4 Flash on a Single AMD MI300X

#100
post #68

Earlier quoted context omitted.

Just put a fan on it. It's just 600W, so nothing super-special is needed. Or add a water cooler.

You're going to need at least a couple loud ass IPPC 3000s to usefully move that kind of heat if you don't want it to throttle. And then another normal sized fan for the doorway of the room it's in. Not exactly super special, but a ~constant 600W+ of heat tends to be a learning experience. It's worse than a high end gaming rig, much closer to a literal space heater. I don't work during the summer because it sucks fig…

I've got two 300W cards, that when fully running work at about 400W - 550W. It's like a space heater at a low setting, the room gets noticeably warmer.
Post reply on HN