Live data from Hacker News

DeepSeek V4 Flash on a Single AMD MI300X

github.com

21–30 of 114 posts

Re: DeepSeek V4 Flash on a Single AMD MI300X

#21

Earlier quoted context omitted.

Give it an AI-bubble pop and these will be flooding the market.

When is it popping? Is the AI bubble in the room with us now?

Tomorrow? Next year? In 5 years? Nobody can say. But we do know that AI is overvalued, so it WILL pop.

Re: DeepSeek V4 Flash on a Single AMD MI300X

#22
post #15
post #7

Earlier quoted context omitted.

At 830tok/s * 1 hour that's almost 3M tokens which is just $0.54 worth of tokens at Deepseeks current output price.

830t/s is burst aggregate. ~500 is sustained and it's for 8 concurrent users. Meaning for $1.99/hour if you serve 8 users it's 8*$0.54, not just $0.54. You shouldn't rent one out if you're just serving it for yourself, but from a financial standpoint if you sell to users you can take a 100% margin.

Bro 500 Aggregate. so that's 500 * 60 * 60 = 1.8M output which is .5$ at best... Not including pre-fill and stuff.

This is not the real margins, even if you are selling to 8 users it's 90 tps per median stream. So assuming that .6-.7$

This is not even remotely worth it.

You need to 3x this tps(~1500 tps) to be worth it, and that's what most providers are doing, at 20-30 users at 50-60 tps with better optimized batch processing and kernels you can make some profit.

Re: DeepSeek V4 Flash on a Single AMD MI300X

#23

Earlier quoted context omitted.

Give it an AI-bubble pop and these will be flooding the market.

When is it popping? Is the AI bubble in the room with us now?

The AI bubble will pop when China gets access to EUV, so the earliest it could happen is 2030

Re: DeepSeek V4 Flash on a Single AMD MI300X

#24

Earlier quoted context omitted.

When is it popping? Is the AI bubble in the room with us now?

Tomorrow? Next year? In 5 years? Nobody can say. But we do know that AI is overvalued, so it WILL pop.

Well, if no body can say when it will pop, then can we really say it's a bubble and it's overvalued?

I can tell you it will pop in 10 years and when it pops, it will still be 20x bigger than in 2026. Does that even make any sense?

People said AI bubble will pop soon in 2024 and that it was overvalued. Turns out, many AI stocks 10x, 20x since 2024. Actual usage has gone exponential as well. Anthropic revenue went from $100m ARR at start of 2024 to $80b ARR today.

Re: DeepSeek V4 Flash on a Single AMD MI300X

#25
post #22
post #15

Earlier quoted context omitted.

830t/s is burst aggregate. ~500 is sustained and it's for 8 concurrent users. Meaning for $1.99/hour if you serve 8 users it's 8*$0.54, not just $0.54. You shouldn't rent one out if you're just serving it for yourself, but from a financial standpoint if you sell to users you can take a 100% margin.

Bro 500 Aggregate. so that's 500 * 60 * 60 = 1.8M output which is .5$ at best... Not including pre-fill and stuff. This is not the real margins, even if you are selling to 8 users it's 90 tps per median stream. So assuming that .6-.7$ This is not even remotely worth it. You need to 3x this tps(~1500 tps) to be worth it, and that's what most providers are doing, at 20-30 users at 50-60 tps with better optimized batch…

You are right, it's not 500x8 it's 90x8.

Re: DeepSeek V4 Flash on a Single AMD MI300X

#26
post #7

Earlier quoted context omitted.

At 830tok/s * 1 hour that's almost 3M tokens which is just $0.54 worth of tokens at Deepseeks current output price.

This is exactly what I came to say. The price of Flash is so cheap that trying to run it locally or with your own hardware is pointless. I was using it about a month ago to program some stuff and ran it for 4 days non-stop and it cost me about $2.

> trying to run it locally or with your own hardware is pointless.

Serving local models has advantages other than price. If you work in restricted industries, or have a strong need to protect your IP, or if you just value privacy more than cost, you now have options.

Re: DeepSeek V4 Flash on a Single AMD MI300X

#27
post #7

Earlier quoted context omitted.

At 830tok/s * 1 hour that's almost 3M tokens which is just $0.54 worth of tokens at Deepseeks current output price.

This is exactly what I came to say. The price of Flash is so cheap that trying to run it locally or with your own hardware is pointless. I was using it about a month ago to program some stuff and ran it for 4 days non-stop and it cost me about $2.

With the cost of electricity, hardware depreciation and tok/s it rarely makes sense to run locally.

Re: DeepSeek V4 Flash on a Single AMD MI300X

#28
This is still quite a bit away from the performance that deepseek gets on their H800. In their DSpark paper they report a throughput of 15k tokens/s/gpu. The MI300 should be able to compete with the H800 so there are probably still quite a few optimizations that can be made.

Re: DeepSeek V4 Flash on a Single AMD MI300X

#30

Earlier quoted context omitted.

Tomorrow? Next year? In 5 years? Nobody can say. But we do know that AI is overvalued, so it WILL pop.

Well, if no body can say when it will pop, then can we really say it's a bubble and it's overvalued? I can tell you it will pop in 10 years and when it pops, it will still be 20x bigger than in 2026. Does that even make any sense? People said AI bubble will pop soon in 2024 and that it was overvalued. Turns out, many AI stocks 10x, 20x since 2024. Actual usage has gone exponential as well. Anthropic revenue went from…

Many are saying July 2027, as in the past these market corrections have correlated with Shrek movie releases.

Debt-backed investors have to pay up eventually. =3

Post reply on HN