Live data from Hacker News

Bringing Up DeepSeek-V4-Flash on AMD MI300X

fergusfinn.com

21–29 of 29 posts

Re: Bringing Up DeepSeek-V4-Flash on AMD MI300X

#22
post #20

Earlier quoted context omitted.

yes,sir, any possibility to find 1000pcs or more

They are all over ebay. https://www.ebay.com/sch/i.html?_nkw=bc-250 I'm super curious what you would use them for.

now many guys want to buy this, I am reseller AMD BC-250,it is popular now

Re: Bringing Up DeepSeek-V4-Flash on AMD MI300X

#24
post #22

Earlier quoted context omitted.

They are all over ebay. https://www.ebay.com/sch/i.html?_nkw=bc-250 I'm super curious what you would use them for.

now many guys want to buy this, I am reseller AMD BC-250,it is popular now

thank you sir, actually it is not easy to find on ebay, because they are small seller, hard to find hundred, also ebay price a bit high for me

Re: Bringing Up DeepSeek-V4-Flash on AMD MI300X

#25
post #12

Checked out this company about a year ago and they only offered small models. Now I see they have GLM-fp8/Kimi and DeepSeek V4 Pro. Since workloads are predominantly cached input, I'm surprised to see no separate price for cached input vs uncached. I hope the prices will drop significantly; with these prices you'll end up with thousands in monthly costs quickly. Hopefully more hardware companies will be on the market…

Hi! Co-founder of Doubleword here - we've hugely increased the number of models that we offer (partly thanks to work that we've done on hotswapping https://blog.doubleword.ai/fast-sglang-starts.

We're kind of known for our low prices - our prices (our main usage is for our high throughput API - the async tier) is significantly below average openrouter prices - but cached prices is coming soon which will lower them even more :)

Re: Bringing Up DeepSeek-V4-Flash on AMD MI300X

#27
post #25
post #12

Checked out this company about a year ago and they only offered small models. Now I see they have GLM-fp8/Kimi and DeepSeek V4 Pro. Since workloads are predominantly cached input, I'm surprised to see no separate price for cached input vs uncached. I hope the prices will drop significantly; with these prices you'll end up with thousands in monthly costs quickly. Hopefully more hardware companies will be on the market…

Hi! Co-founder of Doubleword here - we've hugely increased the number of models that we offer (partly thanks to work that we've done on hotswapping https://blog.doubleword.ai/fast-sglang-starts . We're kind of known for our low prices - our prices (our main usage is for our high throughput API - the async tier) is significantly below average openrouter prices - but cached prices is coming soon which will lower them e…

What kind of workloads are you primarily seeing from users? I´d guess coding harness-type stuff where you have repeated calls with lots of cache hits. Or is it more like bulk OCR or invoice processing?

Re: Bringing Up DeepSeek-V4-Flash on AMD MI300X

#28

Earlier quoted context omitted.

I wish you guys could partner with Modular to get Mojo inference working on your hardware, e.g. https://www.modular.com/models/deepseek-v4-pro

Not sure I understand. If they support MI300x, their self-hosted will run on our hardware.

If it was that easy, I wouldn't have commented.

It's not, which is why it would be nice if they did the actual work (on your hardware).

I would 100% pay $16/hr to run a self-hosted instance, but I won't spend thousands of dollars to (maybe) get it working (my time + the hardware).

Re: Bringing Up DeepSeek-V4-Flash on AMD MI300X

#29

Earlier quoted context omitted.

Not sure I understand. If they support MI300x, their self-hosted will run on our hardware.

If it was that easy, I wouldn't have commented. It's not, which is why it would be nice if they did the actual work (on your hardware). I would 100% pay $16/hr to run a self-hosted instance, but I won't spend thousands of dollars to (maybe) get it working (my time + the hardware).

Ok, sure. Valid. Have you asked them to support V4?

https://docs.modular.com/max/models/

I agree with you though, serving up inference is secret sauce for a lot of teams and not everyone publishes how to do it because of the costs involved in doing so. They need an ROI.

Post reply on HN