DeepSeek-V4-Flash Update
311–320 of 362 posts
Re: DeepSeek-V4-Flash Update
#312I wonder when the antirez/ds4 group will have an update to their high accuracy 2 bit quant. Although it's funny that I am thinking about that at all because I have a 2060 :P . My local inference is playing with Gemma 4 E2B and MiniCPM 5 1B.
Re: DeepSeek-V4-Flash Update
#313Earlier quoted context omitted.
He explicitly said he wants open weight models at his recent speech at an AI conference in Shanghai: https://news.ycombinator.com/item?id=48970449#48970784
The Chinese equivalent of the expression 'open weights' appears nowhere in this talk. You believed the Western press. Chinese AIs are not typically open. The most important by far is Bytedance.
Re: DeepSeek-V4-Flash Update
#314Re: DeepSeek-V4-Flash Update
#315Oh my goodness what an update. I need these weights. It's an incredible model for the size. The improved tool calling etc. should be able to make my harness way simpler. This runs at mega-speed on prosumer hardware (2x RTX Pro 6000).
Re: DeepSeek-V4-Flash Update
#316Can someone please explain how these models aren’t just fine tuned for benchmarks? I’m not plugged in to this space much but it seems like such an obvious problem…
They definitely are - but also people are using them pretty extensively for work. So ultimately you can't really fake "is it good". But there's no real measurements of that when a model is released, so we are stuck with benchmarks.
Re: DeepSeek-V4-Flash Update
#317Earlier quoted context omitted.
Sounds like it'll replace v4-flash, v4.1 would be nice to keep both available. On the other hand, it's nice to just get an improvement on anything that asks for "deepseek-v4-flash" without having to change the model string.
I think that's backwards. Anything that changes the performance of a model deserves a minor version bump. A new model has to be qualified before being pushed to production; but we don't get the choice here, just cross your fingers there are no regressions at all on all possible tasks the model might be asked to do.
Re: DeepSeek-V4-Flash Update
#318Earlier quoted context omitted.
Weird, I'm using 5.6 Sol through Chinese resellers and it reverse engineers stuff just fine
I think I’m missing something, what do you mean Chinese resellers? As I understand it, it’s difficult to even access OpenAI in China, how could they be reselling it? Do you mean something like openrouter but a Chinese version or something?
Re: DeepSeek-V4-Flash Update
#319Earlier quoted context omitted.
Weird, I'm using 5.6 Sol through Chinese resellers and it reverse engineers stuff just fine
Do you have any suggestions for resellers? My Gmail username is the same as my HN username if you're not comfortable posting that here. Thanks.