Live data from Hacker News

DeepSeek v4.1 Flash

twitter.com

371–380 of 405 posts

Re: DeepSeek v4.1 Flash

#371
post #20

As I also said on Twitter - it really amazes me how fearless Deepseek are. Every single model release is packed with new and crazy clever ideas and somehow, they always commit to training them at near frontier scale. I know everybody wants the tell all story of the clever ideas that were developed over the last ~3 years at Anthropic and OpenAI, but what I really want to thumb through is DeepSeek's notebook of "brilli…

>it really amazes me how fearless Deepseek are

Reel it in a bit, man.

Re: DeepSeek v4.1 Flash

#372

Earlier quoted context omitted.

> At this point it's just marketing stunts. If you have access to a SOTA model without guardrails, provide a prompt that lets the agent come up with "creative" solutions to problems, and don't properly isolate it, they can end up inadvertently hacking 3rd party companies. Even if it was a mistake or "mistake", the part where the agent can exploit things across multiple levels like that, isn't just marketing. It seems…

But aren't there plenty of uncensored/unrestricted models out there? Where is all the collateral damage? Also, I think if Claude and OpenAI are just doing industry standard guardrails that everyone else is doing including DeepSeek, the fact that they are talking about it more than other companies makes it part of the marketing campaign. As an analogy, if Apple were to talk up their phones having fast charging but the…

I agree that fears are overblown. But we have definitely seen some attacks, especially in the crypto space. Three major ones just in the past month: Coldcard wallet, Liquid, and Trezor email compromised.

They're almost certainly a result of more competent models finding exploits.

Re: DeepSeek v4.1 Flash

#373

Earlier quoted context omitted.

The speed of GLM 5.3 Flash on OpenRouter seems to vary considerably by provider. Some are fast and some are slow. OpenRouter does provide some tuning knobs, but not enough for my taste. It’s also token-heavy with reasoning, though I found it better than Deepseek V4 Flash previously.

> though I found it better than Deepseek V4 Flash previously Same experience here. But man, switch to V4.1 now! It is much better. I don't event need to test it for long run and I believe it's crazy good. I call it "AI era model taste" when I judge the model by it's output without reading the bench scores.

I’ve just tried, it’s now my new favorite Flash model

Re: DeepSeek v4.1 Flash

#374
post #331

Seems just slightly better than last v4 release, considerably (3x) more expensive, but also faster and slightly more token efficient. https://aibenchy.com/compare/deepseek-deepseek-v4-1-flash-hi...

Ignoring the obvious ai slop webpage. I don't really trust the benchmark. It seems either pretty saturated, or inconsistent just based on the results.

Re: DeepSeek v4.1 Flash

#375
post #20

As I also said on Twitter - it really amazes me how fearless Deepseek are. Every single model release is packed with new and crazy clever ideas and somehow, they always commit to training them at near frontier scale. I know everybody wants the tell all story of the clever ideas that were developed over the last ~3 years at Anthropic and OpenAI, but what I really want to thumb through is DeepSeek's notebook of "brilli…

>it really amazes me how fearless Deepseek are Reel it in a bit, man.

The circle jerking of Chinese models on this site never ceases to amuse me.

Re: DeepSeek v4.1 Flash

#376
post #374
post #331

Seems just slightly better than last v4 release, considerably (3x) more expensive, but also faster and slightly more token efficient. https://aibenchy.com/compare/deepseek-deepseek-v4-1-flash-hi...

Ignoring the obvious ai slop webpage. I don't really trust the benchmark. It seems either pretty saturated, or inconsistent just based on the results.

What seems inconsistent?

The coverage is quite small, only 22 tests.

It's more to compare the cost/speed/consistency between models, given the same tasks.

Re: DeepSeek v4.1 Flash

#377
post #331

Seems just slightly better than last v4 release, considerably (3x) more expensive, but also faster and slightly more token efficient. https://aibenchy.com/compare/deepseek-deepseek-v4-1-flash-hi...

Your benchmark is a curious one. I didn't see you included Muse Spark 1.3 contributor even though its price is much lower even than DeepSeek. The low price changes many recommendations completely. And FWIW, DeepSeek retain and train on your data, too.

Re: DeepSeek v4.1 Flash

#378

OpenCode Go is currently running a 4x usage promo on DeepSeek v4.1 flash, not a bad way to get your feet wet (even if their cache hit prices are probably still very sub-optimal)

For small projects and hobby-programming, OpenCode Go is great, and its model performance is quite strong in my experience. Every time it's mentioned, there are people loudly claiming that it has terrible, quantized models, though this is never backed with data. I'm suspicious that this is being propagated by those whose financial interests are harmed by the existence of a cheap and decent coding subscription.

Re: DeepSeek v4.1 Flash

#379
post #163

Amazing Cyberbench scores. Holy shit. Too bad that DeepSeek AI went beyond 470b weights (which is a somewhat realistic limit for a 2x 128GB unified memory machine cluster like Strix Halo or Nvidia Spark). That means that to make the model fit into memory there you need a quantisation of lower than 4bits per weight (which is usually bad) to fit it into the available memory.

what would happen if you ran it off the SSD? Would it just wear it out or take weeks to run a simple prompt?

Re: DeepSeek v4.1 Flash

#380

Earlier quoted context omitted.

I was talking with a friend from the medical industry about it today. 30-50% of r&d spend in his sector is spent on safety, and for good reason. Proper trials, safety reviews and checkpoints and so on. Given the potential harm that could come from AI, we should probably be mandating something similar. Why wait to focus on safety until it’s too late.

In this case "safety" means how to restrict access to good models for working class. You can be sure the rich have access to unrestricted and uncensored models.

I really don't think this is true at all.

Do you have any evidence to suggest fully unrestricted frontier models are available for a price? Or...even exist?

Post reply on HN