Live data from Hacker News

DeepSeek v4.1 Flash

twitter.com

291–300 of 426 posts

Re: DeepSeek v4.1 Flash

#291

Earlier quoted context omitted.

Humans are not living creatures. They're just bipedal meat shells being operated by a 20W electrochemical computer running a suite of chemically signalled, electrically actuated modellable functions, much of which is wasted on homeostatic regulation of the meat shell, which is capable of incredible things, but it's still just a result of simple electrochemical functions like action potential generation, dendritic int…

It is amusing to see those trapped in extreme HAAD, the source and whole content of religious illusion, pretend that it is their opponents who are in a state of religious fantasia.

HAAD?

Re: DeepSeek v4.1 Flash

#292
post #27

Earlier quoted context omitted.

…are you sure a brave stance against safety and welfare is what we need in this moment? Why do you think your conception of the dangers are more accurate than all the scientists who have spent their lives studying this?

Model welfare is wishy washy bullshit. It's software, it doesn't have feelings. > Why do you think your conception of the dangers are more accurate than all the scientists who have spent their lives studying this? Do the Chinese have no such scientists?

Alas, the Chinese scientists have not read Harry Potter fanfiction, and thus their minds are inundated with cognitive biases

Re: DeepSeek v4.1 Flash

#293

I was talking with a friend from the medical industry about it today. 30-50% of r&d spend in his sector is spent on safety, and for good reason. Proper trials, safety reviews and checkpoints and so on. Given the potential harm that could come from AI, we should probably be mandating something similar. Why wait to focus on safety until it’s too late.

But here you are in the text generating industry, the worst that can happen is bad grade because AI will mess up John Keats with John Cleese or your React application will have bugs. Inconvenient, but mostly harmless.

Re: DeepSeek v4.1 Flash

#294

Earlier quoted context omitted.

> At this point it's just marketing stunts. If you have access to a SOTA model without guardrails, provide a prompt that lets the agent come up with "creative" solutions to problems, and don't properly isolate it, they can end up inadvertently hacking 3rd party companies. Even if it was a mistake or "mistake", the part where the agent can exploit things across multiple levels like that, isn't just marketing. It seems…

> Even if it was a mistake or "mistake", the part where the agent can exploit things across multiple levels like that, isn't just marketing. When we say "safety" people do not think we are protecting them from accidental automated crime at scale being committed on their behalf.

I'm fairly sure most "safety" people consider "large scale automated crime" part of the threat model, as the agents could accidentally fall into such a trap, if optimized for some misunderstood goal.

Re: DeepSeek v4.1 Flash

#295

Earlier quoted context omitted.

Humans are not living creatures. They're just bipedal meat shells being operated by a 20W electrochemical computer running a suite of chemically signalled, electrically actuated modellable functions, much of which is wasted on homeostatic regulation of the meat shell, which is capable of incredible things, but it's still just a result of simple electrochemical functions like action potential generation, dendritic int…

> What about when they watch adult video in VR, or have waifus? Humans engage in voluntary suspension of disbelief for pleasure and recreation all the time. Categorically different. People have killed themselves or others due to conversations they had with LLMs, but those are just the extreme cases. Most schizophrenics don't commit suicide or kill others, they are mentally ill nonetheless.

>People have killed themselves or others due to conversations they had with LLMs, but those are just the extreme cases. Most schizophrenics don't commit suicide or kill others, they are mentally ill nonetheless.

Do you think the kind of person who was already psychologically unhinged enough to kill themselves or another person because a chatbot told them to would be completely harmless and totally safe if only chatbots had never been invented? Or is it possible that close to all of the risk posed by this person comes from the person's mental illness, and not the pixels on the screen they're looking at?

Chatbots don't make people murderous any more than "satanic music", violent video games, or cannabis do. LLMs are just the latest entry on a list of hysterical moral panics that conflate coincidence with causation.

Re: DeepSeek v4.1 Flash

#297
post #287

I'm surprised more people aren't talking about the cache hit price: $0.003 per million tokens. I have a feeling that the price of 1 million tokens transmitted over the internet is more expensive than cache hit. Are we close to making the chat completion API obsolete because the cost of context transfer over network is going to dominate the task total cost? Here's the same token usage priced at different rates: a real…

$0.003 off-peak, not 0.003 cents.

Re: DeepSeek v4.1 Flash

#298
post #287

I'm surprised more people aren't talking about the cache hit price: $0.003 per million tokens. I have a feeling that the price of 1 million tokens transmitted over the internet is more expensive than cache hit. Are we close to making the chat completion API obsolete because the cost of context transfer over network is going to dominate the task total cost? Here's the same token usage priced at different rates: a real…

[deleted]

Re: DeepSeek v4.1 Flash

#299
post #161

I've run some evals on my puzzle game https://redactle.net/llm-leaderboard Deepseek v4.1 flash is able to solve it some of the time. I've found it burns through more reasoning tokens than any other model. Google models like Gemini 3.8 Flash are still dominating and is able to one-shot most evals while being the cheapest. I'm curious what other unique evals people are running.

Does it move the needle on high reasoning?

Re: DeepSeek v4.1 Flash

#300
post #27

It's so refreshing to see DeepSeek's tech report[1] full of juicy details; meanwhile, something like Fable's system card[2] is like 70% "safety", 10% "model welfare" to make sure little Claude isn't distressed, and 20% benchmark numbers. [1]: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/... [2]: https://www.anthropic.com/claude-fable-5-1-mythos-5-1-system...

…are you sure a brave stance against safety and welfare is what we need in this moment? Why do you think your conception of the dangers are more accurate than all the scientists who have spent their lives studying this?

> …are you sure a brave stance against safety and welfare is what we need in this moment?

Is it out of convenience to not see the hypocrisy? "Safety and welfare" for you and me. Yet if you work at Anthropic or OAI, or are a partner of them then you can let it rip!

Oh, and when they illegally do just that - you get a "we're sorry bro" blog post that's designed to drum up FOMO and, most importantly, zero accountability. Yet, if anyone else abuses a model in that same manner? Illegal! You're defending a very slippery slope here.

Also, who do you think trained these models to have these capabilities? It sure as shit wasn't content that OAI or Anthropic had by default. Why should I trust them with these skills when they "have not spent their lives studying this"?

Maybe start looking around before it's being used against you [0].

[0] https://www.gadgetreview.com/anthropic-is-building-ai-to-pre...

Post reply on HN