OpenCode Go is currently running a 4x usage promo on DeepSeek v4.1 flash, not a bad way to get your feet wet (even if their cache hit prices are probably still very sub-optimal)
I wouldn't trust OpenCode Go with my lunch after the shit-sandwich they served us with horrible V4 quants and 0 transparency. Then blaming it on their partners.
DeepSeek v4.1 Flash
421–430 of 497 posts
Re: DeepSeek v4.1 Flash
#422It's so refreshing to see DeepSeek's tech report[1] full of juicy details; meanwhile, something like Fable's system card[2] is like 70% "safety", 10% "model welfare" to make sure little Claude isn't distressed, and 20% benchmark numbers. [1]: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/... [2]: https://www.anthropic.com/claude-fable-5-1-mythos-5-1-system...
Re: DeepSeek v4.1 Flash
#423Got some fun if slightly janky looking pelicans out of this one: https://tools.simonwillison.net/markdown-svg-renderer?url=ht... I ran it on all seven reasoning levels supported by OpenRouter, but the reasoning token counts suggest to me that it doesn't actually support seven different levels. This is one of my biggest problems with OpenRouter - their abstraction layer makes reasoning levels harder to reason about. r…
If you're going to use OpenRouter to test reasoning levels, always make sure you are locking to the official provider instead of third party providers.
Re: DeepSeek v4.1 Flash
#424Tune has changed finally, but damn, for HN , embarrassingly slow, has to be said
Re: DeepSeek v4.1 Flash
#425It's so refreshing to see DeepSeek's tech report[1] full of juicy details; meanwhile, something like Fable's system card[2] is like 70% "safety", 10% "model welfare" to make sure little Claude isn't distressed, and 20% benchmark numbers. [1]: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/... [2]: https://www.anthropic.com/claude-fable-5-1-mythos-5-1-system...
When Chinese tech report is tech report and US tech report is bible scripture.
Re: DeepSeek v4.1 Flash
#426Quite a flex calling their GPT-6 competitor "Flash"! But it is faster than their last flash model due to a combination of architectural innovations including engrams and a new encoder/decoder design that uses 8B parameters for prefill and 16B for generation.
This is definitely not on par with GPT-6 astra. Not with GPT-5.6 sol either. But probably will set as a new baseline for modern API based LLM because it's so cheap.
Re: DeepSeek v4.1 Flash
#427Got some fun if slightly janky looking pelicans out of this one: https://tools.simonwillison.net/markdown-svg-renderer?url=ht... I ran it on all seven reasoning levels supported by OpenRouter, but the reasoning token counts suggest to me that it doesn't actually support seven different levels. This is one of my biggest problems with OpenRouter - their abstraction layer makes reasoning levels harder to reason about. r…
Some OpenRouter providers do not implement reasoning levels for these models correctly at all: https://www.reddit.com/r/DeepSeek/comments/1vdqjwr/openroute... If you're going to use OpenRouter to test reasoning levels, always make sure you are locking to the official provider instead of third party providers.
Re: DeepSeek v4.1 Flash
#428Earlier quoted context omitted.
I always talk to models using grugspeak, like 'where getcontext used' I felt a bit bad about it, then I learned yday that model's internal thinking traces are also like this
words like "is" "the" etc are filler words anyway. they won't be changing the meaning that much. I asked AI whether it hurts to read ill formed sentence as it does to a human. It replied, "it doesn't"
Re: DeepSeek v4.1 Flash
#429It's so refreshing to see DeepSeek's tech report[1] full of juicy details; meanwhile, something like Fable's system card[2] is like 70% "safety", 10% "model welfare" to make sure little Claude isn't distressed, and 20% benchmark numbers. [1]: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/... [2]: https://www.anthropic.com/claude-fable-5-1-mythos-5-1-system...
When Chinese tech report is tech report and US tech report is bible scripture.
Re: DeepSeek v4.1 Flash
#430Earlier quoted context omitted.
$0.003 off-peak, not 0.003 cents.
Yep, but even 0.006 is quite a big improvement. I'm curious now to test the model on some token-heavy tasks, like code exploration before a coding session, to see whether it will decrease the total cost of the task in the end or not.