Live data from Hacker News

DeepSeek V4 Flash 0731

arcprize.org

431–440 of 474 posts

Re: DeepSeek V4 Flash 0731

#431
post #318
post #154

Earlier quoted context omitted.

It won't be long before I can just stay home, and have my robot ride my bike for me.

I'm going to send mine to visit my mom. It's so hot in August.

You reminded me of a short story published over a century ago.

https://americanliterature.com/author/em-forster/novella/the...

Re: DeepSeek V4 Flash 0731

#432

Earlier quoted context omitted.

I think you're overlooking the fact that for long-horizon tasks, even small errors compound over time and can lead to catastrophic outcomes. For simple queries, we have reached the threshold since the beginning of the year, and models are good enough from every provider to make a meaningful difference between one another. (ChatGPT, Claude, Gemini, Grok, MuseSpark, Kimi, DeepSeek, GLM...) The real unlock will be, and…

I have a silly (but honest) question. What's an example or two of a > 24hr task that people are actually asking something to do? Like real life ones.

I wondered this too. Also, the duration of the task depends on the quality of the prompt and the model used. I have a feeling a lot of these day-long tasks are bogged down by suboptimal tool use and on-demand python slop.

Re: DeepSeek V4 Flash 0731

#433
post #292

Seeing everyone spend like 200USD a month seems kind of mad. I have £20/month Gemini and £20 a month claude for a bunch of personal projects. Yes I have to wait sometimes, it's probably a good thing.

I max out my $200/month Claude plan. You are obviously just not taking advantage of it to the same level as others. Which is fine. Don't pay for something you don't need. But I would definitely take a massive productivity hit if I had 1/20th the usage.

Same here and my codex plan. I have to wait until next week to get back to work.

Re: DeepSeek V4 Flash 0731

#434

Earlier quoted context omitted.

Why? It's open weight, there are plenty providers on open router that are serving the latest v4 flash at 0.14/0.28 $.

This would be more convincing if those providers had converged on a number that was not the exact pricing of DeepSeek themselves. Clearly DeepSeek is setting the price here and without them holding it down I expect increases.

Before this release, there were providers undercutting deepseek on the preview version.

I suspect that once the hype dies down, or the field gets more competitive we will see the same on 0731.

Re: DeepSeek V4 Flash 0731

#435

Earlier quoted context omitted.

I've reverse engineered multiple classic games and turned them into popular, browser-based MMO-like experiences. I'm also creating a free platform that replaces extremely out-of-date software, some of it only available with mutli-million dollar contracts, to help medical physics professionals with cutting-edge radiotherapy devices used to treat cancer. https://brynnbateman.com/ for a list of projects

Nice! I tried to play https://eternalsagas.com/play but it stalls at 45% loading with the progress bar always at 0 but the music playing. And out of curiosity, how do you automate testing the porting in the browser that's actually playable etc? And aren't you a bit scared of hosting and serving the "hairy bits" such as full assets? Nice job anyway!

Yeah that's my weakest browser game since I've actually transitioned to using a desktop Godot client I haven't released yet. It will load eventually it just takes forever bc the assets are quite large, which is part of the reason I've had to move to Godot. That project doesn't exactly have an active user base which is why it's lower on priority list to fix.

I automate testing playable parts in browser by adding a dev mode that allows text commands for everything instead of having to rely on clicking UI or 3D elements. It can see the full game state in JSON and interact in any way via commands.

And re: IP - I just accept that I might get a C&D any day and have to take it all down. I'm careful to not accept a single penny for any reason and don't even have Patreon. Usually monetizing is what makes IP owners unhappy. And for the Pokemon MMO I just don't advertise it anywhere meaningful since they'll C&D the second they see it. Largely made it for my nephew and we play it together.

Re: DeepSeek V4 Flash 0731

#436

Earlier quoted context omitted.

Majority of white collar work absolutely does not require sota models

White collar work will require SOTA models up until the point where said models can automate white collar work entirely. Then, we might finally see a meaningful commoditization crunch.

I genuinely don't know any work that has been taken over completely by AI without humans in the loop managing things

At this point, I don't even know if its possible

Re: DeepSeek V4 Flash 0731

#437

Earlier quoted context omitted.

I've reverse engineered multiple classic games and turned them into popular, browser-based MMO-like experiences. I'm also creating a free platform that replaces extremely out-of-date software, some of it only available with mutli-million dollar contracts, to help medical physics professionals with cutting-edge radiotherapy devices used to treat cancer. https://brynnbateman.com/ for a list of projects

Time to make creative software for linux that can replace adobe for Video/photo editing :P

GIMP isn't bad these days, I use it constantly

Re: DeepSeek V4 Flash 0731

#438
Did anyone else experience a change in verbosity? I've been playing around with an agent that holds your hand in a Jupiter notebook and it felt like it started writing essays versus nice, concise, helpful paragraphs like before. My gut was correct because I checked my Deepinfra usage and it was almost a 2x out-token usage for every in-token.

Not a huge deal since it's still cents per session, but my bigger issue was the weird change in tone. It became a lot more pretentious and over-explanatory.

Heavy prompt reworking helped but maybe that's just the cost of being better at coding and ARC-AGI?

Re: DeepSeek V4 Flash 0731

#439
post #162

I've been using it extensively since the release and the best summary I can give is that it's good enough to use it for (almost) everything and cheap enough that the cost are irrelevant. I'm running it in Oh My Pi with a second instance running as "advisor" and even with 5-6 active sessions (effectively 12 streams) I'm struggling to spend more than 5 bucks per day. OpenCode Go even has double limits temporarily so fo…

If what you're saying is true and accurate, then US-based AI labs are in big trouble. The only saving grace might be some sort of a 'national security' proclamation banning the use of state-of-the-art Chinese (and non-US) models across US federal and state governments and large enterprises (especially ones with federal government contracts), but even still, US AI labs will probably lose out massively on international…

Sample of 1…

I downgraded my Claude subscription and delegated my Claude Opus access to serve the role of an Architect to brainstorm and plan every step of development.

I leave the development to Deepseek.

Claude gets to review at many layers. It is often just as good as if I let Opus develop it by itself(the architect session will find similar number/level of gaps).

Deepseek flash as an architect and problem solver is not as thorough as Opus5 + high. Codex sol+ high is even better than Opus 5 at this moment for my needs.

Re: DeepSeek V4 Flash 0731

#440

Earlier quoted context omitted.

I think you're overlooking the fact that for long-horizon tasks, even small errors compound over time and can lead to catastrophic outcomes. For simple queries, we have reached the threshold since the beginning of the year, and models are good enough from every provider to make a meaningful difference between one another. (ChatGPT, Claude, Gemini, Grok, MuseSpark, Kimi, DeepSeek, GLM...) The real unlock will be, and…

That doesn’t make sense. It’s not like SOTA models are error free, yet we still use them. You use Fable 5 right? If that’s good enough for you now, why wouldn’t a Chinese model that’s as good as Fable 5 but at 10% the cost be good enough in 6 months?

"It's only real AI if it comes from the silicon region of USA"
Post reply on HN