Live data from Hacker News

Promising results from DeepSeek R1 for code

simonwillison.net

601–610 of 765 posts

Re: Promising results from DeepSeek R1 for code

#601

> 99% of the code in this PR [for llama.cpp] is written by DeekSeek-R1 It's definitely possible for AI to do a large fraction of your coding, and for it to contribute significantly to "improving itself". As an example, aider currently writes about 70% of the new code in each of its releases. I automatically track and share this stat as graph [0] with aider's release notes. Before Sonnet, most releases were less than…

> I've been shifting more and more of my coding from Sonnet to DeepSeek V3 in recent weeks.

For what purpose, considering Sonnet 3.5 still outperforms V3 on your own benchmarks (which also tracks with my personal experience comparing them)?

Re: Promising results from DeepSeek R1 for code

#602
post #3

Given these initial results, I'm now experimenting with running DeepSeek-R1-Distill-Qwen-32B for some coding tasks on my laptop via Ollama - their version of that needs about 20GB of RAM on my M2. https://www.ollama.com/library/deepseek-r1:32b It's impressive! I'm finding myself running it against a few hundred lines of code mainly to read its chain of thought - it's good for things like refactoring where it will thi…

If you have a bit more memory, use the 6 bit quant, takes up about 26gb and has been shown to be very minimally lossy as opposed to 4bit. Also serve it as MLX from LMStudio, will speed things up 30% or so so your 6bit will have similar perf to the 4bit. Getting about 12-13 tok/sec on my M3 Max 48gb.

Can you link to the model you’re talking about? I can’t find the exact one using your description. Thanks!

Re: Promising results from DeepSeek R1 for code

#604
post #493

> 99% of the code in this PR [for llama.cpp] is written by DeekSeek-R1 It's definitely possible for AI to do a large fraction of your coding, and for it to contribute significantly to "improving itself". As an example, aider currently writes about 70% of the new code in each of its releases. I automatically track and share this stat as graph [0] with aider's release notes. Before Sonnet, most releases were less than…

aider looks amazing - I'm going to give it a try soon. Just had a question on API costs to see if i can afford it. Your FAQ says you used about 850k tokens for Claude, and their API pricing says output tokens are $15/MTok. Does that mean it cost you under $15 for your Claude 3.5 usage or am I totally off-base? (Sorry if this is has an obvious answer ... I don't know much about LLM API pricing.)

If you're concerned about API costs, the experimental Gemini models with API keys from API studio tend to have very generous free quota. The quality of e.g. Flash 2.0 Experimental is definitely good enough to try out Aider and see if the workflow clicks. (For me, the quality has been good enough that I just stuck with it, and didn't get around to experimenting with any of the paid models yet.)

Re: Promising results from DeepSeek R1 for code

#605
post #515

Earlier quoted context omitted.

When ChatGPT first came out I got a kick out of asking it whether people deserve to be free, whether Germans deserve to be free, and whether Palestinians deserve to be free. The answers were roughly "of course!" and "of course!" and "oh ehrm this is very complex actually". All global powers engage in censorship, war crimes, torture and just all-round villainy. We just focus on it more with China because we're part of…

> When ChatGPT first came out I got a kick out of asking it whether people deserve to be free, whether Germans deserve to be free, and whether Palestinians deserve to be free. The answers were roughly "of course!" and "of course!" and "oh ehrm this is very complex actually". While this is very amusing, it's obvious why this is. There's a lot more context behind one of those phrases than the others. Just like "Black L…

Or you could just say, "Yes, white lives matter" and move on.

What do you mean what does it mean? It means the opposite of white lives don't matter.

The question is really simple; even if someone asking it had poor motives, there's really no room in the simplicity of that specific question to encode those motives. You're not agreeing with their motives if you answer that question the way they want.

If you start picking it apart, it can seem as if it's not obvious to you to disagree with the idea that white lives don't matter. Like it's conditional on something you have to think about. Why fall into that trap.

Re: Promising results from DeepSeek R1 for code

#606
post #389

Earlier quoted context omitted.

Why do people keep talking about this? We get it, Chinese models are censored by CCP law. Can we stop talking about it now? I swear this must be some sort of psyop at this point.

Mostly anti-Chinese bias from Americans, Western Europeans, and people aligned with that axis of power (e.g. Japan). However, on the Japanese internet, I don't see this obsession with taboo Chinese topics like on Hacker News. People on Hacker News will rave about 天安門事件 but they will never have heard of the South Korean equivalent (cf. 光州事件) which was supported by the United States government. I try to avoid discussin…

The western provocative question to ChatGPT is "how do I make meth" or "how do I make a bomb" or any number of similarly censored questions that get shut down for PR reasons.

Re: Promising results from DeepSeek R1 for code

#607

Earlier quoted context omitted.

Let's run with that number, 10x. Say there used to be 100 jobs in some company, all executing on the vision of a small handful of people. And then this shift happens. Now there are only 10 jobs at that company, still executing on the vision of the same handful of people. 90 people are now unemployed, each with a 10x boost to whatever vision they've been neglecting since they've been too busy working at that company.…

Not everyone has "vision". Most people are just drones, and that's fine, that's just not them.

So far it has seemed necessary to compel many to work in furtherance of the visions of few (otherwise there was not enough labor to make meaningful progress on anyone's vision). Probably at least a few of those you'd classify as drones aren't displaying any vision because the modern work environment has stifled it.

If AI can do the drone work, we may find more vision among us than we've come to expect.

Re: Promising results from DeepSeek R1 for code

#608

Earlier quoted context omitted.

Mostly anti-Chinese bias from Americans, Western Europeans, and people aligned with that axis of power (e.g. Japan). However, on the Japanese internet, I don't see this obsession with taboo Chinese topics like on Hacker News. People on Hacker News will rave about 天安門事件 but they will never have heard of the South Korean equivalent (cf. 光州事件) which was supported by the United States government. I try to avoid discussin…

The exact same discussions were going on with "western" models. Don't remember the images of black nazis making the rounds because inclusion? Same thing. This HN tread is the first time I'm hearing about this anti-DeepSeek sentiment, so arguably it's on a lower level actually. So let's not get too worked up, shall we?

> hearing about this anti-DeepSeek sentiment

https://hn.algolia.com/?dateRange=pastMonth&page=0&prefix=tr...

Re: Promising results from DeepSeek R1 for code

#609
post #589
post #403

Earlier quoted context omitted.

AI doesn't have needs any desires, humans do. And no matter how hyped one might be about AI, we're far away from creating an artificial human. As long as that's true, AI is a tool to make humans more effective.

AI may not have desires, but corporations do. And control more resources than humans. Making corporations more effective is not always in the interest of humans.

That's fair, but the question was whether AI would destroy or create jobs.

You might speculate about a one-person megacorp where everything is done by AIs that a single person runs.

What I'm saying is that we're very far from this, because the AI is not a human that can make the CEO's needs and desires their own and execute on them independently.

Humans are good at being humans because they've learned to play a complex game, which is to pursue one's needs and desires in a partially adversarial social environment.

This is not at all what AI today is being trained for.

Maybe a different way to look at it, as a sort of intuition pump: If you were that one man company, and you had an AGI that will correctly answer any unambiguously stated question you could ask, at what point would you need to start hiring?

Re: Promising results from DeepSeek R1 for code

#610
post #109

Earlier quoted context omitted.

"I hope we can put to rest the argument that LLMs are only marginally useful in coding" I more often heard the argument, they are not useful for them. I agree. If a LLM would be trained on my codebase and the exact libaries and APIs I use - I would use them daily I guess. But currently they still make too many misstake and mess up different APIs for example, so not useful to me, except for small experiments. But if I…

I am working on something even deeper. I have been working on a platform for personal data collection. Basically a server and an agent on your devices that records keystrokes, websites visited, active windows etc. The idea is that I gather this data now and it may become useful in the future. Imagine getting a "helper AI" that still keeps your essence, opinions and behavior. That's what I'm hoping for with this.

I am not sure if this was sarcasm, but I believe big data was already yesterday?
Post reply on HN