Live data from Hacker News

Promising results from DeepSeek R1 for code

simonwillison.net

571–580 of 765 posts

Re: Promising results from DeepSeek R1 for code

#571
post #416

Earlier quoted context omitted.

Maybe it would work better if it used an IDE rather than having to write flawless code without ever testing it?

I tried something related today with Claude, who'd messed up a certain visualization of entropies using JS: I snapped a phone photo and said 'behold'. The next try was a glitch mess, and I said hey, could you get your JS to capture the canvas as an image and then just look at the image yourself? Claude could indeed, and successfully debugged zir own code that way with no more guidance. This was all in the default web…

Holy shit.

Re: Promising results from DeepSeek R1 for code

#572
post #549

Earlier quoted context omitted.

I think what you're describing is going to be a very short transitional period. Once current AI gets good enough, the people micromanaging parts of it will do more to hinder the process than to help it. One person setting the objectives and the AI handling literally everything else including brainstorming issues etc, is going to be all that's needed.

What are you going to do when the output is wrong? You're not expecting it to always be right, are you?

I said this in another comment but look at the leading chess engines. They are already so far above human level of play that having a human override the engines choice will nearly always lead to a worse position.

> You're not expecting it to always be right, are you?

I think another thing that gets lost in these conversations is that humans already produce things that are "wrong". That's what bugs are. AI will also sometimes create things that have bugs and that's fine so long as they do so at a rate lower than human software developers.

We already don't expect humans to write absolutely perfect software so it's unreasonable to expect that AI will do so.

Re: Promising results from DeepSeek R1 for code

#573
post #389

Earlier quoted context omitted.

Why do people keep talking about this? We get it, Chinese models are censored by CCP law. Can we stop talking about it now? I swear this must be some sort of psyop at this point.

Mostly anti-Chinese bias from Americans, Western Europeans, and people aligned with that axis of power (e.g. Japan). However, on the Japanese internet, I don't see this obsession with taboo Chinese topics like on Hacker News. People on Hacker News will rave about 天安門事件 but they will never have heard of the South Korean equivalent (cf. 光州事件) which was supported by the United States government. I try to avoid discussin…

Great job using your voice for the voiceless.

> Nobody does that for GPT, Claude, etc

Flat out not true.

> companies will generally follow local laws

And people are doing the right thing by talking about it according to their local laws, and their own values, not those others have or may forced to abide by.

Re: Promising results from DeepSeek R1 for code

#574

> 99% of the code in this PR [for llama.cpp] is written by DeekSeek-R1 It's definitely possible for AI to do a large fraction of your coding, and for it to contribute significantly to "improving itself". As an example, aider currently writes about 70% of the new code in each of its releases. I automatically track and share this stat as graph [0] with aider's release notes. Before Sonnet, most releases were less than…

Can you make a plot like HISTORY but with axis changed? X: date Y: work leverage (i.e. 50%=2x, 90%=10x, 95%=20x, leverage = 1/(1-pct) )

Re: Promising results from DeepSeek R1 for code

#575
post #547

Earlier quoted context omitted.

Is that censorship or just the AI reflecting the training data? I feel like that answer is given because that is how people write about Palestine generally.

That’s irrelevant. The models are censored for “safety”. One man safety is another man censorship.

I think you are missing my point... I am saying the example wasn't censorship from the model, but were reflective of the source material.

You can argue the source material is censored, but that is still different than censoring the model

Re: Promising results from DeepSeek R1 for code

#576
post #549

Earlier quoted context omitted.

I think what you're describing is going to be a very short transitional period. Once current AI gets good enough, the people micromanaging parts of it will do more to hinder the process than to help it. One person setting the objectives and the AI handling literally everything else including brainstorming issues etc, is going to be all that's needed.

What are you going to do when the output is wrong? You're not expecting it to always be right, are you?

I don't expect any code to be right the first time. I would imagine if it's intelligent enough to ask the right questions, research, and write an implementation, it's intelligent enough to do some debugging.

Re: Promising results from DeepSeek R1 for code

#577

Earlier quoted context omitted.

This assumes that prompts do not evolve to the point where grandma can mutter some words to AI that produces an app that solves a problem. Prompts are an art form and a friction point to great results. Was only some months before reasoning models that CoT prompts where state of the art. Reasoning models take that friction away. Thinking it out even further, programming languages will likely go away altogether as ulti…

> programming languages will likely go away altogether As we know them, certainly. I haven't seen discussions about this (links welcome!), but I find it fascinating. What would a PL look like, if it was not designed to be written by humans, but instead be some kind of intermediate format generated by an AI for humans to review? It would need to be a kind of formal specification. There would be multiple levels of abst…

I find it similarly fascinating.

Take for example neuralink. If you consider that interface 10 years, or further 1000 years out in the future, it's likely we will have a direct, thought-based human computer interface. Which is interesting when thinking of this for sending information to the computer, but even more so (if equally alarming) for information flowing from computer to human. Whereas today, we read text on web pages, or listen to audio books, in that future, we may instead receive felt experiences / knowledge / wisdom.

Have you had a chance to read 'Metaman: The Merging of Humans and Machines into a Global Superorganism' from 1993?

Re: Promising results from DeepSeek R1 for code

#578
post #403

Earlier quoted context omitted.

AI doesn't have needs any desires, humans do. And no matter how hyped one might be about AI, we're far away from creating an artificial human. As long as that's true, AI is a tool to make humans more effective.

> AI doesn't have needs any desires, humans do. I fear that this won't age well. But to shamelessly riff on Marx, those who control the means of computation will control society.

In the AI age, those who own the problems stand to own the AI benefits. Utility is in the application layer, not the hosting or development of AI models.

Re: Promising results from DeepSeek R1 for code

#579
post #531

Earlier quoted context omitted.

While this may or may be not the reason of why it behaves like this, there's no doubt that ChatGPT (as well as any other model, released by a major company, open or not) undergoes a lot of censorship and will refuse to produce many types of (often harmless) content. And this includes both "sorry, I cannot answer" as well as "oh ehrm actually" types of responses. And, in fact, nobody makes a secret out of it, everyone…

At one point chatGPT censored me for asking: "What is a pannus?"

It's the handle of a frying pan, obviously.

Re: Promising results from DeepSeek R1 for code

#580

Earlier quoted context omitted.

> make the balance between capital and labor even more uneven. I think it's interesting to note that as opens source models evolve and proliferate, the capital required for a lot of ventures goes down - which levels the playing field. When I can talk to one agent-with-a-CAD-integration and have it design a gadget for me and ship the design off to a 3D printer and then have another agent write the code to run on the g…

What value do you bring to the venture, though? What makes your venture more likely to succeed than anybody else's, if the barrier is that low? I mean, I'll tell you: if anyone can spend $100 to design the same new gadget, the winner is going to be whoever can spend a million in production (to get economy of scale) and marketing. Currently, financial capital needs your brain, so you can leverage that. But if they can…

Since everyone has AI, then it stands that humans still make the difference. That is why I don't think companies will be able to automate software dev too much, they would be cutting the one advantage they could have over their competition.
Post reply on HN