Live data from Hacker News

A study on robustness and reliability of large language model code generation

arxiv.org

211–220 of 229 posts

Re: A study on robustness and reliability of large language model code generation

#211

Earlier quoted context omitted.

I would peg gpt as between junior and mid. Mids (around 5-10 years of experience) tend to produce overly complex code for the problem at hand.

Dude, chatgpt as a junior or mid level programmer? Have you actually met junior or mid level programmers? I’ve got 20 years of experience in the tech industry and chatgpt is far far better than any mid level engineer I’ve ever met and I include FAANG engineers too. Chatgpt can answer any leetcode programming question in essentially every programming language in existence in less than 20 seconds. Chatgpt can analyze c…

Can you share a link to the game you made with ChatGPT?

Re: A study on robustness and reliability of large language model code generation

#212
post #200

Earlier quoted context omitted.

I see humans do this all the time, especially abuse of exception handling, even among “Senior” developers. They don’t have a semantic understanding of what they are doing or why they are creating a race conditions or creating perverse control flow logic N layers down in the stack. The fact that researchers get it wrong is, well, unsurprising. LLMs might actually be an improvement.

What’s a “senior” developer these days? 3-4 years experience on average?

Knows how to build a feature or small system without handholding.

Re: A study on robustness and reliability of large language model code generation

#213
post #187

Earlier quoted context omitted.

>My believe is that there are exactly zero such jobs. If your claim is that we're at least at 1%, then you're claiming that there are at least 269k fully automatable programmer jobs [1]. It shouldn't be hard for you to find at least one to start. If I built an entire car but I am missing the key. I cannot drive the car but the car is 99% complete. It's just missing the key. So what happens is, it can replace 0% of hu…

[flagged]

[flagged]

Re: A study on robustness and reliability of large language model code generation

#214
post #186

Earlier quoted context omitted.

>You assume that getting from “48% as good” as a human programmer (on an undoubtedly biased sample) to “100% as good” is about equally hard as getting from 0% to 48%. That seems a pretty big assumption. I meant 38% aka 40% made a mistake on my math. It's not a big assumption. It's the most reasonable assumption. When you drive 50% of 10 miles the next 50% takes the same amount of time. It's the default assumption. I…

> It's not a big assumption. It's the most reasonable assumption. When you drive 50% of 10 miles the next 50% takes the same amount of time. It's the default assumption. And it's wrong, as anyone with experience in ML will tell you. 60% is easy. 70-85% is not too bad. 95% is hard. 100% is effectively impossible. This is the core problem of ML systems. They're good enough for 90% of cases but that last 10% can be incr…

The topic at hand is the pace of AI technology. This one is hard to quantify as it's been mostly linear in terms of technological progress. Example: From gans to LLMs. That progress has not been diminishing.

You seem to think the topic is about the current state of training algorithms and how it's hard to bring the model to 100 percent. That's a different topic with a subtle distinction.

Re: A study on robustness and reliability of large language model code generation

#215
post #191

Earlier quoted context omitted.

Your hypothesis is "our company will destroy the world" is a marketing move? Look. People in reality are basically never this gigabrained. Maybe consider as your first-line theory, that when people say "AI will destroy the world", that they mean to express that they believe that AI will destroy the world?

It's definitely a marketing strategy. "Our technology is so good it might destroy the world. So we're letting you use it for 19.99 a month." It's just a completely inconsistent position. Who benefits from the government saying only openAI and a few other companies can make this stuff? Would openAI rather talk about the (non-existent) existential threat of AGI or actual problems with their technology like data privacy…

Yeah, I think that's well said. I think the doomerism is more of a sell to investors than users sorta situation but it cuts both ways, which makes me sad for the tech industry having just made this same mistake with crypto. Who benefits from this? Anyone invested who doesn’t want to disappoint an investor (LPS, VCs, and founders) and/or anyone who’s ego is invested in the space and/or anyone who is raising money in the space and reliant on the technology and/or AI companies who benefit from the publicity/legitimacy of having their CEO talk to congress and/or AI companies who get an entire NYT piece run for a week about how an LLM is going to kill us all and/or any public company that can benefit from a stock rally by pushing the narrative that they’re in a land grab market and have the coveted goods and/or any large company that could benefit from blocking out nimbler competitors by getting arduous regulation passed… I could go on. You realize that like every 10th article on HN is an AI existentialism piece? Reddit is even worse. It also diverts the conversation from LLMs being shitty to the machines are coming. I’m fine with doomerism (getting hysterical can be fun) but I think it’s completely unwarranted with where we’re at. I don’t think we should overlook the risks associated with intellectual property, impersonation, data privacy, the ever growing body of bot spam destroying the internet etc. but worrying about the more dramatic risks has proven to be a pretty effective diversion tactic from the obvious acute problems — namely that LLMs have very little practical utility as they are now and the little utility they do have is primarily beneficial to spammers/scammers. I can already feel the person chiming in to tell me they love co-pilot (please just learn the language you’re working in).

But to be clear I’m not saying this is a methodical highly coordinated ad campaign amongst multiple companies. I don’t think human beings are that competent. I think it’s way more grassroots and feeds on cultural tropes that have existed since at least the 1950s but probably the early 1900s. I think Anthropic’s leadership probably genuinely believes their own bullshit but I also think they understand that that same bullshit has raised over a billion in funding.

Re: A study on robustness and reliability of large language model code generation

#216

Earlier quoted context omitted.

I would peg gpt as between junior and mid. Mids (around 5-10 years of experience) tend to produce overly complex code for the problem at hand.

Dude, chatgpt as a junior or mid level programmer? Have you actually met junior or mid level programmers? I’ve got 20 years of experience in the tech industry and chatgpt is far far better than any mid level engineer I’ve ever met and I include FAANG engineers too. Chatgpt can answer any leetcode programming question in essentially every programming language in existence in less than 20 seconds. Chatgpt can analyze c…

Yes, we employ them regularly. And yes, I use both ChatGPT and copilot regularly.

In my experience, if you compare the unedited output of ChatGPT to the unedited output of someone with a couple years of experience in the given domain, ChatGPT will have more subtle bugs than the human provided output.

Hence my between junior & mid assessment. It doesn't mean that it's useless though. Having a tool at that level of experience for every domain in the world that can cook up code near instantaneously is damn useful. And I assume it'll only get better.

Re: A study on robustness and reliability of large language model code generation

#217

Earlier quoted context omitted.

First off, you did not address most of my point. To answer your complaint, I am absolutely looking at reality. Please point us to an actual real-world professional programming job, matching the criteria I list above, for which a current LLM can economically replace a human for over a 5-year period. My believe is that there are exactly zero such jobs. If your claim is that we're at least at 1%, then you're claiming th…

>My believe is that there are exactly zero such jobs. If your claim is that we're at least at 1%, then you're claiming that there are at least 269k fully automatable programmer jobs [1]. It shouldn't be hard for you to find at least one to start. If I built an entire car but I am missing the key. I cannot drive the car but the car is 99% complete. It's just missing the key. So what happens is, it can replace 0% of hu…

Your "1%" is a fantasy measure. It is purely subjective. It creates a false sense of linearity when the progress curve is not only not linear, it may not be possible to complete. You just handwave away any criticism with a "clearly" when it isn't clear at all.

Mine is clearly measurable. You wanted to talk about job replacement; I'm measuring jobs replaced. You believe that we're at 48% complete (or 38% or whatever) but you can't point to even one programming job that ChatGPT can take over. If your fantasy 48% measure means 0% real-world success, I think the 0% is the more useful number to look at.

Humans may not be smart enough to build human-grade minds, in the same way that cats will never learn to program no matter how much they paw keyboards. People mistaking ChatGPT for AGI strikes me as the same kind of error where they mistook Eliza for a person, or 1970s/1980s computers as the things that would develop consciousness and maybe take over. It doesn't tell us much about the power of computers, but rather the inability of some humans to understand how complex and powerful minds are.

Re: A study on robustness and reliability of large language model code generation

#218
post #189

Earlier quoted context omitted.

> 48% if half way to 100% … It's halfway their to taking your job and your underwhelmed. You’re implying 100% reliability is actually attainable. If that were the case, wouldn’t that mean the halting problem would have been solved by AI? I’m not an expert but I’ve heard that’s like one of those fundamental laws of information theory that really can’t be broken. > It's typical. It's like an indie band is only popular…

I agree with the majority of your well considered comment. I'm okay with LLMs being really unreliable though. I think if they actually worked, it would be an unmitigated disaster for workers globally.

I think there’s arguments to be made both ways. I‘m generally of the belief that technology that works is a net win for humanity. But I think we need to stop being vague about what it means for an LLM to work. For me, I just want to tell an LLM to replace all the hard coded strings in my app with translation tags. Sadly, this isn’t possible to do reliably with what we have today. I’m not sure who’s job this would eliminate other than my own. I think the more existential questions are about if we made a god box that could perform all the white collar jobs. I think we’re so far away from that despite being sold that vision, that we might as well consider the moral and philosophical implications of a time machine or teleportation device if we’re going to continue to entertain AI doomerism for LLMs.

Re: A study on robustness and reliability of large language model code generation

#219

Earlier quoted context omitted.

This argument is not really valid anymore, especially in a paper that talks about LLMs.

It's an observation, not an argument, but note that this is a paper about LLM mistakes . But if the language bothers you and you view LLMs as the solution, you're certainly free to feed it into an LLM yourself. To be frank, I find it a little off-putting to suggest being a non-native English speaker is "no longer valid."

Just a grammar check or a proofreading pass through an LLM is all it needs. It's an English-language paper, so we e.g. separate words by spaces.

Hate this modern kneejerk reaction to get preachy and supercilious whenever anything vaguely cultural comes up.

Re: A study on robustness and reliability of large language model code generation

#220

While I do see LLMs helping programmers a lot, I'm not super impressed by seeing it write what looks like a lot of boilerplate. If things become so common that it can be replicated by an LLM, it seems like we need to be abstracting it away.

The worst programmers I know are the busiest. We’re already in this world. Seriously develop and app with Django and tell me how much actually “code” you write ? I’ve been a developer for almost 2 decades and the coding part of the job is about 5% of my time :) The rest is stopping people from building stupid shit they don’t need to build and trying to make sense of “the business” With generative AI, you can have all…

Yeah I agree. Webdev in particular could use a lot more work. I think serverless does a good job, but platform lock-in prevents it from becoming the standard. I think databases are a great example of abstraction done right. I rarely see database code(sql, configs) that seems copy-pasted.
Post reply on HN