Live data from Hacker News

A study on robustness and reliability of large language model code generation

arxiv.org

171–180 of 229 posts

Re: A study on robustness and reliability of large language model code generation

#171

Earlier quoted context omitted.

This argument is not really valid anymore, especially in a paper that talks about LLMs.

It's an observation, not an argument, but note that this is a paper about LLM mistakes . But if the language bothers you and you view LLMs as the solution, you're certainly free to feed it into an LLM yourself. To be frank, I find it a little off-putting to suggest being a non-native English speaker is "no longer valid."

I'm not a native english speaker. For serious work, that is, anything that may be printed or published with my real name on it, I use grammar checks. That's the least I can do.

Re: A study on robustness and reliability of large language model code generation

#172

Earlier quoted context omitted.

I don't see how it added any extra information to the parent comment

They asked whether this paper was worth reading, because the English wasn't conventional in their dialect. I'm pointing out that that isn't a good way to screen papers, because most English speakers don't share the same dialect. Perhaps unstated or understated is the implication that, were you to discard papers in this way, you're liable to miss out. I'll leave it to you whether that added information. That wasn't re…

> I'm pointing out that that isn't a good way to screen papers, because most English speakers don't share the same dialect.

You did not make that clear at all.

> Perhaps unstated or understated is the implication that, were you to discard papers in this way, you're liable to miss out.

Indeed, that implication was totally missing.

> Feel free to downvote my comment if you don't feel it was appropriate or informative or otherwise didn't uphold whatever criteria you're judging it by.

I would certainly do so. Like their paper, it is quite poor, and not helpful for this forum. Unfortunately, /u/dang has removed my ability to downvote.

Re: A study on robustness and reliability of large language model code generation

#173

Earlier quoted context omitted.

They asked whether this paper was worth reading, because the English wasn't conventional in their dialect. I'm pointing out that that isn't a good way to screen papers, because most English speakers don't share the same dialect. Perhaps unstated or understated is the implication that, were you to discard papers in this way, you're liable to miss out. I'll leave it to you whether that added information. That wasn't re…

> I'm pointing out that that isn't a good way to screen papers, because most English speakers don't share the same dialect. You did not make that clear at all. > Perhaps unstated or understated is the implication that, were you to discard papers in this way, you're liable to miss out. Indeed, that implication was totally missing. > Feel free to downvote my comment if you don't feel it was appropriate or informative o…

Alrighty. Thanks for the feedback. Have a great day.

(For what it's worth, I often feel my comments are too long and wordy, and occasionally I've gotten feedback to that effect, so I was trying to be brief in my original comment.)

Re: A study on robustness and reliability of large language model code generation

#174

I'm sympathetic to the viewpoint that GPT-4 is prone to mistakes when writing code. Unfortunately, the analysis in this paper is pretty bad and doesn't support that conclusion. The authors assume that for any given method under consideration, it must only occur within a particular pattern of other method calls and control flow instructions. But the templates they have chosen are clearly only applicable in certain sit…

I see humans do this all the time, especially abuse of exception handling, even among “Senior” developers. They don’t have a semantic understanding of what they are doing or why they are creating a race conditions or creating perverse control flow logic N layers down in the stack. The fact that researchers get it wrong is, well, unsurprising. LLMs might actually be an improvement.

Hi, do you have a recommended read for those of us who might inadvertently create race conditions?

Re: A study on robustness and reliability of large language model code generation

#175

Earlier quoted context omitted.

Not true. They are well past 1%. That's not looking at reality. You can lower it below 38% but 1% is pure fantasy.

First off, you did not address most of my point. To answer your complaint, I am absolutely looking at reality. Please point us to an actual real-world professional programming job, matching the criteria I list above, for which a current LLM can economically replace a human for over a 5-year period. My believe is that there are exactly zero such jobs. If your claim is that we're at least at 1%, then you're claiming th…

>My believe is that there are exactly zero such jobs. If your claim is that we're at least at 1%, then you're claiming that there are at least 269k fully automatable programmer jobs [1]. It shouldn't be hard for you to find at least one to start.

If I built an entire car but I am missing the key. I cannot drive the car but the car is 99% complete. It's just missing the key. So what happens is, it can replace 0% of human locomotion but it's 99% complete? get it?

Just because the tool can't be used doesn't mean it's zero percent of the way to being drive able. Same with your job. The LLM can't replace any job yet, but it doesn't mean it's 1% of the way there.

But you see what I just explained to you is obvious. You already know this just like I and every one else on the face of the earth already knows this fact.

You're taking the discussion into a play on words. What does 1% apply to? Rather then use common sense and derive what I mean you prefer to redirect the conversation into a very specific definition of 1% that serves your own purpose.

We can play this game all day. We can discuss whether your application of 1% is more fitting than my application of 1%. What a waste of everyone's time.

I think it's better if rather then playing these games to "win" discussions, use your common sense to move the discussion past these games. Otherwise we're going to be talking about obvious things all day and getting all worked up about personal definitions.

Clearly the LLM is more than just 1%.

Re: A study on robustness and reliability of large language model code generation

#176

Earlier quoted context omitted.

> I'm pointing out that that isn't a good way to screen papers, because most English speakers don't share the same dialect. You did not make that clear at all. > Perhaps unstated or understated is the implication that, were you to discard papers in this way, you're liable to miss out. Indeed, that implication was totally missing. > Feel free to downvote my comment if you don't feel it was appropriate or informative o…

Alrighty. Thanks for the feedback. Have a great day. (For what it's worth, I often feel my comments are too long and wordy, and occasionally I've gotten feedback to that effect, so I was trying to be brief in my original comment.)

i understood your original comment :)

totally agree that dismissing the content of a paper for its grammatical mistakes alone is foolish.

this isn’t a paper about classic roman literature

Re: A study on robustness and reliability of large language model code generation

#177

Earlier quoted context omitted.

>You’re implying 100% reliability is actually attainable. If that were the case, wouldn’t that mean the halting problem would have been solved by AI? I’m not an expert but I’ve heard that’s like one of those fundamental laws of information theory that really can’t be broken. Nah even 100% reliability isn't attainable by a human. 100% obviously doesn't imply solving the halting problem. 100% as in 100% as reliable as…

I'm not going to rebut everything but I want to point out that I think Geoffrey Hinton is way more in your camp in regards to faith in the technology (which is what we're discussing). I'm in the camp that thinks LLMs are at best a viable competitor to furby -- I have zero existential fears and have very little confidence in them. I'm not opposed to AIs, I just think the state of the art is really really bad (sorry) a…

>I'm not going to rebut everything but I want to point out that I think Geoffrey Hinton is way more in your camp in regards to faith in the technology (which is what we're discussing). I'm in the camp that thinks LLMs are at best a viable competitor to furby

Your point was on whether everybody wants to see AI take over. To that Geoffrey Hinton doesn't want AI to take over, while you do. I for one am not sure.

As for whether Geoffrey Hinton thinks AI is legit or not is a different story. On that topic he's on my side, but that was besides the point I brought him up to point out that your take on "everybody" wanting AI to develop further is incorrect.

> and the doom and existentialism is a marketing ploy to a world imbued in conspiracy theory and institutional distrust in order to compensate for a wildly over promised and under delivered product.

Other way around. With all the money and business interests going into LLMs business interests are promoting LLMs in a future that you want and that is not apocalyptic. The conspiracy theories aren't a thing. It makes no sense as those theories don't align with where the money is being thrown.

>But hey, when you gotta raise money, you gotta raise money, and as you put it, the "trend line is going up" and that's all that matters.

Bro. I am not saying "trendline" as if it's something I have to keep throwing money at to support.

I am talking about a mathematical projection based on data. The pace of technology in the past when graphed points to an ever increasing line on a line graph. When you take the slope of that line and use it to do a quantitative prediction, that line just points up. That's just a fact of reality. The logical outcome of the data we see.

Re: A study on robustness and reliability of large language model code generation

#178
post #129
post #44

Earlier quoted context omitted.

I just don’t see how it could save time. Programmers nearly universally agree that reading code is harder than writing it. So when you have chat GPT writing your code, you have to read and understand it to ensure it’s actually doing what you need it to without awkward bugs or problems. Given that it’s harder to read than write code, it seems to stand to reason that the “time savings” must be nil.

"I just don’t see how it could save time." Have you tried it much? It's saved me a ton of time over the past 8 months. I don't know what else to say - I'm not the kind of person to deceive myself into thinking something is saving me time when it isn't. The time saved is mainly in the micro-research you no longer have to do - the bits where you have to go and look up how to write a for loop in Bash, or how to call sup…

I have found it to save time when I'm learning a new, widely used framework. (Presumably this would hold for a language as well.) I used it to learn Keras and React + Redux, and it was helpful there. For instance I was able to ask it about what sort of messages would be emitted by Redux in a given circumstance, and the precise values were wrong but it didn't really matter. That was very cool.

Once I have a good sense for how things work, I'd much rather go to the documentation. Most of the time I can do this straight from my editor by typing the method I want and querying the LSP for documentation. In those cases it's more or less instant, there's no improvement to be made there.

Maybe in some of the other cases it would be faster via an LLM integrated with my editor, but I don't think it would justify the cognitive load of considering whether or not it's correct. I'd like to worry about whether the application is correct. And I don't really think I'm wasting time waiting for the Python documentation to load, I'm continuing to chew on the problem.

Additionally, when I go to the documentation, I'm looking for callouts about the safety of an API, and I take the lack of such callouts as authoritative. Eg, recently I was reading that in node-postgres you have to return connections to the pool, which was a good refresher for me, because previously I'd been using Rust where resources are generally cleaned up automatically. I just don't trust a lack of a callout from GPT-4 as being a confirmation of a lack of a safety issue.

Re: A study on robustness and reliability of large language model code generation

#179

Earlier quoted context omitted.

Bear in mind that most researchers are not native English speakers.

This argument is not really valid anymore, especially in a paper that talks about LLMs.

> not really valid anymore

You have to express reasons when not evident.

In case, just in case (and to also show how lack of definition may trigger any weak interpretation), after suggestion from a nearby member, the idea would had been that today people could have sentences reformulated by language models: the composer intends to express a logic, so any translator would have to grasp said logic primarily and first.

Re: A study on robustness and reliability of large language model code generation

#180

Earlier quoted context omitted.

my mistake. 40% of the way there, point still stands.

> point still stands No? I assume you actually don't know the concept of diminishing return, and it's totally fine. Every other comment here is explaining it to you already. I won't bother to repeat. Please read the sibling comments carefully.

"Hey do you know what exponential growth is?" You think I just say that sentence and suddenly all your arguments are suddenly flushed down the toilet because I stated some random concept out of nowhere? Come on. You need evidence.

You brought up this question of diminishing returns out of nowhere and offer no evidence for it. So why would my point not stand?

I read the sibling comments. One person brought it up but offered no proof that this is what's happening. Simply did what you did in a manner that was less rude, just stated the concept of diminishing returns applies to LLMs without offering a shred of evidence that indicates this is what is happening.

There is a difference between introducing the concept of diminishing returns out of nowhere, and providing evidence that diminishing returns is WHAT is actually HAPPENING with LLMS.

We're about a year out from the introduction of chatgpt, that's not enough time to know if all the gains in the past decade of AI has suddenly hit wall of diminishing returns.

Post reply on HN