Live data from Hacker News

Study finds that 52% of ChatGPT answers to programming questions are wrong

futurism.com

91–100 of 104 posts

Re: Study finds that 52% of ChatGPT answers to programming questions are wrong

#91
post #88

Earlier quoted context omitted.

If someone showed me this solution I'd have quite a few questions. Like, why is there a 'newline in the example section, and why isn't that part in a comment? Why introducing the "helper" to enforce that the execution always begins at 1? Could there be some other way to design the program so that the four conditions don't all end with the same twenty or so characters?

Those are questions of style. The allegation was that ChatGPT produces answers that are wrong .

No, the context here is interns. If someone came to me and asked for an internship and showed that to me I'd signal that they need to show me something else that is more impressive, or somehow find a way to explain and motivate that code that convinces me that it is a decent solution.

I'm not sure how you're discerning "style" from "wrong". Would using some esolang be a matter of "style" as long as the asserts on the output pass?

Re: Study finds that 52% of ChatGPT answers to programming questions are wrong

#92
post #91

Earlier quoted context omitted.

Those are questions of style. The allegation was that ChatGPT produces answers that are wrong .

No, the context here is interns. If someone came to me and asked for an internship and showed that to me I'd signal that they need to show me something else that is more impressive, or somehow find a way to explain and motivate that code that convinces me that it is a decent solution. I'm not sure how you're discerning "style" from "wrong". Would using some esolang be a matter of "style" as long as the asserts on the…

> I'm not sure how you're discerning "style" from "wrong".

"Wrong" = "produces incorrect results".

What else?

> or somehow find a way to explain and motivate that code that convinces me that it is a decent solution.

You don't actually know Scheme, do you?

Re: Study finds that 52% of ChatGPT answers to programming questions are wrong

#93
post #72
post #9

Earlier quoted context omitted.

> I expect it to get me in the ballpark faster than I would have on my own. This is great if you are an experienced developer who can tell the difference between "in the ballpark" and fixable and "in the ballpark" but hopeless.

If 52% of responses have a flaw somewhere, then 48% of responses are flawless. That is amazing and important. The headline should be "LLM gives flawless responses to 48% of coding questions."

There are articles every day about how AI is replacing programmers, coding is dead etc, including from the nvidia ceo this week. This kind of thing shows we are not quite there yet. There are lots of folks on twitter etc who rave about how genAI built a full app for them, but in my experience that comes with a huge amount of human trial and error and understanding to know what needs to be tweaked.

Re: Study finds that 52% of ChatGPT answers to programming questions are wrong

#94
post #14

Earlier quoted context omitted.

How is this any different than the age old "googling stack overflow" method everyone's been using for years?

Stack Overflow is a community with an answer-rating system and there is often some level of review from other people commenting on the answer's advantages and shortcomings. You often have multiple answers to choose from too. Those features build trust in an answer or prompt you to look elsewhere. The UI for an LLM answer would have difficulty replicating the same thing since every answer is (probably) a new one and y…

I've never been that fond of SO, but I find chatgpt very useful. SO tends to be simpler one time questions. My interactions with gpt are conversations working towards a solution.

LLMs totally have a beginner problem. It's much like the problem where a beginner knows they need to look something up but can't figure out the right keywords to search for.

Also chatgpt has never called me an idiot for asking a stupid question, having not read the question properly, and making the assumption it was the same as an existing question after skimming it. I wouldn't ask SO a question these days, the response is more likely to be toxic than helpful.

Re: Study finds that 52% of ChatGPT answers to programming questions are wrong

#95
post #72

Earlier quoted context omitted.

If 52% of responses have a flaw somewhere, then 48% of responses are flawless. That is amazing and important. The headline should be "LLM gives flawless responses to 48% of coding questions."

There are articles every day about how AI is replacing programmers, coding is dead etc, including from the nvidia ceo this week. This kind of thing shows we are not quite there yet. There are lots of folks on twitter etc who rave about how genAI built a full app for them, but in my experience that comes with a huge amount of human trial and error and understanding to know what needs to be tweaked.

> This kind of thing shows we are not quite there yet

I think you need to probably consider the time it took to go from 100% wrong, to 90, to 80, etc…. My guess is that interval is probably shrinking from milestone to milestone. This causes me to suspect that folks starting SWE careers in 2024 will not likely not be SWEs before 2030.

That’s why I tell my grandkids they should consider plumbing and HVAC trades instead of college. My bet is within 10 years nearly every vocation that you require a college degree will be made partially or completely obsolete by AI.

Re: Study finds that 52% of ChatGPT answers to programming questions are wrong

#96
post #91

Earlier quoted context omitted.

No, the context here is interns. If someone came to me and asked for an internship and showed that to me I'd signal that they need to show me something else that is more impressive, or somehow find a way to explain and motivate that code that convinces me that it is a decent solution. I'm not sure how you're discerning "style" from "wrong". Would using some esolang be a matter of "style" as long as the asserts on the…

> I'm not sure how you're discerning "style" from "wrong". "Wrong" = "produces incorrect results". What else? > or somehow find a way to explain and motivate that code that convinces me that it is a decent solution. You don't actually know Scheme, do you?

If it helps you answer my question, assume that I don't.

Re: Study finds that 52% of ChatGPT answers to programming questions are wrong

#97
ChatGPT isn’t the best coding LLM. Claude Opus is.

Also as you can always tell if a coding response works empirically mistakes are much more easily spotted than in other forms of LLM output.

Debugging with AI is more important than prompting. It requires an understanding of the intent which allows the human to prompt the model in a way that allows it to recognize its oversights.

Most code errors from LLMs can be fixed by them. The problem is an incomplete understanding of the objective which makes them commit to incorrect paths.

Being able to run code is a huge milestone. I hope the GPT5 generation can do this and thus only deliver working code. That will be a quantum leap.

Re: Study finds that 52% of ChatGPT answers to programming questions are wrong

#98
post #96

Earlier quoted context omitted.

> I'm not sure how you're discerning "style" from "wrong". "Wrong" = "produces incorrect results". What else? > or somehow find a way to explain and motivate that code that convinces me that it is a decent solution. You don't actually know Scheme, do you?

If it helps you answer my question, assume that I don't.

You seem determined to shift the goalposts, so I think we're done here. It's not difficult:

1) Many (some say most) job applicants cannot write a working implementation of FizzBuzz.

2) ChatGPT can write a working implementation of FizzBuzz.

∴ ChatGPT is a better programmer than many (most) job applicants, at least on this specific problem.

QED.

What you are doing part of a long tradition of AI denial. Take (e.g.) chess. First it was "a computer cannot play chess". Then it was "a computer cannot play chess well enough to beat a human being." Then it was "a computer cannot play chess well enough to beat a grandmaster." (you are somewhere between this stage and the previous one) Then it was "a computer cannot play chess well enough to beat the world champion." Then it was "playing chess is not a measure of intelligence." Notice how the goalposts gradually move so that eventually the criterion is "it doesn't count unless the computer is better than the best person in the world." followed by "Ehh... those grapes were probably sour anyway."

For some reason, we don't get so defensive about machinery in other areas of expertise. No one tries to deny that a D11 Caterpillar can shift dirt faster than a human being with a shovel. No one tries to deny that the Webb telescope can see distant galaxies better than any human being's naked eye.

But people freak out when it comes to "intellectual" accomplishments.

Denial is rarely a good strategy in the long run.

Re: Study finds that 52% of ChatGPT answers to programming questions are wrong

#99
post #96

Earlier quoted context omitted.

If it helps you answer my question, assume that I don't.

You seem determined to shift the goalposts, so I think we're done here. It's not difficult: 1) Many (some say most) job applicants cannot write a working implementation of FizzBuzz. 2) ChatGPT can write a working implementation of FizzBuzz. ∴ ChatGPT is a better programmer than many (most) job applicants, at least on this specific problem. QED. What you are doing part of a long tradition of AI denial. Take (e.g.) che…

Why did you write all these words and still not answer my question?

The context is interns, and I responded from the perspective that someone came to me with that code and asked to be an intern. You seem very convinced that most people that apply for software development jobs can't "write a working implementation of FizzBuzz", but I fail to see the relevance. Why count all the hairdressers and kids that have spent a few weeks on HTML and whatnot that might apply for one software related job and be refused?

It's a rhetorical question, don't waste time on it.

I think it's more interesting to look at the output from the machine as if a person had produced it and offered it in an internship process. That's a good way to at least partially neuter the influence of advertising and so on when we evaluate what you got out of it.

As for the chess part, I'm not so sure computers can play chess. Can they carry the board and pieces to the place where a match will be played? Can they unpack it, push the clock button, move the pieces? When the match takes place on e.g. Lichess, can they put a finger on a screen and move a piece, or do they need some special interface that is incompatible with humans to participate? Do they need a human to help them get to the game and initiate participation, or can they do this on their own, because they previously said to someone that they will or because they feel like it?

You're treating simulacra as ultimately real, and see me as stupid because I still have some contact with the material and don't confuse it with the virtual.

Re: Study finds that 52% of ChatGPT answers to programming questions are wrong

#100
post #95

Earlier quoted context omitted.

There are articles every day about how AI is replacing programmers, coding is dead etc, including from the nvidia ceo this week. This kind of thing shows we are not quite there yet. There are lots of folks on twitter etc who rave about how genAI built a full app for them, but in my experience that comes with a huge amount of human trial and error and understanding to know what needs to be tweaked.

> This kind of thing shows we are not quite there yet I think you need to probably consider the time it took to go from 100% wrong, to 90, to 80, etc…. My guess is that interval is probably shrinking from milestone to milestone. This causes me to suspect that folks starting SWE careers in 2024 will not likely not be SWEs before 2030. That’s why I tell my grandkids they should consider plumbing and HVAC trades instead…

I tell my grandkids that vocational school is a perfectly decent and honorable way to get into a trade that pays better than retail.

I also tell them that a good university is a perfectly decent and honorable way to begin a life of the mind. It's not the only way, and a life of the mind isn't the life everyone wants.

I also tell them that the purpose of a university education is mostly not about training for a job.

Post reply on HN