A study on robustness and reliability of large language model code generation
1–10 of 229 posts
Re: A study on robustness and reliability of large language model code generation
#2Re: A study on robustness and reliability of large language model code generation
#3Re: A study on robustness and reliability of large language model code generation
#4Re: A study on robustness and reliability of large language model code generation
#5So it's already beating humans at this too?
Like can Liutenant Commander Data be excluded from human competitions? Yes. But can a replicant from Blade Runner? I guess the answer is that they will not feel the need to compete. But what if they do?
Level 1: https://en.m.wikipedia.org/wiki/The_Measure_of_a_Man_(Star_T...
Re: A study on robustness and reliability of large language model code generation
#6Re: A study on robustness and reliability of large language model code generation
#7Re: A study on robustness and reliability of large language model code generation
#8 We collect 1208 coding questions from StackOverflow on 24 representative Java APIs. We summarize thecommon misuse patterns of these APIs and evaluate them oncurrent popular LLMs. The evaluation results show that evenfor GPT-4, 62% of the generated code contains API misuses,which would cause unexpected consequences if the code isintroduced into real-world software.
Should have used ChatGPT to correct those sentences. /sWhat do we consider an API misuse at the end of the day? And also against the Java APIs... Which are some of the oldest.
I'm not saying ChatGPT answers work, 50% of the time they don't work out of the box. But overal, the time savings are incredible, compared to the fossil digging way of googling. And I value the time savings towards the equally correct way of googling for answers. You reap what you sow. The knowledge the LLM is based upon is the same stuff you used to distill manually. It shouldn't be more or less correct, if you think of it that way.
Surprise, ChatGPT will not replace devs. But it is like having the electrical drill vs. screwdriver type of situation.
Re: A study on robustness and reliability of large language model code generation
#9So it's already beating humans at this too?
I wonder what would happen in 39 years if an AI which is designed to feel and emulate humans enters human competitions and beats them at everything. Like can Liutenant Commander Data be excluded from human competitions? Yes. But can a replicant from Blade Runner? I guess the answer is that they will not feel the need to compete. But what if they do? Level 1: https://en.m.wikipedia.org/wiki/The_Measure_of_a_Man_(Star_…
Re: A study on robustness and reliability of large language model code generation
#10Earlier quoted context omitted.
I wonder what would happen in 39 years if an AI which is designed to feel and emulate humans enters human competitions and beats them at everything. Like can Liutenant Commander Data be excluded from human competitions? Yes. But can a replicant from Blade Runner? I guess the answer is that they will not feel the need to compete. But what if they do? Level 1: https://en.m.wikipedia.org/wiki/The_Measure_of_a_Man_(Star_…
If there's a human only competition and the AI "identifies" as a human, you'd expect it still wouldn't be let in. Anything else would be absurd, right?