Live data from Hacker News

OpenAI Researchers Find That AI Is Unable to Solve Most Coding Problems

futurism.com

161–170 of 174 posts

Re: OpenAI Researchers Find That AI Is Unable to Solve Most Coding Problems

#161
post #156

Earlier quoted context omitted.

I'm not really a frontend developer. I'm using Elm for toy projects, in fact I did one recently.[0] Elm is my favourite language! > You spend about three days trying to get it to build then say fuck it and rewrite it. What are the problems you encounter? I can't quite imagine in what way an Elm project could be hard to build! (Also not trying to be offensive, but I almost don't believe you!) And into which language d…

typescript usually, the elm frontends tend to be in some abandoned repo which hasn't had a ci run in like 2 years and which instantly fail on missing deps or security controls etc.

Yes, right, I forgot, this also happened to me once: someone deleted their repo for an Elm library. It was salvageable through whatever archive and publishing my own copy.

It happens less often in Elm than in JavaScript though! I'll take "abandoned for two years" Elm project over "abandoned for two years" typescript project anytime!

The problem, in your case, was not really Elm.

Re: OpenAI Researchers Find That AI Is Unable to Solve Most Coding Problems

#162
post #85
post #56

Earlier quoted context omitted.

> unless non-SWE people using LLMs become technical enough to get what they need out of them Non-SWE person here. In the past year I've been able to use LLMs to do several tasks for which I previously would have paid a freelancer on Fiverr. The most complex one, done last spring, involved writing a Python program that I ran on Google Colab to grab the OCR transcriptions of dozens of 19th-century books off the Interne…

> I didn't put any full-time SWEs out of work, but I did take one job away from a Fiverr freelancer. I think this is the nuance most miss when they think about how AI models will displace work. Most seem to think “if it can’t fully replace a SWE then it’s not going to happen” When in reality, it starts by lowering the threshold for someone who’s technical but not a SWE, to jump in and do the work themselves. Or it ma…

Entirely possible. Have you got any numbers and real world examples? Growth? Profits? Actual quantified productivity gains?

The nuance your 'gotcha' scenario miss is that displacing fiverr, speeding up small side project, making scripts fo non-SWE, creating boilerplate, etc is not the trillions of dollars disruption that is needed by now.

Re: OpenAI Researchers Find That AI Is Unable to Solve Most Coding Problems

#163
post #56

Earlier quoted context omitted.

> unless non-SWE people using LLMs become technical enough to get what they need out of them Non-SWE person here. In the past year I've been able to use LLMs to do several tasks for which I previously would have paid a freelancer on Fiverr. The most complex one, done last spring, involved writing a Python program that I ran on Google Colab to grab the OCR transcriptions of dozens of 19th-century books off the Interne…

If the goal is to get something to run correctly roughly once with some known data or input, then that's fine. Actual software development aims to run under 100% of circumstances, and LLMs are essentially cargo culting the development process and entrusting an automation that is unreliable to do mundane tasks. Sadly the quality of software will keep going down, perhaps even faster.

Stop with the realism, one off scripts is going to give trillions in ROI any day now. Personally could easily chip in maybe a million a month in subsbription fees be cause my bolierplate code I write once in a blue moon has speed up infinitely and I will cash out in profits any day now.

Re: OpenAI Researchers Find That AI Is Unable to Solve Most Coding Problems

#164
post #66

I find the framing of this story quite frustrating. The purpose of new benchmarks is to gather tasks that today's LLMs can't solve comprehensively. It an AI lab built a benchmark that their models scored 100% on they would have been wasting everyone's time! Writing a story that effectively says "ha ha ha, look at OpenAI's models failing to beat the new benchemark they created!" is a complete misunderstanding of the r…

Shhh ... you're spoiling everybody's confirmation bias against LLMs. They are obviously terrible at coding, just as we have known all along, and everybody should laugh at them. Nothing to see here!

Since you are one of the cool kids in the know, can you share the road map to profitability and even better the expected/hyped ROI? Without extrpolations into science fiction, please.

Re: OpenAI Researchers Find That AI Is Unable to Solve Most Coding Problems

#165
> SWE-Lancer, built on more than 1,400 software engineering tasks from the freelancer site Upwork

What's not mentioned here (I think) is that tasks in this benchmark are priced and they sum up to million dollars.

And current AIs were able to earn nearly half of that.

So while technically they can't solve most problems (yet) they are already perfectly capable of taking about 40% of the food off your plate.

Re: OpenAI Researchers Find That AI Is Unable to Solve Most Coding Problems

#166
To all those devs saying "i tried it and it wasnt perfect the first time, so I gave up", I am reminded of something my father used to say: "A bad carpenter blames his tools"

So AI not going to answer your question right on its first attempt in many cases. It is forced to make a lot of assumptions based on the limited info you gave it, some of those may not match your individual case. Learn to prompt better and it will work better for you. It is a skill, just like everything else in life.

Imagine going into a job today and saying "i tried google but it didnt give me what I was looking for as the first result, so I dont use google anymore". I just wouldnt hire a dev that couldnt learn to use AI as a tool to get there job done 10x faster. If that is your attitude, 2026 might really be a wake-up call for your new life.

Re: OpenAI Researchers Find That AI Is Unable to Solve Most Coding Problems

#168

To all those devs saying "i tried it and it wasnt perfect the first time, so I gave up", I am reminded of something my father used to say: "A bad carpenter blames his tools" So AI not going to answer your question right on its first attempt in many cases. It is forced to make a lot of assumptions based on the limited info you gave it, some of those may not match your individual case. Learn to prompt better and it wil…

One of the reasons SO works is that the correct or best answer tends to move toward the top of the list. AI struggles to do this reliably - and so its closer to SO where the answers are randomly selected and you have to try a few to get the one that's correct.

Also, I bet your father imagined the carpenter's toolbox full of well accepted useful tools. For many of us, non-bad carpenters, AI hasn't made the cut yet.

Re: OpenAI Researchers Find That AI Is Unable to Solve Most Coding Problems

#169
post #58

Earlier quoted context omitted.

I mean it’s been a couple of years! It may or may not happen but “scam” means intentional deceit. I don’t think anyone actually knows where LLMs are going with enough certainty to use that pejorative.

>“scam” means intentional deceit. Yes. I'm pretty sure any engineer working on this knows it's not "a few years away". But it doesn't stop product teams from taking adcvantadge of the hype cycle. Hence, "use deception to deprive (someone) of money or possessions.".

I'm equally sure there are true believers who've drunk the Kool-aid and really believe AGI is right around the corner, just need to fix a few bugs and wait just a few more generations of Moore's law. What difference does the beliefs of a nameless engineer at an AI company make?

Re: OpenAI Researchers Find That AI Is Unable to Solve Most Coding Problems

#170

Earlier quoted context omitted.

>“scam” means intentional deceit. Yes. I'm pretty sure any engineer working on this knows it's not "a few years away". But it doesn't stop product teams from taking adcvantadge of the hype cycle. Hence, "use deception to deprive (someone) of money or possessions.".

I'm equally sure there are true believers who've drunk the Kool-aid and really believe AGI is right around the corner, just need to fix a few bugs and wait just a few more generations of Moore's law. What difference does the beliefs of a nameless engineer at an AI company make?

>What difference does the beliefs of a nameless engineer at an AI company make?

Hopefully a manager and proper task scheduling. If I made these promises every sprint and kept saying "yea the task is only a week away from completion!" I'd be fired unless I fell down the rabbit hole to Alice in Wonderland. I'm using good faith to assume a lot of those AI engineers are smarter and better schedulers than I am.

But that's what managers and proper scoping and perspective is for. Maybe they're okay with that, but I'd wager any profit motivated company would not keep exploring unless the gains are enormous.

Post reply on HN