Live data from Hacker News

What we learned in 6 months of working on an AI Developer

blog.pythagora.ai

11–20 of 51 posts

Re: What we learned in 6 months of working on an AI Developer

#11
post #7

Even though I don't think GPT-4 is up to the task, it does seem like now is the right time to be working on these things. Pretty soon GPT-4 will not be the best in the field. The next generation will perform much better. Possibly the most frustrating thing I find about GPT-4 is how close it gets with it's wrong answers. It's easy to dismiss a lesser answer when it responds with a laughably out-of-band idea. GPT-4 oft…

> Pretty soon GPT-4 will not be the best in the field. The next generation will perform much better.

What makes you believe that progress is linear, or at least a line forever going up?

I keep seeing people predicting rapidly improving AI, based on how rapid it improved over the last x months.

But why is that not an outlier? How do we know we haven't hit a ceiling and stagnating? Isn't progress typically very bumpy and sudden?

Re: What we learned in 6 months of working on an AI Developer

#12
post #7

Even though I don't think GPT-4 is up to the task, it does seem like now is the right time to be working on these things. Pretty soon GPT-4 will not be the best in the field. The next generation will perform much better. Possibly the most frustrating thing I find about GPT-4 is how close it gets with it's wrong answers. It's easy to dismiss a lesser answer when it responds with a laughably out-of-band idea. GPT-4 oft…

Oh man. When it’s so close but wrong it’s amazing for creative endeavors! For technical ones, it is quite a bad thing. It’s like being a Star Wars fan but the AI just wants to talk about Star Trek.

I think this is why the non-tech people see AI as so amazing. For anything human and non-technical, the “almost but not quite” nature is a good thing.

I was using an AI to help me debug a weird thing (mainly summarizing log splats hundreds of lines long) and I eventually got pretty close to identifying the issue when I asked “wtaf is this message. Never seen anything like it.” It then went on about how it was offended that I used vulgar language. I had to apologize for saying “wtaf!” Anyway, I found a bug in a linker, so that was fun; thanks Al.

Re: What we learned in 6 months of working on an AI Developer

#13
One of the things they seem to have figured out is the requirement to at least model a sort of actor-critic architecture with their agents. It helps quite a bit.

They seem to badmouth Aider a tad (not cool) but I do wonder how a full-stack of this + Aider might work? There needs to also be some sort of good test generator involved.

All that said, any time someone actually demonstrates progress on the automated Software Engineer problem and it makes it to HN, I am deeply reminded of the old quote:

"It is difficult to get a man to understand something, when his salary depends on his not understanding it."

Just read through this comments section and check out the pure copium. Yes, ChatGPT can do basic sysadmin tasks with ./configure and make.

Yes it does make sense to work on this now, assuming LLMs will get better, because LLMs have continued to get better on any metric you can imagine.

Finally, yes, AI devs will make landing pages and basic APIs. I didn't realize we were all hardcore world-class 0.01% programmers? I have certainly written a landing page and basic API before, in fact I do that sort of thing a lot more than I write uber1337 hax0r code. You probably do too!

Re: What we learned in 6 months of working on an AI Developer

#14
post #11
post #7

Even though I don't think GPT-4 is up to the task, it does seem like now is the right time to be working on these things. Pretty soon GPT-4 will not be the best in the field. The next generation will perform much better. Possibly the most frustrating thing I find about GPT-4 is how close it gets with it's wrong answers. It's easy to dismiss a lesser answer when it responds with a laughably out-of-band idea. GPT-4 oft…

> Pretty soon GPT-4 will not be the best in the field. The next generation will perform much better. What makes you believe that progress is linear, or at least a line forever going up? I keep seeing people predicting rapidly improving AI, based on how rapid it improved over the last x months. But why is that not an outlier? How do we know we haven't hit a ceiling and stagnating? Isn't progress typically very bumpy a…

>What makes you believe that progress is linear, or at least a line forever going up?

I assume neither of those things. I have however read a lot of the papers published since GPT-4 was trained. There have been a lot of advances since then, so much so that simply saying "a lot" seems to be a massive understatement.

I think it is a reasonable assumption that at least a portion of those advancements would be able to build upon the existing technology of GPT-4 to produce something greater.

I am not assuming discoveries yet to be made. I am considering existing discoveries that have not yet made it into the top level of production.

Re: What we learned in 6 months of working on an AI Developer

#16
post #8

Maybe AI developers can make landing pages and basic APIs. But, taking front end as an example, I just don't see how an AI can reproduce exact design specifications and interactivity to the point where it wouldn't just be faster to write the code yourself or search for some human verified snippet that does what you want. And programmers who do know how to actually write efficient code without AI seem like they'd be e…

> When the dishwasher was invented, everyone thought the human dish washer would be obsolete. And yet, restaurants still employ dish washers because they are much more efficient and thorough than a dishwashing machine.

This is a good example of both job destruction and job retention by technology.

Job destruction - the total number of potential hand dishwasher jobs has reduced because the vast majority of commodity dishwashing is machine driven.

Job enhancement - machine dishwashers just can't produce the quality/dexterity of hand dishwashers.

I feel like generative AI will do the same. It will replace a large number of commodity jobs - editors, translators, copy producers, website designers, app prototypes, paper pushers but it will also reveal the value of skilled producers.

Too risky to let chatGPT write code for your backend that destroys your production database and crashes your company forever.

Re: What we learned in 6 months of working on an AI Developer

#17
post #6

Until I see an AI sysadmin that can help with basic configure/make problems, I don't have high hopes for an AI developer.

That should be quite easy compared to software development, which is much more open-ended since the requirement are usually more nebulous, potentially contradictory, and at times simply wrong.

Re: What we learned in 6 months of working on an AI Developer

#18
post #11
post #7

Even though I don't think GPT-4 is up to the task, it does seem like now is the right time to be working on these things. Pretty soon GPT-4 will not be the best in the field. The next generation will perform much better. Possibly the most frustrating thing I find about GPT-4 is how close it gets with it's wrong answers. It's easy to dismiss a lesser answer when it responds with a laughably out-of-band idea. GPT-4 oft…

> Pretty soon GPT-4 will not be the best in the field. The next generation will perform much better. What makes you believe that progress is linear, or at least a line forever going up? I keep seeing people predicting rapidly improving AI, based on how rapid it improved over the last x months. But why is that not an outlier? How do we know we haven't hit a ceiling and stagnating? Isn't progress typically very bumpy a…

I would say that Microsoft/OpenAI’s attacks on open source, whether it be through “AGI safety” BS as a front for regulating their way to a monopoly, or attempting Embrace, Exentend, Extinguish on companies like Mistral, and Cold War-style fear mongering about China, are the greatest near-term risks to linear progress. And it’s worth noting on that latter point that China is not similarly constrained and so could end up outcompeting the U.S., regardless

Re: What we learned in 6 months of working on an AI Developer

#19
post #4

The focus on upfront specs feels a bit off. Since it's apparently cheap to generate running code, as a user, I'd much rather be able to just iterate really fast and use output to refine my requirements rather than having to laboriously state them all up front. Agile rather than waterfall if you will.

In that case, it might be easier to start over with fixed specs. That might not be as much work as it sounds like since most of the existing code would have been produced by the LLM, and only the human feedback and interventions would have to be redone. It would be almost like backtracking to an earlier point in a chat history and changing path there. TA talks about that another LLM could provide insight about where to change what.

It might also be possible to change an existing history without abandoning all which has happened afterwards. Of course, this could lead to conflicts, sort of like when rebasing a branch, and it would be useful to have another LLM look for it.

GPT Copilot might or might not be able to start from existing code as well and one would approach it as one would a legacy codebase that has to be adapted to new requirements.

Post reply on HN