Live data from Hacker News

Testing GPT 4's code-writing capabilities with some real world problems

tylerglaiel.substack.com

471–480 of 677 posts

Re: Testing GPT 4's code-writing capabilities with some real world problems

#471
post #183

I want to see GPT-4 dealing with this situation: - they: we need a new basic POST endpoint - us: cool, what does the api contract look like? URL? Query params? Payload? Response? Status code? - they: Not sure. Third-party company XXQ will let you know the details. They will be the ones calling this new endpoint. But in essence it should be very simple: just grab whatever they pass and save it in our db - us: ok, cool…

Generally agreed, although I think LLM's in the near- and medium-term will end up being useful for things like:

* checking if code will be impacted by breaking changes in a library upgrade

* converting code to use a different library/framework

* more intelligent linting/checking for best practices

* some automated PR review, e.g. calling out confusing blocks of code that could use commenting or reworking

Re: Testing GPT 4's code-writing capabilities with some real world problems

#472
post #8

In before all the comments about how “most code is trivial” or “most programming is stuff that already exists” or “you’re missing the point look how it’s getting better”. I really am in awe of how much work people seem willing to do to justify this as revolutionary and programmers as infantile, and also why they do that. It’s fascinating. Thinking back to my first job out of college as a solid entry level programmer.…

For me, its mostly that I have used GPT-3.5 a little for programming C++, and I wasnt impressed. For one, it made horrible, glaring mistakes (like defining extern functions which dont exist, using functions which are specific to a platform im not using, etc.), stuff beginners would do. It also decided to sneak in little issues, such as off-by-one errors (calling write() with a buffer and a size that is off by one in…

GPT4 is very different from 3.5. I've asked it today to write some unit tests given the code of the class (~200 lines) and the methods I wanted to cover and it did that just perfectly. It put asserts where it made sense (without me asking to do it), and the unit test code was better written than some code I've seen written by (lazy) humans. It's not perfect sure and it's easy to get a bad response but give OpenAI a few more iterations and my job will be simply to copy paste the requirement to GPT and paste the generated code back to compile.

Re: Testing GPT 4's code-writing capabilities with some real world problems

#473
post #439

Earlier quoted context omitted.

> If increased productivity equaled job loss there would be two programmers alive today, doing the same job as the fewer than 10000 programmers using punch cards as we entered the year 1950. The only reason it's not the case in this example is because computers at the time were a tiny early adopter niche, which massively multiplied and expanded to other areas. Like, only 1 in 10,000 businesses would have one in 1950,…

Have no doubt, we will find new places to put computers. In the 80's and even the 90's everyone said the same thing, "Why do I need a computer? I can do everything I do already without a problem?" Well, turns out with computers you could do 12 more things you can never considered. Consider the interoffice memo: it'd take what, 1-2 hours to get a document from one floor to another through the system? Cool, you can wor…

>In the 80's and even the 90's everyone said the same thing, "Why do I need a computer? I can do everything I do already without a problem?" Well, turns out with computers you could do 12 more things you can never considered

Yeah. Also, unfortunately, it turns out those people in the 80s and 90s got it right. They didn't really need a computer - they'd better off without one. But as soon as we got them, we'd find some things to use them for - mostly detrimental to our lives!

Re: Testing GPT 4's code-writing capabilities with some real world problems

#474
post #8

In before all the comments about how “most code is trivial” or “most programming is stuff that already exists” or “you’re missing the point look how it’s getting better”. I really am in awe of how much work people seem willing to do to justify this as revolutionary and programmers as infantile, and also why they do that. It’s fascinating. Thinking back to my first job out of college as a solid entry level programmer.…

For me, its mostly that I have used GPT-3.5 a little for programming C++, and I wasnt impressed. For one, it made horrible, glaring mistakes (like defining extern functions which dont exist, using functions which are specific to a platform im not using, etc.), stuff beginners would do. It also decided to sneak in little issues, such as off-by-one errors (calling write() with a buffer and a size that is off by one in…

A Toolformer model integrated with Intellisense is the obvious next step.

Re: Testing GPT 4's code-writing capabilities with some real world problems

#475

Earlier quoted context omitted.

Two points: GPT4 is significantly better in this regard, and you should be concerned about the rate of progress more than it’s actual capabilities today.

I suspect (and sort of hope) they're close to hitting a local maxima presumably they've already fed it all the code in the world (including a load they probably aren't licensed to)

If you tried GPT4, you probably understood it's not about feeding all the code in the world. GPT4 analyzes and "understands" your code and will answer based on this. Clearly, it will read the variable names and make deductions based on this. It will actually read the comments, the function names and make decisions based on this. And it knows the rules of the language. I mean, I'm writing this because this is what I've witnessed since the time I spent playing with it.

The problem I've seen is that, maybe like the author has been writing, it's making sh*t up. That's not untrue, sometimes I didn't give it all dependent classes and it tried to think sometimes correctly, sometimes incorrectly what those were (such as method signatures, instance members, etc.) I wish it would have asked me some details rather than trying to figure things out. The guys at OpenAI have still a lot to do, but the current status is very impressive

Re: Testing GPT 4's code-writing capabilities with some real world problems

#476
post #8

In before all the comments about how “most code is trivial” or “most programming is stuff that already exists” or “you’re missing the point look how it’s getting better”. I really am in awe of how much work people seem willing to do to justify this as revolutionary and programmers as infantile, and also why they do that. It’s fascinating. Thinking back to my first job out of college as a solid entry level programmer.…

While I think there's truth to what you say, I'd also point our that workers in many pre-automated industries with an "artisan" approach also considered themselves irreplaceable because they figured, correctly, that nobody could build a machine with the capability of reproducing their workflow, with all its inherent uncertainty, flexibility and diverse physical and mental skills. What they failed to predict was that…

The artisan of old created one bowl at a time and made artisan pots. A machine that makes bowl can automate the artisan's job away.

However, the next challenge is that the machine itself is now an "artisan" device.

I'm sure the first bowl printing machine ( https://youtu.be/bD2DNSt8Wb4 ) was purely artisan... but now you can buy them on Alibaba for a few thousand dollars ( https://www.alibaba.com/product-detail/Printing-Machine-Cera... )

I am sure there is a (bowl printing machine) machine out there.

But if you say "I want a bowl printing machine that can do gradient colors" that first one (and probably the first few until it gets refined) are all going to be artisanal manufacturing processes again.

This all boils down to that at some point in the process, there will be new and novel challenges to overcome. They're moving further up the production chain, but there is an artisan process at the end of it.

The design of a new car has changed over time so that it is a lot more automated now than it was back then ( https://youtu.be/xatHPihJCpM ) but you're not going to get an AI to go from "create a new car design" to actually verifying that it works and is right.

There will always be an artisan making the first version of anything.

Re: Testing GPT 4's code-writing capabilities with some real world problems

#477
post #314

Earlier quoted context omitted.

If as GP says AI could automate 1% of a single programmer's job (the boring part where you write code), then how on earth can you derive that a single programmer could do the job of 5 with AI? It's completely illogical.

The GP estimate is way off... If they indeed wait for input from other departments/companies 99% of the time (so they just need to actually program 5 minutes in their 8-hour workday), then they can be already thrown out of a job and have the company do with 1/10 the programmers, no AI required...

Exactly!

Re: Testing GPT 4's code-writing capabilities with some real world problems

#478
post #183

I want to see GPT-4 dealing with this situation: - they: we need a new basic POST endpoint - us: cool, what does the api contract look like? URL? Query params? Payload? Response? Status code? - they: Not sure. Third-party company XXQ will let you know the details. They will be the ones calling this new endpoint. But in essence it should be very simple: just grab whatever they pass and save it in our db - us: ok, cool…

Obviously you just formalize the interface for exchanges API contracts... and do pre-delivery validation...

Also, ChatGPT would likely be able to extrapolate. It would just need to write an email to XXQ to confirm the change.

Cope harder... the fact that you can write an email won't save you.

Re: Testing GPT 4's code-writing capabilities with some real world problems

#479
post #183

I want to see GPT-4 dealing with this situation: - they: we need a new basic POST endpoint - us: cool, what does the api contract look like? URL? Query params? Payload? Response? Status code? - they: Not sure. Third-party company XXQ will let you know the details. They will be the ones calling this new endpoint. But in essence it should be very simple: just grab whatever they pass and save it in our db - us: ok, cool…

I've been working as a freelance software developer for about 5 years now, and my billing model is such that I only bill for hours spent writing code. Time spent communicating with people is non-negligible, which means that it has to be baked into my hourly rate. So I'm very cognizant of how much time I spend communicating with people, and how much time I spend writing code.

I strongly disagree that 99% of the effort is not writing code. Consider how long these things actually take:

> - they: we need a new basic POST endpoint

> - us: cool, what does the api contract look like? URL? Query params? Payload? Response? Status code?

> - they: Not sure. Third-party company XXQ will let you know the details. They will be the ones calling this new endpoint. But in essence it should be very simple: just grab whatever they pass and save it in our db

> - us: ok, cool. Let me get in contact with them

That's a 15 minute meeting, and honestly, it shouldn't be. If they don't know what the POST endpoint is, they weren't ready to meet. Ideally, third-party company XXQ shows up prepared with contract_json to the meeting and "they" does the introduction before a handoff, instead of "they" wasting everyone's time with a meeting they aren't prepared for. I know that's not always what happens, but the skill here is cutting off pointless meetings that people aren't prepared for by identifying what preparation needs to be done, and then ending the meeting with a new meeting scheduled for after people are prepared.

> - company XXQ: we got this contract here:

> - us: thanks! We'll work on this

This handoff is probably where you want to actually spend some time looking over, discussing, and clarifying what you can. The initial meeting probably wants to be more like 30 minutes for a moderately complex endpoint, and might spawn off another 15 minute meeting to hand off some further clarifications. So let's call this two meetings totalling 45 minutes, leaving us at an hour total including the previous 15 minutes.

> - us: umm, there's something not specified in . What about this part here that says that...

That's a 5 minute email.

> - company XXQ: ah sure, sorry we missed that part. It's like this...

Worst case scenario that's a 15 minute meeting, but it can often be handled in an email. Let's say this is 20 minutes, though, leaving us at 1 hour 15 minutes.

So your example, let's just round that up into 2 hours.

What on earth are you doing where 3 hours is 99% of your effort?

Note that I didn't include your "one week later" and "2 days later" in there, because that's time that I'm billing other clients.

EDIT: I'll actually up that to 3 hours, because there's a whole other type of meeting that happens, which is where you just be humans and chat about stuff. Sometimes that's part of the other meetings, sometimes it is its own separate meeting. That's not wasted time! It's good to have enjoyable, human relationships with your clients and coworkers. And while I think it's just worthwhile inherently, it does also have business value, because that's how people get comfortable to give constructive criticism, admit mistakes, and otherwise fix problems. But still, that 3 hours isn't 99% of your time.

Re: Testing GPT 4's code-writing capabilities with some real world problems

#480
post #183

I want to see GPT-4 dealing with this situation: - they: we need a new basic POST endpoint - us: cool, what does the api contract look like? URL? Query params? Payload? Response? Status code? - they: Not sure. Third-party company XXQ will let you know the details. They will be the ones calling this new endpoint. But in essence it should be very simple: just grab whatever they pass and save it in our db - us: ok, cool…

Consider creating an AI stakeholder that speaks for the client. This approach would allow the client to provide input that is wordy or scattered, and the consultant could receive immediate responses by asking the AI model most questions. Better yet, they can ask in a just-in-time manner, which results in less waste and lower mental stress of collecting all possible critical information upfront.

As the project progresses, the AI model would likely gain a better understanding of the client's values and principles, leading to improved results and potentially valuable insights and feature suggestions.

Post reply on HN