Live data from Hacker News

Devin: AI Software Engineer

cognition-labs.com

371–380 of 604 posts

Re: Devin: AI Software Engineer

#372

Earlier quoted context omitted.

I'll give you examples of how it helps me: 1) copilot is a terrific auto complete, and writes tremendous amounts of repetitive boilerplate 2) copilot can help me kickstart writing some complex functions starting from a comment where I tell it what is the input and expected output. Is the implementation always perfect or bug free? No. But in general I just need to review and check rather than come up with the instruct…

I ended up getting annoyed with the autocomplete feature taking over things such as snippet expansion in vscode, so I turned it off personally. I felt that the battling against the assistant made around a break even productivity gain overall. Except for regular expressions, that I have basically offloaded to AI almost in its entirety for non trivial things.

Completely agree. I turned it off and realized I can absolutely fly writing code when copilot stops getting in the way. I only turn on for writing tests now.

Re: Devin: AI Software Engineer

#373
post #312

Earlier quoted context omitted.

If you ignore how much energy you're burning while searching for dozens and dozens of articles that may or may not give you the answer you're looking for. I'd say the electricity that LLMs burn is nothing compared to my energy and time in that regard.

Id bet $50 the inference is more expensive

I feel worthless now :)

Re: Devin: AI Software Engineer

#374
We’re still at the “rhyming not reasoning” phase of LLMs. The question of whether we move past rhyming and onto reasoning is a good one, and I’m not sure what I think about it. But I am pretty sure that coding is a lot more like reasoning than it is like rhyming, at least for de novo problems above a certain level of complexity (intellectual challenge) and complication (moving parts).

I remain open minded about what’s next and at the rate things are changing, I wouldn’t rule anything out a priori for now.

Re: Devin: AI Software Engineer

#375

Earlier quoted context omitted.

Great point. I have no idea. What are your banking use cases? For me, it's mostly information I want. I don't really need a full app for that. I want to know: 1) How much money do I have? 2) Did a check I cashed clear yet? 3) Are there any unusual charges? How is my spending this month? 4) Anything I should look into? For actions I'd want to take: 1) Deposit check 2) Transfer money from account to account 3) Make a p…

This sounds like hell. Why on earth would you want that? Open the app, go to the account, enter the money with specificity, select the account to transfer it to, click the button. Sure, being able to say "Transfer 200 to Steve" is nice and all but.. I just don't... consider it much better than the process we have today?

Just because the process today is "good enough" doesn't mean that it's a worthwhile use of developer time to create that process.

But, you're highlighting one of the challenges of changing the status quo. The "new thing" has to be significantly better than the old, otherwise people won't immediately want to switch.

Similar arguments were probably made when the iPhone removed the physical keyboard.

Re: Devin: AI Software Engineer

#376

Earlier quoted context omitted.

It's worth pointing out that on their eval set for "issues resolved" they are getting 13.86%. While visually this looks impressive compared to the others, anything that only really works 13.86% of the time, when the verification of the work takes nearly as much time as the work would have anyway, isn't useful. The problem with this entire space is that we have VC hype for work that should ultimately still be being do…

> I've personally seen teams working nights and weekends, implementing solutions to never before seen problems in a few weeks, and still getting a thumbs down when they cross the finish line. This is an important lesson that all SWEs should take to heart. Nobody cares about your novel algorithm. Nobody cares about your high availability architecture. Nobody cares about your millisecond network latency optimizations.…

I wish this were the case. The amount of time I spend trying to talk principal engineers out of massive refractors because we want to get this out soon is near criminal.

Re: Devin: AI Software Engineer

#377
post #73

As a developer but also product person, I keep trying to use AI to code for me. I keep failing, because of context length, because of shit output from the model, because of lack of any kind of architecture etc etc etc. I'm probably dumb as hell, because I just can't get it to do anything remotely useful, more than helping me with leetcode. Just yesterday I tried to feed it a simple HTML page to extract a selector, I…

It's worth pointing out that on their eval set for "issues resolved" they are getting 13.86%. While visually this looks impressive compared to the others, anything that only really works 13.86% of the time, when the verification of the work takes nearly as much time as the work would have anyway, isn't useful. The problem with this entire space is that we have VC hype for work that should ultimately still be being do…

Agreed on the lack of value for 13.86% correctness — I noticed that too. This reminds me a little of last year's hype around AutoGPT et al (at around the same time of year, oddly enough); it's very promising as a measure of how far we've come since just a few years ago when that metric would be 0%, but it doesn't seem super usable at the moment.

That being said, something is definitely coming. 50% correctness would probably be well worth using — simple copy/paste between my editor and GPT4 has been useful for me, and that's much less likely to completely solve an issue in one shot — and not only will small startups doing finetunes be grinding towards better results... The big labs will be too, and releasing improved foundation models that the startups can then continue finetuning. I don't think a new AI winter is on the horizon yet; Meta has plenty of reason to keep pushing out better stuff, both from a product perspective (glasses) and from an efficiency perspective (internal codegen), and OpenAI doesn't seem particularly at risk of stopping since Microsoft is using them both to batter Google on search (by having more people use ChatGPT for general question answering than using Google search), and to claw marketshare from Amazon in their cloud offerings. Similarly, some AI products have already found product/market fit; Midjourney bootstrapped from 0 to $200MM ARR (!) for example, purely on the basis of monthly subscriptions, by disrupting the stock image industry pretty convincingly.

Re: Devin: AI Software Engineer

#378

Earlier quoted context omitted.

I like it for condensing long stack traces and very very simple requests, but it really falters when you try to do anything domain specific. Library documentation? Yeah, it doesn't really save time when GPT makes up functions and libraries, making me check the docs anyways... I was initially hopeful but I find it gets in my way for anything not trivial.

Yeah, it doesn't really save time when GPT makes up functions and libraries, making me check the docs anyways... That behavior is now vanishingly-rare, at least in GPT4.

So Copilot uses GPT-4 under the hood, and about half the time I use it to generate anything bigger than a couple of lines it doesn't even compile, let alone be correct. It hallucinates constantly.

Re: Devin: AI Software Engineer

#379
post #218

Earlier quoted context omitted.

You really need to try Opus. Try a provider that works across models (one in my bio).

It's incredible how far behind HN of all places is w.r.t. what the current best tech is. So many people talking about GPT-4 here, or even 3.5 when the SOTA has moved way along. Gemini Advanced is also a great model, but for other reasons. That thing really knows a boat load of low level optimization tricks.

> It's incredible how far behind HN of all places is w.r.t. what the current best tech is.

Good fucking Christ, Opus has been out for like 8 days, I've had holidays longer than that!

Re: Devin: AI Software Engineer

#380

I must say, I'm not HUGELY impressed with a website that lets me, unauthenticated, upload files of an arbitrary size. Just posted a 500mb dmg file to their server. If anyone is practicing for their B1 Dutch exam, feel free to use this link to get the practice paper. https://usacognition--serve-s3-files.modal.run/attachments/4...

Looks like they deleted it and restricted file uploads :/
Post reply on HN