Live data from Hacker News

Devin: AI Software Engineer

cognition-labs.com

521–530 of 604 posts

Re: Devin: AI Software Engineer

#521
post #249
post #238

From their twitter: > When evaluated on the SWE-Bench benchmark, which asks an AI to resolve GitHub issues found in real-world open-source projects, Devin correctly resolves 13.86% of the issues unassisted, far exceeding the previous state-of-the-art model performance of 1.96% unassisted and 4.80% assisted. While it is a progress, its far away from being useful to be a software engineer.

13% unassisted is crazy. That's probably half the performance of an intern that costs ~100k/year.

But how do you know if you are in the remaining 87%?

Re: Devin: AI Software Engineer

#522
post #350

Earlier quoted context omitted.

> It's something a clever fourth-grader would write. This level of cope and denial is amazing to witness. The most powerful (multi trillion dollar) companies on the planet are pouring practically infinite resources into developing systems that will ultimately make you redundant. An early version of AGI is staring you in the face while you call it a "fourth-grader". It won't stay in fourth grade forever.

I don't think I'm particularly in denial about the prospects of AI. I think it's going to be hugely disruptive and could possibly put me out of a job. But I'd like to posit a hypothetical counterpoint, just to get you thinking. So far, all of the work on AGI has been the result of brute forcing. We've tried to develop a structural understanding of how the human brain works, and we've failed. So we've fallen back to t…

Part of me hopes this is true, that AGI (or even worse - ASI) will never be fully realized. Too disruptive.

A counter example to nuclear power or space travel is integrated circuits. This technology has transformed our society and we haven't reached the end of it yet.

Our own brains are living proof that intelligence is possible with lower power consumption. I watched a recent lecture by Geoffrey Hinton where he mentioned future AI hardware based on analog integrated circuits could reduce the power consumption by orders of magnitude [1].

It is possible that we will hit a wall and never achieve anything more than Chat GPT++++, but the smartest people in town mostly believe that we will create machines that exceed human intelligence and capability.

We have some understanding of how neural networks work under the hood. The scale of the current models are too vast to comprehend in their specific details, but I think we understand them in principle.

[1] Prof. Geoffrey Hinton - "Will digital intelligence replace biological intelligence?" https://www.youtube.com/watch?v=N1TEjTeQeg0

Re: Devin: AI Software Engineer

#523

Earlier quoted context omitted.

> It's worth pointing out that on their eval set for "issues resolved" they are getting 13.86%. While visually this looks impressive compared to the others, anything that only really works 13.86% of the time, when the verification of the work takes nearly as much time as the work would have anyway, isn't useful. Yeah, I remember speech recognition taking decades to improve, and being more of a novelty - not useful at…

You can't compare the accuracy of speech recognition to LLM task completion rates. A nearly-there yet incomplete solution to a Github issue is still valuable to an engineer who knows how to debug it.

> A nearly-there yet incomplete solution to a Github issue is still valuable to an engineer who knows how to debug it.

Not sure if I can agree. There would definitely be a value in looking at what libraries the solution uses, but otherwise it may be easier to write it oneself, especially when the mistakes are not humanlike.

Re: Devin: AI Software Engineer

#524

Earlier quoted context omitted.

> It's worth pointing out that on their eval set for "issues resolved" they are getting 13.86%. While visually this looks impressive compared to the others, anything that only really works 13.86% of the time, when the verification of the work takes nearly as much time as the work would have anyway, isn't useful. Yeah, I remember speech recognition taking decades to improve, and being more of a novelty - not useful at…

You can't compare the accuracy of speech recognition to LLM task completion rates. A nearly-there yet incomplete solution to a Github issue is still valuable to an engineer who knows how to debug it.

Sure, and no doubt people paying for speech recognition 25 years ago were finding uses for it too. It depends on your use case.

A 13% success rate is both wildly impressive and also WAY below the level where I would personally find something like this useful. I can't even see reaching for a tool that I knew would fail 90% of the time, unless I was desperate and out of ideas.

Re: Devin: AI Software Engineer

#525

Earlier quoted context omitted.

>What I'm talking about is something you don't even seem to realize exists in professional writing I've read hundreds of books, fiction and otherwise. This isn't a brag, it's just to say, believe me, I know what professional writing looks like and I know where LLMs currently stand because I've used them a lot. I know the quality you can squeeze out if you're willing to let go of any presumptions. You'll notice that n…

>Rie Kudan won an award on a novel she used GPT to verbatim ghostwrite (no edits essentially) 5% of. Her words, not mine. Who knows how much more of the novel is edited GPT. That a professional human novelist was able to leverage GPT for their book isn't disproving the grandparent's post. They knew what good looks like, and if it wasn't good they wouldn't have kept it in the book.

That's my point. Good writing can come out of LLMs and nobody has to take my word for it.

When part of the OPs point seems to be that LLMs can't write good stuff then that's proof enough.

If you're talking about replacing professionals wholesale then I never made that argument.

Re: Devin: AI Software Engineer

#526
post #73

As a developer but also product person, I keep trying to use AI to code for me. I keep failing, because of context length, because of shit output from the model, because of lack of any kind of architecture etc etc etc. I'm probably dumb as hell, because I just can't get it to do anything remotely useful, more than helping me with leetcode. Just yesterday I tried to feed it a simple HTML page to extract a selector, I…

It's worth pointing out that on their eval set for "issues resolved" they are getting 13.86%. While visually this looks impressive compared to the others, anything that only really works 13.86% of the time, when the verification of the work takes nearly as much time as the work would have anyway, isn't useful. The problem with this entire space is that we have VC hype for work that should ultimately still be being do…

It’s like talking to zip file sometimes. Very difficult to make it actually do something you don’t expect it already to can do. Like a smarty index decorated with language.

Re: Devin: AI Software Engineer

#527

Earlier quoted context omitted.

The mechanism for this is charity. I understand people like to see social welfare programs, but I believe they are inferior to private sector charity because they ignore moral hazard, tend to produce worse outcomes, and are less economically efficient. At the core, welfare steals from all to provide provide for some. This introduces deadweight loss. I see a future where those displaced by automation are able to retra…

"At the core, welfare steals from all to provide for some." Yep, I want a system of progressive taxation that provides negative pressure on wealth inequality while uplifting the poorest. I like this idea because I like liberty, and a rich person loses almost no agency by having some portion of their income or wealth taxed while a small stipend can make a huge difference in the choices available to a poor person. Ther…

Overall I think we just fundamentally disagree about government's role as well as for what to optimize: individual or society.

> I like this idea because I like liberty, and a rich person loses almost no agency by having some portion of their income or wealth taxed while a small stipend can make a huge difference in the choices available to a poor person.

I also like liberty, but to me the opposite of liberty is achieved when you redistribute for non public goods. Although you may have elevated the poorest, you've dampened the richests' purchasing power, however marginal. It is not pareto optimal, and to me the atomic unit is the individual and not the collective society.

> Do you have any evidentiary basis for the counter? Why would the system of voluntary charity suddenly improve its outcomes or have more money to spend?

"Crowd-out was small as a share of total New Deal spending (3%), but large as a share of church spending: our estimates suggest that church spending fell by 30% in response to the New Deal, and that government relief spending can explain virtually all of the decline in charitable church activity observed between 1933 and 1939."

https://www.nber.org/papers/w11332

This is an interesting problem due to its nature: good evidence can really only be collected once policies are in effect. Regression discontinuities around the New Deal are likely good candidates to study. The above paper estimates a 30% drop in religious charitable giving. I didn't look to see if they are able to discern whether this is due to the perception that the poor now get money from the gov so don't need extra or whether the income tax cut disposable income and thusly the charitable contribution budget.

Another instance is the Texas Seed Bill, where Grover Cleveland vetoed the disaster relief bill after finding no power enumerated to the federal government to provide aid. Private donations exceeded the Congressionally approved sum (or so I've heard but cannot find a source atm).

https://en.m.wikipedia.org/wiki/Texas_Seed_Bill

The logical basis is the following: by taxing someone you remove their purchasing power and thus naturally cut their ability to provide charitable contributions. Keep in mind the level of welfare we are already providing far exceeds what the wealthiest provide in their 50%+ tax rates. A lot of burden comes from "well to do but vunerable to economic shocks" folk, for which taxation rates have material impacts.

> If social welfare causes a significant change in the average person's risk tolerance, a claim that definitely requires substantiation, then I still might argue this isn't a net negative for the economy or society at large.

This is exactly why I mention moral hazard. We have to ask ourselves what the unintended consequences or the policies are, and how they may be abused.

Risk tolerance shifts due to free money will broadly impact the economy: you will necessarily increase demand as those with jobs suddenly find the opportunity cost of not working preferable. So you either end up with more unemployment or needing to tax at higher rates to meet the new demand for welfare. This might lead to increased economic output from the poorest but also may lead to less output from the richest. One thing is certain: risk is adjusted due to government subsidiaries/theft.

> Perhaps a person who gets the aid of charity benefits more than one who gets welfare, but does the average person?

How will you measure this? Are you quantifying strictly by the total dollars received by the needy? What if you flipped the script and asked whether the outcome is congruent to the expectations of the givers? E.g., are those dollars given charitably providing a better outcome for the givers than through dollars given by the charitable via taxation?

The benefit of the charitable approach is that there can be conditionals. "Do drugs and the money stops" kind of stuff. Society tends to think that is immoral for the government to do, but arguably the outcomes for the needy would be better if conditions were allowed. Charity also has the added benefit of being able to flow to local needs vs flow from the top down.

> If we see unprecedented unemployment due to AI, how exactly do we expect voluntary charity to expand to meet demand?

What is AI doing that will completely eridicate humans? I struggle to understand. Let's say you're displaced from programming. Could you become a farmer? Would machines undercut your price? Probably. But what if you're good at metal work and Joe is good at farming, and both of you have lost your jobs to AI? Maybe you decide to consume from only human, non AI based businesses. Suddenly, costs aside, people create an underlying economic network of non AI businesses and start to thrive again. This is essentially what we see with people trying to buy only from [insert preferences here] businesses.

I don't think AI is the scare people make it out to be. Ultimately, it will be a fun thing to play out so long as regulation is minimized.

Re: Devin: AI Software Engineer

#528

Earlier quoted context omitted.

If the circle you're listening to are "recent grads" stop listening to them . They don't know how the industry works or what it needs, few of them them know how the technology works or what it's realistic near-term prospects are let alone the long-term ones, and none of them have lived long enough to understand the pace of a technology moving from discovery through to maturity and commercialization. If you look to yo…

Experienced, older people are notorious for downplaying and dismissing the implications of revolutionary new technologies. They may be right this time, or they may be wrong. But age and experience aren't as relevant when it comes to revolutions.

And inexperienced people are notorious for falling for unfounded hype. Eventually they either learn to discern value, which is hard, or just write off hype, which usually works fine but occasionally makes them late to the party.

Re: Devin: AI Software Engineer

#529

Earlier quoted context omitted.

Long ago most humans spent their time and energy on farming so they wouldn't starve. We automated that with a couple of Agricultural Revolutions. Then the majority of us spent our time and energy in factories. That got automated during the Industrial Revolution. So we moved to service and knowledge jobs. And now we're potentialy seeing the start of automation in those knowledge jobs. Will this be a painful revolution…

I'd love to live in a Communist society. I think our technological capabilities are to the point where a centrally managed economy is not only possible, but extremely doable. The main issue being, our governments and society have made it clear that the opinions of those who hold Capital are far more important than the unproductive surplus. Our entire society now is based around consuming cheap products we don't need,…

Thank you for the thoughtful post.

I hope that the current circumstances trigger a renewed interest in thinking about how we organise our society. It would be great if that could be done rationally, without taboos. I don't have high hopes though. The people in power want to stay in power, and any hint at socialistic or communistic ideas are quickly shot down.

Communism as implemented by Soviet Russia was pretty bad and failed, but that doesn't mean that the ideas of Karl Marx are all terrible.

I like the idea of a centrally managed economy but I don't think it's a stable system. There's too much opportunity for abuse by the central power.

I hope we explore the spectrum between extreme capitalism and extreme communism.

Does anyone have any reading recommendations that explore these kinds of economic and societal system?

Re: Devin: AI Software Engineer

#530

Earlier quoted context omitted.

> It's worth pointing out that on their eval set for "issues resolved" they are getting 13.86%. While visually this looks impressive compared to the others, anything that only really works 13.86% of the time, when the verification of the work takes nearly as much time as the work would have anyway, isn't useful. Yeah, I remember speech recognition taking decades to improve, and being more of a novelty - not useful at…

Even now, automatic speech recognition is a big timesaver, but you _need_ a human to look through the transcript to pick out the obviously wrong stuff, let alone the stuff that's wrong buy could be right in context.

but* just wanted to mention the error because the comment was about speech recognition errors
Post reply on HN