Live data from Hacker News

A case study in testing with 100+ Claude agents in parallel

imbue.com

41–50 of 57 posts

Re: A case study in testing with 100+ Claude agents in parallel

#41

If this will be future of software in 20 years nobody will understand what the hell software actually does. If nobody will things will get to implode quickly.

Billions of years of evolution and we increasingly understand what the genome does. And that's about as random as it gets.

I think we'll be fine.

This feels more like Y2K panic than grounded in truth. Senior software engineers guide these systems effectively today without creating a mess. I'm sure in some years agents will fill the role of maintainability engineer too. We are not special or irreplaceable.

It's not like we won't be spending an incredible amount of energy to overcome issues with understandably and maintenance. The sheer economic forces will absolutely will this problem solved. It must be solved, because trillions of dollars urgently want it to be solved. That's evolutionary pressure if I've ever seen it.

Also, we ceremoniously ascribe too much value to the software we create. With the exception of a few places, almost all of it gets replaced before our careers are over. At the end of the day, business automation is value creation. It's not sacred. It has a finite life, and then it too dies.

The software artifact just needs to facilitate economic/interest flux long enough to be useful, then it can be replaced with something better or more relevant.

Re: A case study in testing with 100+ Claude agents in parallel

#42

Earlier quoted context omitted.

To an extend you are likely doing something wrong. I understand that the natural instinct is to correct the output when you see your agent doing something wrong. That is not productive. The instinct should be to tweak the agent to do it right. At this point I am almost not writing any code in an enterprise code base.

> The instinct should be to tweak the agent to do it right. I'm extremely doubtful of this. It doesn't save time to tell it "you have an error on line 19", because that's (often) just as much work as fixing the error. Likewise, saying "be careful and don't make mistakes" is not going to achieve anything. So how can you possibly tweak the agent to "do it right" reliably without human intervention? That's not even a so…

Are you seriously interested in the answer, or are you just mad?

I could give you some pointers, but will only type it out if there is a point

Re: A case study in testing with 100+ Claude agents in parallel

#43

Earlier quoted context omitted.

> The instinct should be to tweak the agent to do it right. I'm extremely doubtful of this. It doesn't save time to tell it "you have an error on line 19", because that's (often) just as much work as fixing the error. Likewise, saying "be careful and don't make mistakes" is not going to achieve anything. So how can you possibly tweak the agent to "do it right" reliably without human intervention? That's not even a so…

Are you seriously interested in the answer, or are you just mad? I could give you some pointers, but will only type it out if there is a point

Not GP, but I would love pointers on precisely this problem

Re: A case study in testing with 100+ Claude agents in parallel

#44

Me: has to babysit every feature for hours in Claude Code, building a good plan but then still iterating many many times over things that need to be fixed and tweaked until the feature can be called done. Bloggers: Here's how we use 3,000 parallel agents to write, test, and ship a new feature to production every 17 minutes in an 8M-LOC codebase (all agent-generated!). ... I'm doing something wrong, or other people ar…

To an extend you are likely doing something wrong. I understand that the natural instinct is to correct the output when you see your agent doing something wrong. That is not productive. The instinct should be to tweak the agent to do it right. At this point I am almost not writing any code in an enterprise code base.

> The instinct should be to tweak the agent to do it right.

Ah, yes; must always remember to add "And don't make any mistakes" into the prompt /s

Re: A case study in testing with 100+ Claude agents in parallel

#45

Earlier quoted context omitted.

No. Your data or any derivative of it does not leave RAM unless you are detected as doing something that qualifies as abuse, then it is retained for 30 days.

Even the process of deciding what "qualifies as abuse" does what I'm talking about: they're analyzing your data with their own models and doing whatever they want with the results, including storing it and using it to ban you from the product you paid for, and call the police on you. Either way, I don't believe it.

You are a Star Wars Rebel fighting Darth Vader. Good job!

Re: A case study in testing with 100+ Claude agents in parallel

#46
post #41

If this will be future of software in 20 years nobody will understand what the hell software actually does. If nobody will things will get to implode quickly.

Billions of years of evolution and we increasingly understand what the genome does. And that's about as random as it gets. I think we'll be fine. This feels more like Y2K panic than grounded in truth. Senior software engineers guide these systems effectively today without creating a mess. I'm sure in some years agents will fill the role of maintainability engineer too. We are not special or irreplaceable. It's not li…

I think we are talking about different timespans. I am talking about change in the world after decades of something like that happening. How those Senior Engineers will know how good software looks like if they would never write it themeselves? Imagine looking at someone driving a car for 20 years. Will it be enough for you to drive a car yourself?

Thinking about that always makes me think about Foundation, The Merchant Princess. Mallow travels to the edge of the Empire to look how things are on one of those worlds. He learns that there is the cast of the tech priests and those people have absolutely no idea how those devices actually work.

He said:

> The machines work from generation to generation automatically, and the caretakers are a hereditary caste who would be helpless if a single D-tube in all that vast structure burned out

It was a sign of severe decline of the entire empire. People had no idea how devices work and they would not be able to reproduce it or even repair if one would broke.

It was recurring premise of civilisation decline in the series: no proper maintaince and people loosing interests and knowledge how things are done and how they work.

I just wondering if this is not the same thing starting to happining know with our civilisation.

And evolution? Evolution means mass extinction of species and its normal. I am not sure about you but I would rather avoid any mass extinction regarding humanity.

Re: A case study in testing with 100+ Claude agents in parallel

#47

Earlier quoted context omitted.

Even the process of deciding what "qualifies as abuse" does what I'm talking about: they're analyzing your data with their own models and doing whatever they want with the results, including storing it and using it to ban you from the product you paid for, and call the police on you. Either way, I don't believe it.

You are a Star Wars Rebel fighting Darth Vader. Good job!

Thanks

Re: A case study in testing with 100+ Claude agents in parallel

#48
post #43

Earlier quoted context omitted.

Are you seriously interested in the answer, or are you just mad? I could give you some pointers, but will only type it out if there is a point

Not GP, but I would love pointers on precisely this problem

It is about tweaking inline documentation to make sure that

1. It is not ambiguous 2. It is as complete as possible.

I am surprised that I got down voted for proposing the improve a code base such that agents can run on it as a means to increased productivity.

Re: A case study in testing with 100+ Claude agents in parallel

#49

Earlier quoted context omitted.

To an extend you are likely doing something wrong. I understand that the natural instinct is to correct the output when you see your agent doing something wrong. That is not productive. The instinct should be to tweak the agent to do it right. At this point I am almost not writing any code in an enterprise code base.

> The instinct should be to tweak the agent to do it right. Ah, yes; must always remember to add "And don't make any mistakes" into the prompt /s

I am not entirely sure what you are referring to.

Improving the agent means improving the code base such that the agent can effectively work on it.

It can not Com as a surprise that an agent is better at working on a well documented code base with clear architecture.

On the other hand, if you expect that an agent can add the right amount of ketchup to your undocumented speghatti code, then you will continue to have a bad time.

Re: A case study in testing with 100+ Claude agents in parallel

#50

If this will be future of software in 20 years nobody will understand what the hell software actually does. If nobody will things will get to implode quickly.

(author here) I agree that it's super important to understand what software actually does--that's part of the whole reason we made mngr in the first place!

I believe we can use these types of tools to make software more understandable, and mngr is an example of how to do that.

In our case study, we're using AI to increase our test coverage, and if you look at it, I would argue that we are making it more understandable--now instead of just having 100's of tests, we simply have a document that describes how the software is supposed to work, and the tests are linked to that document, and checked to ensure that they conform.

That means that anyone--not just the author of the software--is now able to read through the high level tutorial description of how the commands work in order to understand what the program should do!

And as for the tests themselves, we've been able to make nice testing infrastructure--like the transcripts and recordings that were highlighted in the post--to make it even easier for us to verify the behavior of the software.

We also have an incredibly detailed style guide and set of tests and guidelines to ensure that the entire code base is consistent, and high quality. You can drop into any of the code and pretty quickly understand what is happening. And if not, claude will do an excellent job of describing how any given component works, and how it relates to the others.

Finally, mngr itself is designed to be fully transparent when it is running--you can literally attach to the coding agent you are running and see exactly what is happening, and the program makes extensive log outputs for everything it does (feel free to open a PR if you'd like to see more!)

It's not perfect formal verification, but it does feel like we're making meaningful progress on making it easier to understand software--not harder.

Post reply on HN