Live data from Hacker News

ProgramBench: Can language models rebuild programs from scratch?

arxiv.org

11–20 of 86 posts

Re: ProgramBench: Can language models rebuild programs from scratch?

#11
post #5

I am not surprised but this one sticks out... > Models favor monolithic, single-file implementations that diverge sharply from human-written code. Well, all of our code is monolithic with some files close 20K lines of code and we do use coding agents - not for the original code but as of late. I've always had that hunch that splitting everything into tiny files does not improve AI coding agent performance although it…

Kinda surprising to me, since i had some trouble with Cursor & Co. once the file went over ~800 lines. It repeatedly failed to edit it, until i split it up into multiple logical components. As it should have been from the beginning...

Though, it was some time ago, so things might have improved?

Re: ProgramBench: Can language models rebuild programs from scratch?

#12
It’s unfortunate that they didn’t eval using subagents/orchestration for such a complex set of tasks (from what I can tell), e.g. analyze program to produce initial spec -> code -> review and rinse&repeat with each of those steps being a separate subagent allocated

I would be interested to see if there’s a significant quantifiable difference.

Re: ProgramBench: Can language models rebuild programs from scratch?

#14
post #5

I am not surprised but this one sticks out... > Models favor monolithic, single-file implementations that diverge sharply from human-written code. Well, all of our code is monolithic with some files close 20K lines of code and we do use coding agents - not for the original code but as of late. I've always had that hunch that splitting everything into tiny files does not improve AI coding agent performance although it…

Kinda surprising to me, since i had some trouble with Cursor & Co. once the file went over ~800 lines. It repeatedly failed to edit it, until i split it up into multiple logical components. As it should have been from the beginning... Though, it was some time ago, so things might have improved?

VSCode basically any model can edit the 20K file without any issues. The coding harness does not read the entire file at once though. It reads chunks of it so the size does not really matter. What matters is how close are the things the agent needs to make the edit.

Re: ProgramBench: Can language models rebuild programs from scratch?

#15
post #5

I am not surprised but this one sticks out... > Models favor monolithic, single-file implementations that diverge sharply from human-written code. Well, all of our code is monolithic with some files close 20K lines of code and we do use coding agents - not for the original code but as of late. I've always had that hunch that splitting everything into tiny files does not improve AI coding agent performance although it…

> Scattering the implementation in various files all over the source tree

If you treat the source tree seriously, you can communicate a lot with how it is structured

Re: ProgramBench: Can language models rebuild programs from scratch?

#17

It’s unfortunate that they didn’t eval using subagents/orchestration for such a complex set of tasks (from what I can tell), e.g. analyze program to produce initial spec -> code -> review and rinse&repeat with each of those steps being a separate subagent allocated I would be interested to see if there’s a significant quantifiable difference.

This might actually be the whole value prop of this benchmark. Forget their initial scores, take open models (so we can be sure the base doesn't change), and test different combinations of harness + prompts + strategies + whatever memthing is popular today. See if the scores improve. Repeat.

Re: ProgramBench: Can language models rebuild programs from scratch?

#18

In before "but they did not use my agent swarm"

It’s the annoying thing about AI. If it works, the AI is magic. If it doesn’t work, you’re using it wrong.

So, would you change your view if someone else runs this bench w/ a different harness and gets better results?

Re: ProgramBench: Can language models rebuild programs from scratch?

#19

In before "but they did not use my agent swarm"

It’s the annoying thing about AI. If it works, the AI is magic. If it doesn’t work, you’re using it wrong.

It was the same thing with OOP, TDD, agile development, C, C++, Rust, ORMs..

Whenever something impacts a ton of people you will get some who gain a lot from it and some who don't, and they're generally unable to relate to the other side.

Maybe the thing works in some domain and not the other. Maybe the two groups are doing different things. Maybe the context around it is different. Maybe they have a different definition of "better".

I think it helps to keep an open mind and not grow attached to either position, but rather inquire, "well we did X with outcome Y, what did you do instead?"

Re: ProgramBench: Can language models rebuild programs from scratch?

#20
post #15
post #5

I am not surprised but this one sticks out... > Models favor monolithic, single-file implementations that diverge sharply from human-written code. Well, all of our code is monolithic with some files close 20K lines of code and we do use coding agents - not for the original code but as of late. I've always had that hunch that splitting everything into tiny files does not improve AI coding agent performance although it…

> Scattering the implementation in various files all over the source tree If you treat the source tree seriously, you can communicate a lot with how it is structured

Well you can communicate organisation structure but not logic or intent. The directory is a tree and the Code is a graph.

You can communicate some information by looking at the org chart of a company but it does not really tell you much how it works.

Arguably a coding agent is less concerned about where the files are at then the code itself.

Post reply on HN