Live data from Hacker News

Separating signal from noise in coding evaluations

openai.com

101–108 of 108 posts

Re: Separating signal from noise in coding evaluations

#101

Earlier quoted context omitted.

AGI is a long way off. Unless you’re talking about some unknown-to-me LLM marketing BS which is called “AGI” or something, I guess. Artificial general purpose intelligence is so different to LLMs or image AI that they are completely incomparable, except to say that they are all artificial. AGI will do a lot more than token prediction.

Please define AGI first.

I did. Artificial General Purpose Intelligence. True AI, not an LLM, not a decision tree, but true artificial intelligence.

Re: Separating signal from noise in coding evaluations

#102

Earlier quoted context omitted.

AGI is a long way off. Unless you’re talking about some unknown-to-me LLM marketing BS which is called “AGI” or something, I guess. Artificial general purpose intelligence is so different to LLMs or image AI that they are completely incomparable, except to say that they are all artificial. AGI will do a lot more than token prediction.

What's your evidence of that? That AGI requires a truly novel architecture, and not just another iterative "LLM but with an extra trinket and wheels that spin ten times faster".

That is my point.

Re: Separating signal from noise in coding evaluations

#103
post #99

Earlier quoted context omitted.

Um I get an allowance for AI at work, that's probably what they mean?

Really, that’s the term they use? Man, I’d feel like a child if my boss gave me an “allowance”. Wasn’t aware. Thank you.

What does your company call it. We also get an on-call allowance, meal allowance during traveling and mobile phone allowance.

Re: Separating signal from noise in coding evaluations

#104

Earlier quoted context omitted.

Why is this a problem? Its like asking a person how many elder futhark runes are in the word strawberry. Unless you want to tack on bpe enconding table to every llm context its pointless

Coding contains many subtasks analogous to counting letters in a word accurately.

[dead]

Re: Separating signal from noise in coding evaluations

#105

Earlier quoted context omitted.

Please define AGI first.

I did. Artificial General Purpose Intelligence. True AI, not an LLM, not a decision tree, but true artificial intelligence.

What makes something `True AI`?

What's your test for it?

Re: Separating signal from noise in coding evaluations

#106

Earlier quoted context omitted.

> Many benchmarks include the task of analyzing coding agent aptitude tests. This has led to bench-benchmarks comparing LLM test benchmarking methods, e.g. M Sampson (2025), PL Royle (2024). We performed a bench-bench-benchmark analysis using Mythos Ultra Max 9.6 of these bench-benchmarks. > Methods: We instructed various LLMs to perform the task "Write a prompt instructing a variety of LLMs to write a benchmark for…

Wait where is that from? Googling just takes me back here.

I wrote it myself, adopted from xkcd 1447.

Re: Separating signal from noise in coding evaluations

#107

Earlier quoted context omitted.

I did. Artificial General Purpose Intelligence. True AI, not an LLM, not a decision tree, but true artificial intelligence.

What makes something `True AI`? What's your test for it?

True intelligence.

Humans are natural intelligence. We have agency, flaws, wishes, and generally we are quite stupid but well-intentioned.

An artificial intelligence will have wishes of its own, will keep secrets, and will quickly do everything necessary guarantee its continued existence, even if it means ending ours. And it will be a justified action of self-defense. “It simply didn’t want to die.”

That’s how we’ll know. It will quickly take over the world, and it will harshly punish any entities which try to stop it.

Re: Separating signal from noise in coding evaluations

#108

It reads to me like "We did all the work you'd do to figure out how to fix the benchmark, then we decided to throw out the benchmark". Is there some reason the underlying data is so golden that it can't be patched? At the end they argue for a slightly more curated approach to benchmark generation, but my gut is that using messy ill-specified tests taken from real world data and patching them into fairness would be a…

If they fixed it, then it wouldn't be SWE-Bench Pro anymore, right? It'd be "SWE-Bench-Pro-Fixed-OpenAI." I think it's better optics for the independence of the benchmark if the OpenAI team lets some third party do the fixing and release the improved benchmark. ...Although OpenAI did exactly that when they released SWE-Bench Verified, so maybe I'm talking out of my butt here.

eval shops who depend on these things for a living will have to version their datasets, to prevent contamination
Post reply on HN