Live data from Hacker News

Top model scores may be skewed by Git history leaks in SWE-bench

github.com

161–167 of 167 posts

Re: Top model scores may be skewed by Git history leaks in SWE-bench

#161

Earlier quoted context omitted.

I challenge you! Try giving this exact prompt to GPT-5-Thinking (medium or high reasoning if API). It is able to (without external code tools) solve a never before seen cypher that is not present in its training data. I think this pretty clearly demonstrates that the “stochastic parrot” is no longer an apt description of its capabilities in generalization: ———— You are given a character-by-character decode table `map…

That's exactly the sort of thing a "stochastic parrot" would excel at. This could easily serve as a textbook example of the attention mechanism.

How about this alternative challenge: ask it to write a poem in IPA (pronunciation language). I’d be surprised if this has ever been done pre-LLM, yet it excels at weird tasks like this.

You could probably just ask it to come up with 100 tasks to prove it’s not a stochastic parrot.

Re: Top model scores may be skewed by Git history leaks in SWE-bench

#162

Everyone on HN is like “yes I knew it! I was so right in 2021 that LLMs were just stochastic parrots!” Strangely one of the most predictable groups of people

This reads like you're ridiculing people for being proved right?

No the point of the comment is that there is no meaningful difference between model performance improvements from before and after this news of a benchmark weakness (spoiler alert, almost all of the benchmarks contain serious problems). The models are improving every quarter whether HN likes it or not.

Re: Top model scores may be skewed by Git history leaks in SWE-bench

#163
post #128

Earlier quoted context omitted.

As others pointed out this problem isn't special. Grok 4 heavy Thought for 4m 17s {"decoded_prefix": "nqxznhzhvqvvjddqiterrqdboctzzmoxmhyzlcfe", "last_10": "kfohgkrkoj", "vowel_counts": {"a": 7, "e": 18, "i": 7, "o": 12, "u": 6}} it did count another e, but that's a known point of failure for LLMs which i assume you put in intentionally. >Counting e's shows at least 10 more, so total e's are 17.

I guess GPT-5 with thinking is still a bit ahead of grok. I wonder what the secret sauce is.

[deleted]

Re: Top model scores may be skewed by Git history leaks in SWE-bench

#164
post #85

Earlier quoted context omitted.

> I didn't mean to imply bad intent > I wouldn't be surprised if they left this loophole on purpose You didn't imply bad intent, you outright suggested it.

He means he doesn't say it was necessarily bad intent, but mentions it as a possibility ("thinking out loud").

Thinking out loud isn't a free pass to say stuff without consequences. Sure we are all protected under free speech, but free speech doesn't remove the meaning and the impact words have in the world.

Re: Top model scores may be skewed by Git history leaks in SWE-bench

#165

Earlier quoted context omitted.

That's exactly the sort of thing a "stochastic parrot" would excel at. This could easily serve as a textbook example of the attention mechanism.

How about this alternative challenge: ask it to write a poem in IPA (pronunciation language). I’d be surprised if this has ever been done pre-LLM, yet it excels at weird tasks like this. You could probably just ask it to come up with 100 tasks to prove it’s not a stochastic parrot.

Yeah, I think "stochastic parrot" is a crappy phrase that obscures the mechanics of the LLM. Of course the LLM is capable of producing novel outputs, for some definition of novel. My only position here is that we can take any apparently magical outputs of the thing and, based on an understanding of how LLMs work, understand how they were likely produced. I think that sort of literacy will take us a long way.

Re: Top model scores may be skewed by Git history leaks in SWE-bench

#166

Earlier quoted context omitted.

I meant it as a hint for anyone inclined to dig deeper. It's a possibility rather than something we can confidently dismiss.

If it's a possibility and you don't want to dig deeper better to sit out and not comment anything at all, lest you risk defamation. Thinking out loud also doesn't make defamation acceptable.

[deleted]

Re: Top model scores may be skewed by Git history leaks in SWE-bench

#167

Earlier quoted context omitted.

It's fine, this is an american site so JAQing is in fact safe under free speech. You're welcome to ask b "would none rid me of this meddlesome priest" with no fear

And I'm protected under free speech to try to educate people about good manners, so it's fine too.

[deleted]
Post reply on HN