Live data from Hacker News

Recreating Minecraft Is Not a Benchmark

kuber.studio

21–30 of 81 posts

Re: Recreating Minecraft Is Not a Benchmark

#21
This sounds unconvincing, because a) pelican test is subjective, there's simply nothing to leak as it has no available direct answers and maybe an extremely faint preference signal, and b) the same small models actually do perform well when you change the subject. Some models are genuinely trained to be better at some domain, such as 2D layouts or vector graphics in this case. It all depends on particular recipes and datasets. Which is the actual reason these tests are poor as vibe checks: they don't do anything to disentangle generalization, memorization, and training preference. One-shotting popular software in particular is definitely not a good test of anything as memorization is going to dominate it.

AAII is also not very useful, neither is any generic score/benchmark. If you want a weather forecast you aren't looking at the average temperature of Earth.

(actually when did the term "one-shot" get hijacked to mean something other than "one example"?..)

Re: Recreating Minecraft Is Not a Benchmark

#22

This sounds unconvincing, because a) pelican test is subjective, there's simply nothing to leak as it has no available direct answers and maybe an extremely faint preference signal, and b) the same small models actually do perform well when you change the subject. Some models are genuinely trained to be better at some domain, such as 2D layouts or vector graphics in this case. It all depends on particular recipes and…

> Some models are genuinely trained to be better at some domain, such as 2D layouts or vector graphics in this case. It all depends on particular recipes and datasets.

I'd argue this is inherently true for every single model today, none of them have completely generalized to be able to solve any task, so whenever people come up with new evaluations and benchmarks, all the models score relatively poorly initially, until researchers start to tune the models to do well in the domains that the evaluations and benchmarks tests, and then we see strong improvements in that domain, which then tapers out to incremental improvements, and some other domain is chosen to be the new focus.

Models aren't better agents today merely by chance, but because it's explicitly part of the training data. They do well with software because we've talked so much about software on the internet until this point and that's part of the training data, but pit them against problems people don't talk so much about, and if the labs creating and training these models didn't consider those problems, then the model will pretty much suck at it.

I guess eventually they will literally cover every single task the model could ever come across, at least some variant/permutation of it, but until then every benchmark/evaluation will just uncover "did the labs consider this and who considered it most important before/during training?" basically.

Re: Recreating Minecraft Is Not a Benchmark

#23

I keep seeing Astra make beautiful 3d stuff online, yet when I feed it some old school RuneScape assets (even tried with some very detailed guidelines) and asked it to generate some new plausible assets it failed horribly. I think there’s still something really off with current (frontier) models when it comes to creating “novel” stuff? Even 2004 style graphics.. Or am promoting it wrong?

It's funny seeing this when I was going to mention runebench

https://maxbittker.github.io/runebench/

Re: Recreating Minecraft Is Not a Benchmark

#24

Most of my and my peers PortCos run their own eval and benchmark sets, simply because they know what they need best. The reality is, capabilities have largely converged across foundation models over the last 18 months, and much of the value add is coming from the harness layer itself now. This has been the operating assumption for me and my peers, and has largely played out that way. That said, this has always been a…

> capabilities have largely converged across foundation models over the last 18 months

For reference, in March '25 the models du jour were Sonnet 3.7, gpt o4 and gemini 2.5 pro. GPT5 was in august '25.

It's been a while since we've heard the old "models have stagnated". Oh well.

Re: Recreating Minecraft Is Not a Benchmark

#26

Most of my and my peers PortCos run their own eval and benchmark sets, simply because they know what they need best. The reality is, capabilities have largely converged across foundation models over the last 18 months, and much of the value add is coming from the harness layer itself now. This has been the operating assumption for me and my peers, and has largely played out that way. That said, this has always been a…

> capabilities have largely converged across foundation models over the last 18 months For reference, in March '25 the models du jour were Sonnet 3.7, gpt o4 and gemini 2.5 pro. GPT5 was in august '25. It's been a while since we've heard the old "models have stagnated". Oh well.

It's not "models have stagnated" but "models released at the same time are on the same level". Improvements are still real but the relative gaps between OpenAI, Anthropic, Meta, Grok, Gemini and open models are closer than ever. That doesn't mean progress is slowing down, it's just more widely distributed.

Re: Recreating Minecraft Is Not a Benchmark

#27
post #26

Earlier quoted context omitted.

> capabilities have largely converged across foundation models over the last 18 months For reference, in March '25 the models du jour were Sonnet 3.7, gpt o4 and gemini 2.5 pro. GPT5 was in august '25. It's been a while since we've heard the old "models have stagnated". Oh well.

It's not "models have stagnated" but "models released at the same time are on the same level". Improvements are still real but the relative gaps between OpenAI, Anthropic, Meta, Grok, Gemini and open models are closer than ever. That doesn't mean progress is slowing down, it's just more widely distributed.

This, and depending on the workflow and usecase, you don't necessarily need the latest and greatest with the right kind of harness engineering.

Like everything in engineering, it's about tradeoffs and what works best for your specific problem.

Re: Recreating Minecraft Is Not a Benchmark

#28
post #26

Earlier quoted context omitted.

> capabilities have largely converged across foundation models over the last 18 months For reference, in March '25 the models du jour were Sonnet 3.7, gpt o4 and gemini 2.5 pro. GPT5 was in august '25. It's been a while since we've heard the old "models have stagnated". Oh well.

It's not "models have stagnated" but "models released at the same time are on the same level". Improvements are still real but the relative gaps between OpenAI, Anthropic, Meta, Grok, Gemini and open models are closer than ever. That doesn't mean progress is slowing down, it's just more widely distributed.

Ah, I see. I misunderstood then. The thing about "gains come from the harness" made me think about it in that way.

Re: Recreating Minecraft Is Not a Benchmark

#29

I keep seeing Astra make beautiful 3d stuff online, yet when I feed it some old school RuneScape assets (even tried with some very detailed guidelines) and asked it to generate some new plausible assets it failed horribly. I think there’s still something really off with current (frontier) models when it comes to creating “novel” stuff? Even 2004 style graphics.. Or am promoting it wrong?

[dead]

Re: Recreating Minecraft Is Not a Benchmark

#30
I use Minecraft (and Warcraft and an arcade flying simulator) as a silly but directionally correct indication of the models' capabilitites.

Compare Astra[0] with GPT 5.4[1] which was OpenAI's state of the art just six months ago.

(all tests on more models with code and prompts available here: https://senko.net/vibecode-bench )

Yes, it's not a scientific benchmark but it's a good heuristic.

For a better eval, create a one-page prompt / mini spec related to whatever you're using the LLMs for, and see how well a particular one works for what's important to you.

0: https://senko.net/vibecode-bench/2026/rts-gpt-6-astra.html

1: https://senko.net/vibecode-bench/2026/rts-gpt-5.4.html

Post reply on HN