Claude Code 2.1 / Opus 4.7 looks best to me: Dome and ceiling structure is correcter than the others. Why is this medium ranked, and not on par with the best two?
Antigravity 2.0 Tops the OpenSCAD Architectural 3D LLM Benchmark
61–70 of 171 posts
Re: Antigravity 2.0 Tops the OpenSCAD Architectural 3D LLM Benchmark
#62Re: Antigravity 2.0 Tops the OpenSCAD Architectural 3D LLM Benchmark
#63The only thing faster moving that AI these days are the goalposts. Three years ago we would have been amazed if models were able to produce anything, now we have the luxury of nitpicking. Even the worst entries in the benchmark are quite impressive.
I guess the wow!->adjust->complain->wow!->... cycle is endless as a human
Re: Antigravity 2.0 Tops the OpenSCAD Architectural 3D LLM Benchmark
#64Re: Antigravity 2.0 Tops the OpenSCAD Architectural 3D LLM Benchmark
#65Earlier quoted context omitted.
I'm not GP, but I am somewhat excited about antigravity CLI. I adopted Gemini CLI early and really liked it, though over time it got dumber and dumber until a point when I realized it was foolish to use it instead of claude/codex. I'm hopefuly that antigravity CLI won't go through that path, but also can't fight a skepticism.
I don’t think it’s the cli that was dumber, just the model it was using. They drastically reduced limits on their best model so that’s likely how you got stuck downgrading model and getting worse results.
This seems very similar to mobile data limits (remember those years?), where there wasn't enough tower bandwidth to serve everyone unlimited data, so telecos were in constant tension between data caps and bandwidth throttling.
It wasn't until 5G came along with 100x network capacity that they could finally give everyone "unlimited" data.
Re: Antigravity 2.0 Tops the OpenSCAD Architectural 3D LLM Benchmark
#66Re: Antigravity 2.0 Tops the OpenSCAD Architectural 3D LLM Benchmark
#67The usage limits are too aggressive, too. I tried to generate a quick Deno Fresh website to act as a a redirect to my GitHub from socials (literally the simplest possible thing I could have asked of it) and it chewed through my five hour limit in tokens from scaffolding.
To me, as a developer of CLI developer tooling, its obvious not a lot of thought or testing went into this product, but as Google has said before: the models are the product".
Re: Antigravity 2.0 Tops the OpenSCAD Architectural 3D LLM Benchmark
#68Re: Antigravity 2.0 Tops the OpenSCAD Architectural 3D LLM Benchmark
#69- Models are very jagged (might excel in one type of 3d model, but not another)
- Gemini models are the least jagged in my experience and have the best image understanding
- Gemini models are also the most creative (which may be undesirable if you want precise CAD part)
- Overall this benchmark doesn't prove much because one 3d model (and one attempt) is just not enough. I am usually testing on at least a dozen models each generated 3 times, but should really do much more, but it's too pricey for a solo dev.
Still, thanks for publishing this. Will be definitely run flash 3.5 soon to see how it performs.
Re: Antigravity 2.0 Tops the OpenSCAD Architectural 3D LLM Benchmark
#70Creating a single real-world object and declaring it a benchmark? No, it doesn't work that way for a robust tool. You need to do something like Iron Chef, with a Greek architecture theme and and a panel or judge that declares the winner. This is just seeing which tool subjectively makes the best looking Pantheon.
Just totally subjective grading criteria of a single poorly defined example with no end use case in mind to guide how to even do evaluation.