Live data from Hacker News

Antigravity 2.0 Tops the OpenSCAD Architectural 3D LLM Benchmark

modelrift.com

61–70 of 171 posts

Re: Antigravity 2.0 Tops the OpenSCAD Architectural 3D LLM Benchmark

#61
post #29

Claude Code 2.1 / Opus 4.7 looks best to me: Dome and ceiling structure is correcter than the others. Why is this medium ranked, and not on par with the best two?

Look at a picture of the Pantheon, the dome isn't as dome-like as you would imagine. It's more like a hump shape.

Re: Antigravity 2.0 Tops the OpenSCAD Architectural 3D LLM Benchmark

#63

The only thing faster moving that AI these days are the goalposts. Three years ago we would have been amazed if models were able to produce anything, now we have the luxury of nitpicking. Even the worst entries in the benchmark are quite impressive.

I remember getting wound up about latency and server issues playing counter-strike in the early '00s. At the same time though, it was hard to justify being angry because playing a multiplayer game with friends who were scattered all over town was something that had to be real magic.

I guess the wow!->adjust->complain->wow!->... cycle is endless as a human

Re: Antigravity 2.0 Tops the OpenSCAD Architectural 3D LLM Benchmark

#65

Earlier quoted context omitted.

I'm not GP, but I am somewhat excited about antigravity CLI. I adopted Gemini CLI early and really liked it, though over time it got dumber and dumber until a point when I realized it was foolish to use it instead of claude/codex. I'm hopefuly that antigravity CLI won't go through that path, but also can't fight a skepticism.

I don’t think it’s the cli that was dumber, just the model it was using. They drastically reduced limits on their best model so that’s likely how you got stuck downgrading model and getting worse results.

I'm sensing in reality that behind the scenes there is a difficult trade-off between quantization and usage limits. You can have a "smart" model but poor limits, or good limits and a "dumb" model.

This seems very similar to mobile data limits (remember those years?), where there wasn't enough tower bandwidth to serve everyone unlimited data, so telecos were in constant tension between data caps and bandwidth throttling.

It wasn't until 5G came along with 100x network capacity that they could finally give everyone "unlimited" data.

Re: Antigravity 2.0 Tops the OpenSCAD Architectural 3D LLM Benchmark

#67
It's crazy how I can see articles like this, but in my practical every day use antigravity is a horrible consumer experience. The TUI is broken. You cannot type input while the model is outputting text, otherwise both get messed up and the the TUI renders a sickly blob of text. There are no keyboard shortcuts to switch between planning and execution mode, or a way to directly load skills.

The usage limits are too aggressive, too. I tried to generate a quick Deno Fresh website to act as a a redirect to my GitHub from socials (literally the simplest possible thing I could have asked of it) and it chewed through my five hour limit in tokens from scaffolding.

To me, as a developer of CLI developer tooling, its obvious not a lot of thought or testing went into this product, but as Google has said before: the models are the product".

Re: Antigravity 2.0 Tops the OpenSCAD Architectural 3D LLM Benchmark

#69
I've run a tons of benchmarks for OpenSCAD for all kinds of models and setups, and what I realised is:

- Models are very jagged (might excel in one type of 3d model, but not another)

- Gemini models are the least jagged in my experience and have the best image understanding

- Gemini models are also the most creative (which may be undesirable if you want precise CAD part)

- Overall this benchmark doesn't prove much because one 3d model (and one attempt) is just not enough. I am usually testing on at least a dozen models each generated 3 times, but should really do much more, but it's too pricey for a solo dev.

Still, thanks for publishing this. Will be definitely run flash 3.5 soon to see how it performs.

Re: Antigravity 2.0 Tops the OpenSCAD Architectural 3D LLM Benchmark

#70

Creating a single real-world object and declaring it a benchmark? No, it doesn't work that way for a robust tool. You need to do something like Iron Chef, with a Greek architecture theme and and a panel or judge that declares the winner. This is just seeing which tool subjectively makes the best looking Pantheon.

Yeah, this is less of a benchmark and more "I like this one guys!".

Just totally subjective grading criteria of a single poorly defined example with no end use case in mind to guide how to even do evaluation.

Post reply on HN