GPT-6 Astra on OpenRouter
141–150 of 256 posts
Re: GPT-6 Astra on OpenRouter
#142Earlier quoted context omitted.
Here's 'Generate an SVG of a ring-tailed lemur riding an electric scooter' at reasoning level max: https://tools.simonwillison.net/markdown-svg-renderer?url=ht... Quote from the thinking trace: > I’m thinking about how a helmet would obscure lemur ears, but using an electric scooter helmet seems responsible. It's pretty solid - face is a little wonky but excellent tail and scooter.
If you want to see something brutal, ask for a zebra riding a scooter. I have yet to see any models do a credible job at that.
Re: GPT-6 Astra on OpenRouter
#143Any tips on using Astra as orchestrator with Luna workers efficiently in codex?
I've created with Sol a skill called Low Quota Mode that intends to reduce the use of tokens usages by the frontier (intelligent model) and delegate the use of bulk reading of docs/code and implementation to a sub-agent running Luna Max. Sol is asked to supervise, read the diffs and approves the commit/pr.
The skill might need some iterations while you use it, for example at the end of a rough session you can ask Sol how did it went, which were the points of conflict with Luna and try to iron them little by little by editing the skill.
Also in difficult tasks, ask to babysit the sub-agent model, I've seen it makes more effort into communication between frontier and sub-agent to guide the task with more care.
So far it has reduced my tokens usage a lot (have not quantified but the quota lasts more).
Re: GPT-6 Astra on OpenRouter
#144That's some crazy SVG generation: https://aibenchy.com/compare/openai-gpt-6-astra-high/google-... It took a while to test it, initially OpenRouter was giving Not Found errors for this model ID.
At the bottom it says "Score 98.58", what measure is used for this score? It's kind of horrible, the perspective is all off (legs of the table makes that very obvious), the mouse/hamster has two mouths, a stub for a right paw, looks like left hand holds a melon on a stick or something, and there are pluses in the background for some reason. Not sure it'd call it "close to perfect" which the score seems to want to ind…
Re: GPT-6 Astra on OpenRouter
#145$10/$50 is incredibly expensive compared to Chinese models which are cents. I think they’re really going to struggle selling these models long-term. My company is already massively cutting down on access because they’ve realised most people don’t actually produce any value using it. All the tokenmaxers have ruined it for the rest of us now that accounting have seen the costs.
Not really comparable IMO. Astra and Fable are not the every day workhorse you reach for to do basic tasks (unless your company has fuck you-money), they’re the tool you break out when you need the absolute strongest performance. There are plenty of tasks where finding and fixing one or two extra edge cases saves the business a lot of money, even if the cost is high. The best example would be scanning for vulnerabili…
Re: GPT-6 Astra on OpenRouter
#146Earlier quoted context omitted.
It is (uses way less tokens)
No it isn't cheaper, any source for task vs price comparison to Sol? Edit: GPT-6 Astra (low): 57 Intelligence Index, $7.70/M tokens GPT-5.6 Sol (high): 57 Intelligence Index, $3.08/M tokens So for the same measured intelligence, Sol costs only 40% as much — i.e. ~60% cheaper, while Astra is ~2.5× more expensive. Why is the burden of proof on me tho!?
Re: GPT-6 Astra on OpenRouter
#147I posted this in the other Astra thread but it's just fallen off the homepage, so... Pelicans from Astra, plus 5.6 Sol, Terra, Luna for comparison: https://static.simonwillison.net/static/2026/gpt-6-and-5.6-p... I think this is a genuinely interesting comparison grid. Astra may be more expensive, but if you have a budget of 10 cents for a Pelican Astra low gives you something SO much better than the other models. Ast…
Oh my goodness. I was not prepared for Luna on "none". Reminded me of https://clocks.brianmoore.com/
Re: GPT-6 Astra on OpenRouter
#148Earlier quoted context omitted.
At the bottom it says "Score 98.58", what measure is used for this score? It's kind of horrible, the perspective is all off (legs of the table makes that very obvious), the mouse/hamster has two mouths, a stub for a right paw, looks like left hand holds a melon on a stick or something, and there are pluses in the background for some reason. Not sure it'd call it "close to perfect" which the score seems to want to ind…
I've replaced "Score" there with model ranking, to reduce confusion, thanks for the feedback!
What panel of judges are you using for scoring/ranking this? Seems subjective enough to not be able to be ranked/scored at all
Re: GPT-6 Astra on OpenRouter
#149Earlier quoted context omitted.
At the bottom it says "Score 98.58", what measure is used for this score? It's kind of horrible, the perspective is all off (legs of the table makes that very obvious), the mouse/hamster has two mouths, a stub for a right paw, looks like left hand holds a melon on a stick or something, and there are pluses in the background for some reason. Not sure it'd call it "close to perfect" which the score seems to want to ind…
And not to mention, the hamster is standing at the wrong end of the table.
Re: GPT-6 Astra on OpenRouter
#150Earlier quoted context omitted.
That's quite shocking, at a sufficiently advanced level we can make all non-realistic graphics purely out of SVGs, as they'd have good scaling for things like logos and app icons. I know it was technically and theoretically possible before AI but most people weren't spending hours tweaking SVG HTML. I remember making an SVG dark mode toggle icon and it took days to get it right, I assume it's one shottable now.
I thought all the designers have been using vector graphics for a long time now.