I've been testing Sol/Terra/Luna now since yesterday, running complex evals on all of them and I feel a bit... mixed on how they perform. The eval is an agent that runs a set of tools and a prompt we can tune separately for different models. The OpenAI version of the prompt was specifically tuned based on their guide[0]. Then we let Opus to run another agent that acts as a user, trying to solve a problem (anonymized…
Yes, this is exactly why I said yesterday that OpenAI does benchmaxxing, seemingly quite a bit [1]. I got a flurry of downvotes for it at first, but I think people came around to it once they tried the model like you did. Ultimately I think the issue is that OpenAI is under tremendous pressure to perform, but GPT-6 is not ready yet, so they had to push GPT-5 to its limits, and the only way they could do it was with r…
Now if you look at where GLM stands, or even DeepSeek v4 Flash, things get really interesting for what they provide.
US labs completely miss a cheap model that can solve problems for 95% of the people. Gemini 3.5 Flash could've been it if it didn't burn so many tokens.