Earlier quoted context omitted.
Many claims but no clear evidence that they actually find significantly more severe issues compared to open models.
Open models and agents can't be trusted without handholding. Astra can one-shot six months of work. Years of work, even. OpenAI just solved Navier-Stokes. Seems like the US is on a takeoff ramp to me.
Astra can confidently one-shot 500k lines of slop, with 800k lines of tests covering it, without testing a single intended product requirement, and none of it actually working.
All models require hand holding. Fable and Astra are no exceptions. The difference is only in the amount of hand holding required, and there's essentially no gap here anymore between American and Chinese models.
I only use Chinese models sparingly because American models are so much cheaper with subscriptions, that it doesn't make economic sense to not use them. If/when that changes, I could simply route to cheapest model that's available at the moment and I wouldn't notice much difference in most applications.