Earlier quoted context omitted.
open models are ahead in speed. they complete tasks as fast as you choose to scale compute. they are more efficient and require less compute for the same thing. they are ahead in specialized tasks. they are ahead in areas closed models refuse to answer. they are ahead in emotional intelligence.
I doubt most of your claims. Maybe the guardrails and emotional intelligence is true. For speed and efficiency, you are most likely wrong. Speed is led by GPT-5.6 Sol on Cerebras Ultrafast at 750 t/s. Afaik you cannot serve a single DeepSeek Flash 4.1 stream at 750 t/s, plus the model is less intelligent as seen on newer benchmarks. I believe OpenAI and Anhropic are at the frontier of efficiency too. There were numer…
5.6 sol ultrafast on cerebras is 750tps, open models readily exceed this. just by using a smaller model cerebras serves qwen 3.8 27b at 1850tps. or even larger models, mimo 2.5 pro was served for a while at 1000tps. and so on. [https://inference-docs.cerebras.ai/models/choose-a-model]
the chinese ai companies have 10% of the total compute resources of the US ones. since the USA tries to stop them from buying nvidia gpus. they maxed out the efficiency.
deepseek v4.1 has engram architecture. it has 550b params instead of 5T+ for astra/fable. it has 8b active instead of potentially hundreds active for astra/fable.
compare input/output/cache: $0.15/$0.60/$0.003 for v4.1 to $10.00/$50.00/$1.00 for astra and $10.00/$50.00/$0.25 for fable.
astra cache reads are over 330 times more expensive.
at the artificial analysis 7:2:1 ratio, deepseek is $0.18/m, fable is $7.18/m, astra is $7.7/m.
but what about intelligence? AA would rate deepseek v4.1 at AA 40, astra is AA 53.
so it cost 4,180% more for 32% more intelligence.
they are serving that at over 250tps at baseten. to get close to that on astra API you are paying double the cost for fast mode.
so it is now 8456% more expensive for a similar speed and 32% more intelligence. 84 times more expensive.