This is going so fast! What a time to be on hackernews: July 16th: The "Kimi K3 moment" - China has caught up to Opus! 4 weeks later: GLM 5.3 - Same performance, but cut the amount of parameters and cost to a third! 12 days later: GLM 5.3 Flash - Almost GLM5.3 performance but cut the parameters in half, cut prices to a fifth and serving on Chinese chips!
Why are these models able to reduce parameters but keep quality? I know the original intuition was scale data + params = quality but it looks like we have hit an s curve on improvements from pure scaling? Is this just because we are in a memory / data crunch? Are we learning how LLMs learn and effectively training better? Do we have a way to derive the amount of intelligence an LLM will have based on size / training…
No? A large model obviously can be dumb, I don't think you can infer much other than by testing it.
These small models are almost certainly worse at some things than the big models. They prize is making them dumber at things no-one cares about while retaining the capabilities people do care about. A model probably does not need to be able to give me a political treatise on the late 19th century "silver question" to be able to write me code.