Earlier quoted context omitted.
You have committed the classic blunder of confusing your abstraction layers. "Probabilistic next word prediction" and "humor" sit about as far apart as "modulating airflow with meat flaps" and "humor" do. One is an interface through which an action is performed and the other is a highly abstract capability. Would you claim that a podcast comedian is fundamentally incapable of being funny because all he ever does is w…
Humor requires a sudden orthogonal leap from context. That's what a punchline is. I think you're saying that you can eventually train models to arrive at that destination by training on existing jokes, effectively encoding these leaps as probabilities. In that case, the model isn't actually making an intuitive/comedic leap; they're just following new probability chains in attempting to approximate examples they've se…
This is what I refer to when it comes to larger models like Fable 5 being funnier. They are more capable of doing that. They can deliver that "sudden orthogonal leap from context" of yours more reliably.
It's not a "fundamental inability" and never was. If you crank the scale up and a capability appears, "current architecture" was never the problem.