Earlier quoted context omitted.
The complete inability to use it to generate spam doesn't make it different to you?
Well, I just said that. Since what is predicated is one of two tokens , . Since spam is not made out of those tokens, it doesn't generate spam. (Even spam could be encoded into those two tokens via binary or whatever, it still wouldn't because it's predicting a classification, and not the next sybmol; i.e. not simply the next bit in a after a string prefix, but a 0 or 1 value indicating whether the bit string is spam…
There is a massive qualitative and quantitative gap between the two. Even if you had a real token predictor for the word spam, you would need a thousand of these to touch LLM capability. But it's not a real predictor. It's not locational. It's something vastly weaker.
> It's all the same sort of thing though: some functions trained to predict a value associated with an input.
Now you made it even worse. If not for the word 'trained' this would be every program. If two programs fall under the umbrella of machine learning in any way, that kind of logic treats them the same. What would you consider a "different kind of AI"? Your definition there fits basically the entire field.