Earlier quoted context omitted.
No, the self-attention for the transformer in GPT means it isn't a Markov Chain. A blog post of you want to read more: https://medium.com/@andrew_johnson_4/are-transformers-markov...
Isn't it though, if you consider the entire context to be part of the state? It seems like his argument is based on an assumption of the Markov model only using the current word as its state.
https://safjan.com/understanding-differences-gpt-transformer...
If we're broadening the scope of a Markov chain to consider the entire system as the current state that's being used to determine the next step of operation, then isn't literally every computer program a Markov chain under that definition?
You can't just include the memory of previous states which the current state is depending on as being part of the "current state" to fit the definition of a Markov chain without having broadened the scope to the point the definition becomes meaningless.