Earlier quoted context omitted.
You're missing the point, it's only a testing excersize for the new model.
No, the point is that you can set up the testing exercise without using an LLM to do a simple find and replace.
It has nothing to do about the performance of the string replacement.
The initial "Find" is to see how well it performs actually find all the "spells" in this case, then to replace them. They using a separate context maybe, evaluate if the results are the same or are they skewed in favour of training data.