Earlier quoted context omitted.
It doesn't work for me with regex101. "The preceding token is not quantifiable" on this part: |
See, this is kinda what I mean. Maybe you can detect tags with regex, but maybe you shouldn't, given the widespread but subtle differences in regex engines. Perhaps the entire approach of "why are you trying to parse X?" Needs to be traced and re-evaluated.
So what do you think would be a more appropriate choice for writing a tokenizer?