This is cool for English text. But once you get Unicode with various ways to represent é, whew lad. This get shitty quickly in the shell.
English is Unicode. Pretending otherwise would be quite naïve. https://www.azabani.com/pages/gbu/#slide4
"Unicode with various ways to represent é" is a shit show to parse using shell tools. e.g. Try scraping Spanish language Twitter feeds. When I have done this kind of work, I made a tool to canonicalize glyphs and had to put it between every step of a pipeline.