The scale of problem seems a bit overblown. Anna's Archive paint a picture like AI companies are some movie villains burning books so no one can see them. but in reality they just disassemble them into pages because it's cheaper and faster to scan. Most of these books is highly specialized, they been collecting dust on shelves for decades and nobody need them. But the problem is real. Even if these books aren't neede…
That's the problem. I tried the other day to find the verbatim quote from one of the Gabriel García Márquez's books - a single effing sentence and I nearly lost my mind - every LLM would be like "this is a copyrighted material and I can't share it verbatim". WTF? These books are widely known, published in many languages, and yet no amount of money spent on tokens can give me the exact fucking sentence, just like the author intended? Really? I need to find, buy, download and search through the actual book to quote something that Márquez wrote at some point? What is even worse is when the LLM just outright lies to you, giving you the quote, but slightly rephrased.