> No one will be allowed to read or download work from the repository, because that would breach publishers’ copyright. Instead, Malamud envisages, researchers could crawl over its text and data with computer software, scanning through the world’s scientific literature to pull out insights without actually reading the text. I find this totally unconvincing. The average scientific article isn't any good, and the NLP a…
> The average scientific article isn't any good, and the NLP algorithms that do tasks like this are even worse. In my day job, I'm often tasked with implementing algorithms from recently published physics papers. In order to do so, I normally have to read through at least 10 related papers (both cited papers and cited by papers) in order to have a clear idea in my head about what is going on. Even then, I often have…
This post makes me feel slightly better about how often I come away from reading a paper feeling like I only have a shallow understanding of the content.