Live data from Hacker News

Internet Archive Scholar: Search Millions of Research Papers

blog.archive.org

31–40 of 51 posts

Re: Internet Archive Scholar: Search Millions of Research Papers

#31

I couldn't find a list of what sources (like which journals) they're archiving from. Does anyone know where to find that? It would be nice to see what subject categories the archive covers.

We are mostly not indexing on a journal-by-journal basis, but try to import from large, broad sources. For example, DOI registrars (Crossref, Datacite, J-Stage), DOAJ article and journal metadata (for OA publications), etc. Some field-specific indexes we have imported from include JSTOR early journals subset, PubMed, and dblp. Some fields/disciplines are probably still systemically under-represented. For example, I b…

I did a vanity search for my own (modest) academic output and found only one paper, which was published by a journal in Europe. The other papers were all published in either Japan or Korea and don’t appear in your search results.

Two large sources in Japan you might consider trying to mirror are the UTokyo Repository [1] and Researchmap [2], through which many researchers in Japan release PDFs of their own papers. Other Japanese universities probably have archives similar to [1].

If you would like me to contact somebody at [1] who might be able to work with you, please let me know in a reply to this comment. (I helped to arrange the IA’s recent tie-up with the University of Tokyo General Library.)

[1] https://repository.dl.itc.u-tokyo.ac.jp/?lang=english

[2] https://researchmap.jp/?lang=en

Re: Internet Archive Scholar: Search Millions of Research Papers

#32

What are the differences and advantages over Sci-Hub?

Sci-Hub exists specifically to exfiltrate paywalled research papers; IA Scholar is for open-access papers that have disappeared off the Internet. They do different things.

Re: Internet Archive Scholar: Search Millions of Research Papers

#33
This seems pretty good.

In computer science we are pretty lucky because open access is the norm.

I checked a few well known exceptions, and this seems to find them ok.

"Mastering the game of Go without human knowledge" (Deepmind in Nature): https://scholar.archive.org/search?q=key:work_yqdj7vjbefg7hh...

"Typing candidate answers using type coercion" (IBM Watson special edition, IEEE IBM Systems Journals): https://scholar.archive.org/search?q=key:work_dym4lqay5fcdxo...

Re: Internet Archive Scholar: Search Millions of Research Papers

#34
post #28

Earlier quoted context omitted.

Good thing is Internet Archive is a nonprofit, so cannot be acquired.

Wikipedia is a not for profit and it's still been acquired, just by people who are insane instead of rich. Until everyone can own their own copy and moderate it the dream of an open network is just that: a dream.

[deleted]

Re: Internet Archive Scholar: Search Millions of Research Papers

#36

This is nice! I just managed to find an article, I couldn’t find with Google. Thus, I was able to solve the PDP-1 "Amherst Mystery" [1]: https://www.masswerk.at/nowgobang/2021/pdp1-spotting#update [1] https://news.ycombinator.com/item?id=26313124

Google fails to index so many good sources nowadays... I think that it has been gotten worst over the last 10 years.

Re: Internet Archive Scholar: Search Millions of Research Papers

#37

This is nice! I just managed to find an article, I couldn’t find with Google. Thus, I was able to solve the PDP-1 "Amherst Mystery" [1]: https://www.masswerk.at/nowgobang/2021/pdp1-spotting#update [1] https://news.ycombinator.com/item?id=26313124

Google fails to index so many good sources nowadays... I think that it has been gotten worst over the last 10 years.

See also https://fatcat.wiki (which is, I think, incorporated into Internet Archive Scholar.)

Re: Internet Archive Scholar: Search Millions of Research Papers

#38
post #9

The internet archive is becoming an alternative good internet. It has a web archive, film archive, software archive, media archive... and now research papers archive. That is the internet as a giant library as we dreamed in early 90's.

Way too centralized (Centranet?), but it is very nice for now. It's a bit like the library of Alexandria, so it could change/disappear at any time.

> It's a bit like the library of Alexandria, so it could change/disappear at any time.

The irony here is that the only second full copy of the Internet Archive is actually hosted at the library of Alexandria.

Source: Digital Amnesia Documentary [1]

[1] https://www.youtube.com/watch?v=NdZxI3nFVJs

Re: Internet Archive Scholar: Search Millions of Research Papers

#39
post #7

Earlier quoted context omitted.

Thank you for the kind words! We are friendly with Semantic Scholar, and have used their "open corpus" dumps as one of several URL seed lists for crawling in the past. Their search and discovery tech is more sophisticated than ours is likely to be any time soon ( https://medium.com/ai2-blog/building-a-better-search-engine-... ). We would love to get to the place where groups like AI2, which are primarily research-ori…

I really like what you have done. One easy improvement "Showing results 16 — 30 out of 26 results" :-) showing below search results... > Hope to include more curated signals, like "won a paper > prize", "journal in DOAJ and other reviewed indices", etc. This would be a great addition.

Fixed, thanks!

Re: Internet Archive Scholar: Search Millions of Research Papers

#40
post #31

Earlier quoted context omitted.

We are mostly not indexing on a journal-by-journal basis, but try to import from large, broad sources. For example, DOI registrars (Crossref, Datacite, J-Stage), DOAJ article and journal metadata (for OA publications), etc. Some field-specific indexes we have imported from include JSTOR early journals subset, PubMed, and dblp. Some fields/disciplines are probably still systemically under-represented. For example, I b…

I did a vanity search for my own (modest) academic output and found only one paper, which was published by a journal in Europe. The other papers were all published in either Japan or Korea and don’t appear in your search results. Two large sources in Japan you might consider trying to mirror are the UTokyo Repository [1] and Researchmap [2], through which many researchers in Japan release PDFs of their own papers. Ot…

Ah, sorry to hear. We in particular want to include content from outside the US/Europe publishing world.

For Japanese publishing, we have done metadata imports from JaLC (Japanese DOI registrar), and crawled a lot of open content from J-Stage (https://www.jstage.jst.go.jp/) and I hoped that coverage was pretty good. If you get a chance, could you try searching for metadata records on https://fatcat.wiki, with both Japanese and English titles and names (if applicable)?

For Korean publishing, the regional DOI registrar (https://www.kisti.re.kr/eng/) does not provide open metadata, which is a known hole in our coverage. IIRC it looked like there might be a way to scrape at least DOIs, titles, and author names, but haven't had time to take a crack at it.

Mainland Chinese publishing is probably the biggest single hole in coverage by absolute numbers. There are two DOI registrars and neither have open metadata.

Regarding the u-tokyo.ac.jp, it looks like we are able to consume metadata and do crawls via the OAI-PMH protocol. We crawled over 112k URLs from that domain via that protocol about a year ago, and they should be preserved/mirrored in web.archive.org but they haven't ended up in fatcat or scholar yet. We want to go slow with pulling in OAI-PMH content, and ensure we de-duplicate records and add filters to ensure we are getting clean metadata and content. Also preserving repository content hasn't been as urgent as getting to small OA publishers which might lack a preservation scheme and vanish off the web.

Post reply on HN