Live data from Hacker News

Internet Archive Scholar: Search Millions of Research Papers

blog.archive.org

41–50 of 51 posts

Re: Internet Archive Scholar: Search Millions of Research Papers

#41
post #7
post #4

Earlier quoted context omitted.

I absolutely love everything about it (the logo Super fast. All my test searches returned what I was looking for. What is your relationship with semantic scholar like? Any plans to integrate ranking signals like references, etc? I'm going to double my monthly donation. This is great.

Thank you for the kind words! We are friendly with Semantic Scholar, and have used their "open corpus" dumps as one of several URL seed lists for crawling in the past. Their search and discovery tech is more sophisticated than ours is likely to be any time soon ( https://medium.com/ai2-blog/building-a-better-search-engine-... ). We would love to get to the place where groups like AI2, which are primarily research-ori…

Thank you for this great datasource!

Do you think this is suitable for bibliometric research? (We don't need citation-graphs). We use Scopus and Web of Science, but I really don't like that we are not able to publish helpful datasets that we extract from these databases.

Re: Internet Archive Scholar: Search Millions of Research Papers

#43
post #3

This service was hinted at back in September, but is now formally announced and live at https://scholar.archive.org Related previous post: https://news.ycombinator.com/item?id=24485444 Much of the catalog functionality can be accessed from the fatcat.wiki API ( https://api.fatcat.wiki/redoc ). Scholar adds a search index over the body content of papers, and we are still thinking through how to make this available thr…

[deleted]

Re: Internet Archive Scholar: Search Millions of Research Papers

#44
post #3

This service was hinted at back in September, but is now formally announced and live at https://scholar.archive.org Related previous post: https://news.ycombinator.com/item?id=24485444 Much of the catalog functionality can be accessed from the fatcat.wiki API ( https://api.fatcat.wiki/redoc ). Scholar adds a search index over the body content of papers, and we are still thinking through how to make this available thr…

Today I discovered "Open Access Diamond journals" in a report (https://zenodo.org/record/4558704/files/OADJS-Findings.pdf). These are low-scale peer-reviewed free-to-read free-to-publish non-commercial journals, typically supported by Universities or Goverment agencies. They serve diverse communities and are not predatory journals (they are free after all).

The bad news is only half of them use DOI or embed licenses in the metadata. Are they indexed or archived somewhere?

To my surprise, there are more than 350,000 papers published in OA Diamond journals every year, and most journals publish fewer than 25 articles a year.

Re: Internet Archive Scholar: Search Millions of Research Papers

#45
How does this compare to BASE and why isn't BASE used as a source?

"BASE is one of the world's most voluminous search engines especially for academic web resources. BASE provides more than 240 million documents from more than 8,000 content providers. You can access the full texts of about 60% of the indexed documents for free (Open Access). BASE is operated by Bielefeld University Library."

https://www.base-search.net/

Re: Internet Archive Scholar: Search Millions of Research Papers

#46
post #41
post #7

Earlier quoted context omitted.

Thank you for the kind words! We are friendly with Semantic Scholar, and have used their "open corpus" dumps as one of several URL seed lists for crawling in the past. Their search and discovery tech is more sophisticated than ours is likely to be any time soon ( https://medium.com/ai2-blog/building-a-better-search-engine-... ). We would love to get to the place where groups like AI2, which are primarily research-ori…

Thank you for this great datasource! Do you think this is suitable for bibliometric research? (We don't need citation-graphs). We use Scopus and Web of Science, but I really don't like that we are not able to publish helpful datasets that we extract from these databases.

I think it is in a good place for simple bibliometric queries. The fatcat elasticsearch API is open at https://api.fatcat.wiki/fatcat_release/ (behind a proxy to filter "unsafe" requests). That works pretty well for jupyter notebook style experimentation if you are willing to learn the elasticsearch query DSL for aggregations and things.

I don't think the catalog has high enough metadata quality today for use in published research. There are some glaring errors and omissions when you actually starting digging in. On the other hand, almost all bibliographic catalogs seem to have such problems. Fatcat, by being open and having an API, does have the potential to aggregate corrections, fixes, and contributions directly from researchers over time.

A particular missing piece today is that there is no categorization or "discipline" metadata of almost any type. This sort of metadata is more subjective, and the catalog currently carefully only includes factual information. We will likely start collecting metadata at the journal ("container") level and can trickle that down to papers. Aggregating, editing, and curating that metadata in Wikidata first, then importing to Fatcat, might be the best and most sustainable path forward.

Re: Internet Archive Scholar: Search Millions of Research Papers

#47
post #9

The internet archive is becoming an alternative good internet. It has a web archive, film archive, software archive, media archive... and now research papers archive. That is the internet as a giant library as we dreamed in early 90's.

Way too centralized (Centranet?), but it is very nice for now. It's a bit like the library of Alexandria, so it could change/disappear at any time.

Exactly. I encourage everyone to become a digital hoarder yourself. See a cool blog post? Assume it will be GONE in 5-10 years. So make a backup PDF copy, and throw it in dropbox. In 5-10 years if you re-encounter that page, and the internet archive is missing the page, you'll be delighted to find it in your own archive, and you can be the one who restores that information to the world.

Re: Internet Archive Scholar: Search Millions of Research Papers

#48
post #31

Earlier quoted context omitted.

I did a vanity search for my own (modest) academic output and found only one paper, which was published by a journal in Europe. The other papers were all published in either Japan or Korea and don’t appear in your search results. Two large sources in Japan you might consider trying to mirror are the UTokyo Repository [1] and Researchmap [2], through which many researchers in Japan release PDFs of their own papers. Ot…

Ah, sorry to hear. We in particular want to include content from outside the US/Europe publishing world. For Japanese publishing, we have done metadata imports from JaLC (Japanese DOI registrar), and crawled a lot of open content from J-Stage ( https://www.jstage.jst.go.jp/ ) and I hoped that coverage was pretty good. If you get a chance, could you try searching for metadata records on https://fatcat.wiki , with both…

Many thanks for the reply. I will contact some colleagues at our university library to ask for suggestions about how to check systematically how comprehensively J-Stage and Fatcat cover research publications in Japan. I will also ask if they have any suggestions about other sources from which the IA might gather such data from Japan. Either I or they will contact you by e-mail.

My subsequent vanity searches at J-Stage and Fatcat weren’t very encouraging. Most of my own papers have appeared in journals published by Japanese university departments or academic societies. While PDFs of the papers appear on the websites of the issuing organizations and show up on Google Scholar, they don’t seem to have DOIs or be listed on J-Stage.

I should mention that my research has mostly been on the humanities side of things, while J-Stage is “an electronic journal platform for science and technology information in Japan” [1].

[1] https://www.jstage.jst.go.jp/static/pages/JstageOverview/-ch...

Re: Internet Archive Scholar: Search Millions of Research Papers

#49

How does this compare to BASE and why isn't BASE used as a source? "BASE is one of the world's most voluminous search engines especially for academic web resources. BASE provides more than 240 million documents from more than 8,000 content providers. You can access the full texts of about 60% of the indexed documents for free (Open Access). BASE is operated by Bielefeld University Library." https://www.base-search.ne…

Base indexes the metadata only I believe.

Re: Internet Archive Scholar: Search Millions of Research Papers

#50
post #48

Earlier quoted context omitted.

Ah, sorry to hear. We in particular want to include content from outside the US/Europe publishing world. For Japanese publishing, we have done metadata imports from JaLC (Japanese DOI registrar), and crawled a lot of open content from J-Stage ( https://www.jstage.jst.go.jp/ ) and I hoped that coverage was pretty good. If you get a chance, could you try searching for metadata records on https://fatcat.wiki , with both…

Many thanks for the reply. I will contact some colleagues at our university library to ask for suggestions about how to check systematically how comprehensively J-Stage and Fatcat cover research publications in Japan. I will also ask if they have any suggestions about other sources from which the IA might gather such data from Japan. Either I or they will contact you by e-mail. My subsequent vanity searches at J-Stag…

This is great feedback, thank you.

For future follow-up, my work email is my handle here (bnewbold) at archive.org

Post reply on HN