Live data from Hacker News

Internet Archive Scholar: Search Millions of Research Papers

blog.archive.org

21–30 of 51 posts

Re: Internet Archive Scholar: Search Millions of Research Papers

#21
(OffTopic) All this talk about the logo here made me check the page out, instead of moving on after reading just the comments as I might otherwise have done. Perhaps that's a HN strategy to use, to get people to actually click through - add a bikesheddy thing to the page that's likely to be divisive, but doesn't require thought. Gives us a cheap way to have an opinion, and thus an incentive to click!

Re: Internet Archive Scholar: Search Millions of Research Papers

#22

Earlier quoted context omitted.

I'm sure they'd be willing to decentralize it if there was a good way to do that. Maybe this can be done with something like IPFS [0]. [0] https://ipfs.io/

Yes, they have very good intentions right now, but what if the leader gets hit by a bus.

Presumably it would be be acquired, paywalled, and monetized by a private equity firm (or some suitably hostile intellectual property rightsholder organization) before going bankrupt and shutting down for good.

Thanks for an incredible journey.

Re: Internet Archive Scholar: Search Millions of Research Papers

#24
post #9

Earlier quoted context omitted.

Way too centralized (Centranet?), but it is very nice for now. It's a bit like the library of Alexandria, so it could change/disappear at any time.

I'm sure they'd be willing to decentralize it if there was a good way to do that. Maybe this can be done with something like IPFS [0]. [0] https://ipfs.io/

Already exists, guys. You're late to the party.

https://dweb.archive.org/archive.html?identifier=home

Re: Internet Archive Scholar: Search Millions of Research Papers

#25
post #9

Earlier quoted context omitted.

Way too centralized (Centranet?), but it is very nice for now. It's a bit like the library of Alexandria, so it could change/disappear at any time.

I'm sure they'd be willing to decentralize it if there was a good way to do that. Maybe this can be done with something like IPFS [0]. [0] https://ipfs.io/

The Internet Archive already stores some big public domain data sets in IPFS/Filecoin: Prelinger Films & Librivox audiobooks. They've been partnering with Protocol Labs for 5 years. https://blog.archive.org/2020/10/22/what-information-should-...

Re: Internet Archive Scholar: Search Millions of Research Papers

#26

Earlier quoted context omitted.

Yes, they have very good intentions right now, but what if the leader gets hit by a bus.

Presumably it would be be acquired, paywalled, and monetized by a private equity firm (or some suitably hostile intellectual property rightsholder organization) before going bankrupt and shutting down for good. Thanks for an incredible journey.

Good thing is Internet Archive is a nonprofit, so cannot be acquired.

Re: Internet Archive Scholar: Search Millions of Research Papers

#27

I couldn't find a list of what sources (like which journals) they're archiving from. Does anyone know where to find that? It would be nice to see what subject categories the archive covers.

We are mostly not indexing on a journal-by-journal basis, but try to import from large, broad sources. For example, DOI registrars (Crossref, Datacite, J-Stage), DOAJ article and journal metadata (for OA publications), etc. Some field-specific indexes we have imported from include JSTOR early journals subset, PubMed, and dblp.

Some fields/disciplines are probably still systemically under-represented. For example, I bet we are missing a bunch of scholarship on art and history published before 1980. We have a couple ideas up our sleeves which we hope will help with "completeness" across more disciplines.

To answer your question directly, you can search journal names here: https://fatcat.wiki/container/search

And click through to see how many articles we know about, and what we think the preservation status is. Click through again to the "coverage" tab for a more detailed breakdown. (improving the usability and ranking on the journal search results is on our short list)

Re: Internet Archive Scholar: Search Millions of Research Papers

#28

Earlier quoted context omitted.

Presumably it would be be acquired, paywalled, and monetized by a private equity firm (or some suitably hostile intellectual property rightsholder organization) before going bankrupt and shutting down for good. Thanks for an incredible journey.

Good thing is Internet Archive is a nonprofit, so cannot be acquired.

Wikipedia is a not for profit and it's still been acquired, just by people who are insane instead of rich.

Until everyone can own their own copy and moderate it the dream of an open network is just that: a dream.

Re: Internet Archive Scholar: Search Millions of Research Papers

#29
post #7
post #4

Earlier quoted context omitted.

I absolutely love everything about it (the logo Super fast. All my test searches returned what I was looking for. What is your relationship with semantic scholar like? Any plans to integrate ranking signals like references, etc? I'm going to double my monthly donation. This is great.

Thank you for the kind words! We are friendly with Semantic Scholar, and have used their "open corpus" dumps as one of several URL seed lists for crawling in the past. Their search and discovery tech is more sophisticated than ours is likely to be any time soon ( https://medium.com/ai2-blog/building-a-better-search-engine-... ). We would love to get to the place where groups like AI2, which are primarily research-ori…

I really like what you have done.

One easy improvement "Showing results 16 — 30 out of 26 results" :-) showing below search results...

> Hope to include more curated signals, like "won a paper > prize", "journal in DOAJ and other reviewed indices", etc. This would be a great addition.

Post reply on HN