The Web is missing an essential part of infrastructure: an open web index
1–10 of 132 posts
Re: The Web is missing an essential part of infrastructure: an open web index
#2A proposal for building an index of the Web that separates the infrastructure part of the search engine - the index - from the services part that will form the basis for myriad search engines and other services utilizing Web data on top of a public infrastructure open to everyone.
Re: The Web is missing an essential part of infrastructure: an open web index
#3For simple "I know TF-IDF, let's build a toy search engine" it will suffice, but apart from that?
Re: The Web is missing an essential part of infrastructure: an open web index
#4Re: The Web is missing an essential part of infrastructure: an open web index
#5I assume that in a world of competitive index users, there is no one size fits all. Presumably application design (and feature) choices will heavily influence how the index should work. For simple "I know TF-IDF, let's build a toy search engine" it will suffice, but apart from that?
Re: The Web is missing an essential part of infrastructure: an open web index
#6> What is Common Crawl?
> Common Crawl is a 501(c)(3) non-profit organization dedicated to providing a copy of the internet to internet researchers, companies and individuals at no cost for the purpose of research and analysis.
> What can you do with a copy of the web?
> The possibilities are endless, but people have used the data to improve language translation software, predict trends, track the disease propagation and much more.
> Can’t Google or Microsoft just do that?
>Our goal is to democratize the data so everyone, not just big companies, can do high quality research and analysis.
Also DuckDuckGo founder Gabriel Weinberg expressed the sentiment that the index should be separate from the search engine many years ago:
> Our approach was to treat the “copy the Internet” part as the commodity. You could get it from multiple places. When I started, Google, Yahoo, Yandex and Microsoft were all building indexes. We focused on doing things the other guys couldn’t do. [2]
From what I remember reading once DuckDuckGo doesn't use Common Crawl though.
[2] https://www.japantimes.co.jp/news/2013/07/28/business/duckdu...
Re: The Web is missing an essential part of infrastructure: an open web index
#7If you look at the original Yahoo Page when Yahoo first started out it attempted to solve this problem.
I believe this index could be regionally or language based...
In the United States one could use
Dewey Decimal
https://en.wikipedia.org/wiki/Dewey_Decimal_Classification
Library of Congress
https://en.wikipedia.org/wiki/Library_of_Congress_Classifica...
Re: The Web is missing an essential part of infrastructure: an open web index
#8Isn't this what Common Crawl[1] is. From their FAQ: > What is Common Crawl? > Common Crawl is a 501(c)(3) non-profit organization dedicated to providing a copy of the internet to internet researchers, companies and individuals at no cost for the purpose of research and analysis. > What can you do with a copy of the web? > The possibilities are endless, but people have used the data to improve language translation sof…
Re: The Web is missing an essential part of infrastructure: an open web index
#9Re: The Web is missing an essential part of infrastructure: an open web index
#10I have been wanting this for years... If you look at the original Yahoo Page when Yahoo first started out it attempted to solve this problem. I believe this index could be regionally or language based... In the United States one could use Dewey Decimal https://en.wikipedia.org/wiki/Dewey_Decimal_Classification Library of Congress https://en.wikipedia.org/wiki/Library_of_Congress_Classifica...
I'm not saying it shouldn't be done but I think it will be way more work than expected and there will be all kinds of issues.