Isn't this what Common Crawl[1] is. From their FAQ: > What is Common Crawl? > Common Crawl is a 501(c)(3) non-profit organization dedicated to providing a copy of the internet to internet researchers, companies and individuals at no cost for the purpose of research and analysis. > What can you do with a copy of the web? > The possibilities are endless, but people have used the data to improve language translation sof…
I dont believe Common Crawl offers a real time search index as its delayed by more than a month (although that could have changed recently). Still useful for research purposes but not that desirable for a search engine that competes with Google, etc.
The Web is missing an essential part of infrastructure: an open web index
91–100 of 132 posts
Re: The Web is missing an essential part of infrastructure: an open web index
#92Earlier quoted context omitted.
I think there are a few factors that would make this idea unworkable. There are two categories of issues, technical and economic, that prevent this from working. I'll go into more detail about the technical issues. The top-level problem is fan-out. If you want to fan the query to the top million domains (far too few to match Google's retrieval depth, but enough to demonstrate the issue), you're going to need to imple…
>"Scoring is at least as important. Modern scorers are multi-level, meaning they do one pass over many documents on a simple, correlated representation of the data." Can you elaborate on what is the "simple correlated representation of the data"? It sounds like you understand this space pretty well might you have any links or literature on how modern crawling architectures and indexing work? Thanks.
[1]: https://nlp.stanford.edu/IR-book/pdf/07system.pdf
[2]: https://pdfs.semanticscholar.org/2795/d9d165607b5ad6d8b97183...
Re: The Web is missing an essential part of infrastructure: an open web index
#93Earlier quoted context omitted.
I dont believe Common Crawl offers a real time search index as its delayed by more than a month (although that could have changed recently). Still useful for research purposes but not that desirable for a search engine that competes with Google, etc.
This is literally why I created my company: http://www.datastreamer.io/ We've been around for about a decade. IBM watson used us as their social data provider during Jeopardy. We provide data to tons of companies and you're probably using our services - just that it's not obvious where we're used since it's SaaS B2B and not B2C. We're not free but the primary reason we exist is that other vendors charge borderline ex…
Re: The Web is missing an essential part of infrastructure: an open web index
#94Earlier quoted context omitted.
Google doesn’t crawl all of the internet very often either. Only sites that have proven to change a lot. So you could presumably supplement commoncrawl with your own more regular crawls.
I'm curious how they track and rank a site's "change velocity" without crawling all of the internet all of the time. It almost seems like a catch 22 no? Might you have any insight into how this works? Any suggested reading or links?
A site that gets lots of links quickly (and is therefore important) will likely garner them from sites you are already frequently visiting.
Re: The Web is missing an essential part of infrastructure: an open web index
#95Earlier quoted context omitted.
This is literally why I created my company: http://www.datastreamer.io/ We've been around for about a decade. IBM watson used us as their social data provider during Jeopardy. We provide data to tons of companies and you're probably using our services - just that it's not obvious where we're used since it's SaaS B2B and not B2C. We're not free but the primary reason we exist is that other vendors charge borderline ex…
This doesn't make any sense. You talk about open data but yours is the opposite. You're just another commercial data hoarder, please don't act like you're not.
Re: The Web is missing an essential part of infrastructure: an open web index
#96Earlier quoted context omitted.
I dont believe Common Crawl offers a real time search index as its delayed by more than a month (although that could have changed recently). Still useful for research purposes but not that desirable for a search engine that competes with Google, etc.
This is literally why I created my company: http://www.datastreamer.io/ We've been around for about a decade. IBM watson used us as their social data provider during Jeopardy. We provide data to tons of companies and you're probably using our services - just that it's not obvious where we're used since it's SaaS B2B and not B2C. We're not free but the primary reason we exist is that other vendors charge borderline ex…
Re: The Web is missing an essential part of infrastructure: an open web index
#97I have been wanting this for years... If you look at the original Yahoo Page when Yahoo first started out it attempted to solve this problem. I believe this index could be regionally or language based... In the United States one could use Dewey Decimal https://en.wikipedia.org/wiki/Dewey_Decimal_Classification Library of Congress https://en.wikipedia.org/wiki/Library_of_Congress_Classifica...
I tried to simplify all data into ~30 categories. My own interests fit into 16, so I drew a visual representation of them. https://github.com/peterburk/sortlikes
Next, I need to figure out the sub-categories. Genres for music, countries for travel, etc.
What interests me most is the cross-cultural connections. For example, Taiwanese punk rock (Fire Ex), or Mongolian folk metal (Nine Treasures, Hanggai). I like that music because it's the same sub-category I'm interested in (Music/Rock).
It's also possible to model the flow of finance around the world through this categorisation. Some of the categories are innately human and don't seem to exist in animals (music, cooking).
Email me if you'd like to chat more about how to categorise culture - I think it's important and I've got lots of ideas about it, but I haven't yet met any other people with this same passion.
Re: The Web is missing an essential part of infrastructure: an open web index
#98Earlier quoted context omitted.
I dont believe Common Crawl offers a real time search index as its delayed by more than a month (although that could have changed recently). Still useful for research purposes but not that desirable for a search engine that competes with Google, etc.
This is literally why I created my company: http://www.datastreamer.io/ We've been around for about a decade. IBM watson used us as their social data provider during Jeopardy. We provide data to tons of companies and you're probably using our services - just that it's not obvious where we're used since it's SaaS B2B and not B2C. We're not free but the primary reason we exist is that other vendors charge borderline ex…
Re: The Web is missing an essential part of infrastructure: an open web index
#99As a user, if some other search engine can serve results that are better than Google, I'd be happy to use it. I've tried duckduckgo, the results are disappointing and often mis-intepreted what I intended to search. So I kept coming back to Google. Will Google be willing to open its indexes? Probably not at their best interest, because it will help its competitors?
https://news.ycombinator.com/item?id=17548623
If you like, I can try implementing that with my next data analysis project. Right now I'm studying the MySpace Dragon Hoard, and I'll soon write a blog post with maps of music genres around the world.
Re: The Web is missing an essential part of infrastructure: an open web index
#100Earlier quoted context omitted.
This is literally why I created my company: http://www.datastreamer.io/ We've been around for about a decade. IBM watson used us as their social data provider during Jeopardy. We provide data to tons of companies and you're probably using our services - just that it's not obvious where we're used since it's SaaS B2B and not B2C. We're not free but the primary reason we exist is that other vendors charge borderline ex…
This doesn't make any sense. You talk about open data but yours is the opposite. You're just another commercial data hoarder, please don't act like you're not.