Ultimately, Wikipedia provides an effective keyword lookup that maps to curated links. Regardless, the notion of a general web index is well nigh moot at this point due to its not having been built into the system from the get-go. Any such attempt at this point will be, by definition, ad hoc and built by some group of individuals, with the vastness of the content, the cost of the project and the intrinsic conflicts t…
The Web is missing an essential part of infrastructure: an open web index
71–80 of 132 posts
Re: The Web is missing an essential part of infrastructure: an open web index
#72Isn't this what Common Crawl[1] is. From their FAQ: > What is Common Crawl? > Common Crawl is a 501(c)(3) non-profit organization dedicated to providing a copy of the internet to internet researchers, companies and individuals at no cost for the purpose of research and analysis. > What can you do with a copy of the web? > The possibilities are endless, but people have used the data to improve language translation sof…
I dont believe Common Crawl offers a real time search index as its delayed by more than a month (although that could have changed recently). Still useful for research purposes but not that desirable for a search engine that competes with Google, etc.
Re: The Web is missing an essential part of infrastructure: an open web index
#73Will Google be willing to open its indexes? Probably not at their best interest, because it will help its competitors?
Re: The Web is missing an essential part of infrastructure: an open web index
#74There are lots of niche directories out there - if you consider Reddit wikis, "awesome" lists and so on. A few of us out there are also working on small directories: * https://href.cool (mine) * https://indieseek.xyz * https://iwebthings.com The thought is that you can actually navigate a small directory - they don't need to be five levels deep - and a network of these would rival a huge directory, avoid centralizati…
Re: The Web is missing an essential part of infrastructure: an open web index
#75Publicly maintained directory that I believe was at least theoretically independent of the larger web companies. It certainly had it share of drama, but was a decent human vetted index of what was out there....
Re: The Web is missing an essential part of infrastructure: an open web index
#76Earlier quoted context omitted.
It's still sooner than "never", which is the current response time for answers that Google cannot provide. Central search indexes like Google are not going away. There are client-side metasearch interfaces that combine Google results with other sources. Those other sources can be much slower, including human responses. You would still have your synchronous sub-second response from centralized search, but there would…
>> Most of the world's knowledge is not public Where is it found ?
I can't find a reference at the moment, but this topic was covered in a professional journal for historians.
Re: The Web is missing an essential part of infrastructure: an open web index
#77I've always thought it would make more sense if each web server could be responsible for indexing the material that it serves (and offer notifications of updates), so instead of having to crawl everything yourself, you could just request the index from each domain, and then merge them.
A significant problem with this is trust. You can't trust websites to reliably or accurately index their sites due to both incompetence and malice. I don't think there's any way around the malicious component. Formal or informal standards may take care of the competence factor with the feature being built into common publishing platforms. XML sitemaps are a microcosm of putting the indexing onus on websites instead o…
Google has been purging large swaths of data from the indexes and they won't say how or why or exactly what criteria they are using. It's difficult to imagine a worse solution for the web than this current model.
Re: The Web is missing an essential part of infrastructure: an open web index
#78Ultimately, Wikipedia provides an effective keyword lookup that maps to curated links. Regardless, the notion of a general web index is well nigh moot at this point due to its not having been built into the system from the get-go. Any such attempt at this point will be, by definition, ad hoc and built by some group of individuals, with the vastness of the content, the cost of the project and the intrinsic conflicts t…
Good observation. This makes me think of all the useful content Wikipedia doesn’t link to though.
Really, I am of the perspective that a machine-grokked indexing system will always be less useful in a significant set of edge-cases than a human-curated index due to such factors as language ambiguity and gaming of such algorithms. As well, the sheer size of the internet requires ranking the pages to ensure the most useful links are properly denoted as such.
WP, being likely the most important and useful crowd-sourced and -built human information system, it is up to us to both keep it funded and add the information we deem important.
Re: The Web is missing an essential part of infrastructure: an open web index
#79Earlier quoted context omitted.
>No one likes child porn. I hope you do realize the contradiction.
Aside from a couple thousand pedophiles, sorry but I'm not gonna take care of their needs...
There are far more than a 'few thousand pedophiles'; that number is more reflective of the number of convictions each year. While drawing statistical inferences is difficult, the stats in the appendices to this report suggest there's perhaps ~100k tips a year to police about child sexual abuse across the US.
Re: The Web is missing an essential part of infrastructure: an open web index
#80Isn't this what Common Crawl[1] is. From their FAQ: > What is Common Crawl? > Common Crawl is a 501(c)(3) non-profit organization dedicated to providing a copy of the internet to internet researchers, companies and individuals at no cost for the purpose of research and analysis. > What can you do with a copy of the web? > The possibilities are endless, but people have used the data to improve language translation sof…
I dont believe Common Crawl offers a real time search index as its delayed by more than a month (although that could have changed recently). Still useful for research purposes but not that desirable for a search engine that competes with Google, etc.
We've been around for about a decade. IBM watson used us as their social data provider during Jeopardy. We provide data to tons of companies and you're probably using our services - just that it's not obvious where we're used since it's SaaS B2B and not B2C.
We're not free but the primary reason we exist is that other vendors charge borderline extortionate pricing and I fundamentally believe that the web MUST remain open.
We've also been providing data for very affordable pricing to researchers for more than a decade.
Search for us as Spinn3r under Google Scholar (our previous name) and we have hundreds and hundreds of PhDs who have access to our data.
We do charge for research usage now but it's very very very affordable.
The entire point is that we're trying to enable innovation.