Live data from Hacker News

The Web is missing an essential part of infrastructure: an open web index

arxiv.org

111–120 of 132 posts

Re: The Web is missing an essential part of infrastructure: an open web index

#112

Earlier quoted context omitted.

This doesn't make any sense. You talk about open data but yours is the opposite. You're just another commercial data hoarder, please don't act like you're not.

You are mistaking between free and open. You can be open without being free. Maintaining web index is extremely expensive. Imagine storing most of the web on your own servers and serving it. Someone has to pay bills for all those disk space and bandwidth. I don’t think web index would ever be free (unless storage, compute and bandwidth were free) but having at reasonably priced is a very good thing. I would hope thes…

> I don’t think web index would ever be free

Yet the company first mentioned does it for free, lol:

https://commoncrawl.org/

I've checked Datastreamer.io for 5 seconds, I don't see any link to their repo. If not "open source" then what does "open" mean?

Re: The Web is missing an essential part of infrastructure: an open web index

#113
post #111

Somebody could try to build own crawler and feed them with 260MM domain names dataset from https://domains-index.com

Is there more like this? Afaik SSL certificates are required to be committed to an open ledger but I can't find anywhere to obtain the ledger..

Re: The Web is missing an essential part of infrastructure: an open web index

#114

Earlier quoted context omitted.

I dont believe Common Crawl offers a real time search index as its delayed by more than a month (although that could have changed recently). Still useful for research purposes but not that desirable for a search engine that competes with Google, etc.

This is literally why I created my company: http://www.datastreamer.io/ We've been around for about a decade. IBM watson used us as their social data provider during Jeopardy. We provide data to tons of companies and you're probably using our services - just that it's not obvious where we're used since it's SaaS B2B and not B2C. We're not free but the primary reason we exist is that other vendors charge borderline ex…

Hacker News guidelines say:

> Please respond to the strongest plausible interpretation of what someone says, not a weaker one that's easier to criticize.

Your comment makes zero sense in this context, because it's just marketing.

> we're trying to enable innovation.

You're trying to make profit, like every other company in the world and that's OK.

Re: The Web is missing an essential part of infrastructure: an open web index

#116

Earlier quoted context omitted.

You are mistaking between free and open. You can be open without being free. Maintaining web index is extremely expensive. Imagine storing most of the web on your own servers and serving it. Someone has to pay bills for all those disk space and bandwidth. I don’t think web index would ever be free (unless storage, compute and bandwidth were free) but having at reasonably priced is a very good thing. I would hope thes…

> I don’t think web index would ever be free Yet the company first mentioned does it for free, lol: https://commoncrawl.org/ I've checked Datastreamer.io for 5 seconds, I don't see any link to their repo. If not "open source" then what does "open" mean?

Commoncrawl is not a company, it's a non-profit. Open means you can access the data, there is no assumption about the data being free or not.

Re: The Web is missing an essential part of infrastructure: an open web index

#117

Earlier quoted context omitted.

This doesn't make any sense. You talk about open data but yours is the opposite. You're just another commercial data hoarder, please don't act like you're not.

You are mistaking between free and open. You can be open without being free. Maintaining web index is extremely expensive. Imagine storing most of the web on your own servers and serving it. Someone has to pay bills for all those disk space and bandwidth. I don’t think web index would ever be free (unless storage, compute and bandwidth were free) but having at reasonably priced is a very good thing. I would hope thes…

Easy to test, though. If they were open, you could download their entire data set under some permissive license. If you can't then they are not open.

Re: The Web is missing an essential part of infrastructure: an open web index

#118
post #116

Earlier quoted context omitted.

> I don’t think web index would ever be free Yet the company first mentioned does it for free, lol: https://commoncrawl.org/ I've checked Datastreamer.io for 5 seconds, I don't see any link to their repo. If not "open source" then what does "open" mean?

Commoncrawl is not a company, it's a non-profit. Open means you can access the data, there is no assumption about the data being free or not.

What? It's a nonprofit organization engaging in nonprofit business. Any organization that engages in business is a "company." Common Crawl is a company. Your comment isn't accurate and it doesn't address the parent's comment.

Re: The Web is missing an essential part of infrastructure: an open web index

#119
post #49
post #16

Earlier quoted context omitted.

Would users notice for many searches? Obviously it wouldn't be useful for news or social media, but for practically everything else a latency of a month would be fine.

It absolutely breaks any use for information produced in the last month. Here's a few things that come to mind: While you say "news" really that covers any information about current-ish events. It's not just "what happened today" but background on things like the Muller report right now. Any technology release, or update. Reviews of any hardware or software. Information about security vulnerabilities. Film reviews. G…

I don't know this for certain but I strongly suspect news (and reviews/criticism, which is editorialised news) doesn't make up a huge percentage of search traffic. People read news sites that align with their preexisting points of view. They don't often go looking for new perspectives. If someone wants to know what's happening they want the filter of their preferred news outlet, if they want a review they want to read or watch their preferred reviewer. They don't want whatever happens to be the top search result.

Although, that said, with Google personalising search results the top result is very likely to be the user's preferred site anyway. We can't have people seeing outside their filter bubble after all.

Re: The Web is missing an essential part of infrastructure: an open web index

#120
The index itself is already separate in a sense that nobody is being stopped to do the indexing task themselves.

Google is a private for-profit company so we cannot realistically expect them to provide something for free to the public without generating profits in return.

The web index is not a locked up proprietary resource by anyone, so people can do the indexing themselves but the real question is how do you fund a service that will keep increasing its workload exponentially and indefinitely? What institution will have the required resources to bare such costs?

Post reply on HN