Live data from Hacker News

The Web is missing an essential part of infrastructure: an open web index

arxiv.org

31–40 of 132 posts

Re: The Web is missing an essential part of infrastructure: an open web index

#31
post #23

While I like the idea, I fear the potential for abuse, conflict and community splits. It will need some sort of moderation, at least to prevent: 1. spam 2. child pornography 3. content against the laws The only thing that is easy to define as policy is #2. No one likes child porn. But even then, there are grey areas with differing legal status - lolicon on the anime side and "barely legal" on the realistic side, plus…

>No one likes child porn. I hope you do realize the contradiction.

Aside from a couple thousand pedophiles, sorry but I'm not gonna take care of their needs...

Re: The Web is missing an essential part of infrastructure: an open web index

#32

While I like the idea, I fear the potential for abuse, conflict and community splits. It will need some sort of moderation, at least to prevent: 1. spam 2. child pornography 3. content against the laws The only thing that is easy to define as policy is #2. No one likes child porn. But even then, there are grey areas with differing legal status - lolicon on the anime side and "barely legal" on the realistic side, plus…

The proposal is for a publicly funded index as base-level infrastructure.

Filtering out spam, pornography, and other undesirable or illegal content would be done at the service level, i.e., but companies/organizations building user-facing search applications on top of the index.

Re: The Web is missing an essential part of infrastructure: an open web index

#33
post #16

Earlier quoted context omitted.

I dont believe Common Crawl offers a real time search index as its delayed by more than a month (although that could have changed recently). Still useful for research purposes but not that desirable for a search engine that competes with Google, etc.

Would users notice for many searches? Obviously it wouldn't be useful for news or social media, but for practically everything else a latency of a month would be fine.

Yes. As the Internet Archive shows, a large corpus of valuable content is no longer changing.

Re: The Web is missing an essential part of infrastructure: an open web index

#34
I didn't see mention of who would pay for this infrastructure. Is it considered a gov't funded or volunteer / donation thing?

There doesn't seem to be a mention of how to alleviate a tragedy of the commons problem (unless I missed it). If common crawl is doing a fine job, who funds them?

Re: The Web is missing an essential part of infrastructure: an open web index

#35
This simply isn't needed, and if it is it can be done by a charity or any group of people, not something that should be built into the infrastructure of the web itself.

You have to remember that while the little that the web provides is also its strongest attraction; it allows the web to be accessed and modified by anyone, they're on little bit of the web can be very different from someone elses.

So by adding on a way that the web must be indexed is kind of like moving closer to communism than liberalism. I guess if we start dictating to google where to get their data then we've moved to the full blown hammer and sickle stage :).

Re: The Web is missing an essential part of infrastructure: an open web index

#36

Earlier quoted context omitted.

It won't work without a central authority. See Soundcloud as an example. People tag their music with whatever they think will get them traffic. So, in order to do this you'll need a mass of volunteers which will lead to politics "XYZ should be classified as G! No it should be F!", "classifying ABC as DEF is racists/sexist/..." and other arguments. You'll also get people lobbying to have things removed (right to be fo…

If anything, that sounds like a solid argument to decentralize it. I don't want China's government, white supremacists, churches, soccer moms, Jihadis or grievance-of-the-month activists controlling how information is indexed; I would rather use multiple indexes that balance out controlling interests and biases.

Unfortunately if it's decentralized, then it becomes controlled by spamlords, SEO artists, advertisers, and anyone else who stands to gain from manipulating the index to their advantage. At least if it's centralized, the fights are out in the open and have a chance of converging on something reasonable (like e.g. wikipedia).

Re: The Web is missing an essential part of infrastructure: an open web index

#37
post #25

There are two entities trying to pull this off: Common Crawl (non-profit): Stores regular, broad, monthly crawls as WARC files. Provides a separate index that can be used to look data up (no a fulltext index though). Used mostly in academia. Mixnode (for-profit): Regularly crawls the web and lets users write SQL queries against the data. Not sure who the primary users are since it's in private beta. There are some se…

I think CC used to provide full-text indices. Not sure though and can't find any posts on it.

Re: The Web is missing an essential part of infrastructure: an open web index

#38

This simply isn't needed, and if it is it can be done by a charity or any group of people, not something that should be built into the infrastructure of the web itself. You have to remember that while the little that the web provides is also its strongest attraction; it allows the web to be accessed and modified by anyone, they're on little bit of the web can be very different from someone elses. So by adding on a wa…

Putting together an accurate picture of the state of the world is not contradictory with liberalism.

Re: The Web is missing an essential part of infrastructure: an open web index

#39

Earlier quoted context omitted.

Google needs sub-second response to show ads. Some users may be happy to wait for hours or days to get high quality answers not available from commercial companies. Can still be faster than emailing a human friend or consultant or tasking an employee or department.

No idea where you got this statisti, but I guarantee no user would be happy to wait hours or (god forbid) days on a search result.

I would gladly wait hours or days for certain long-tail searches; the kind that I revisit every few days/weeks/months for half an hour trying various search terms to see if I can crack the code and find the content that I know is out there somewhere.

I imagine getting status updates with intermediate search results, and I annotate each one with 'warm' or 'cold' and maybe add some more search terms into the hopper to forcibly narrow or expand the search.

Post reply on HN