Live data from Hacker News

The Web is missing an essential part of infrastructure: an open web index

arxiv.org

131–132 of 132 posts

Re: The Web is missing an essential part of infrastructure: an open web index

#131
post #72

Earlier quoted context omitted.

Apart from that Common Crawl respects robots.txt (which makes sense) so many sites you expect to see there are not indexed. Netflix, Facebook LinkedIn and many more. If common-crawl sees serious adoption those sites will modify their robots.txt but it's and chicken/egg problem.

There is a simple solution: if companies do not respect do-not-track then why should we respect robots.txt?

Because then you end up in an arms race that the little guy usually does not win.

There are a significant number of crawlers out there that don't respect robots.txt. The usual response to them isn't to roll over dead, it's to get CloudFlare (on the technological end) and/or sic the lawyers on them (for CFAA, IP, or ToS violations).

Re: The Web is missing an essential part of infrastructure: an open web index

#132
post #109

Earlier quoted context omitted.

I prefer

Arguably no one invests in the title tag anymore because it's not user-visible in the way a heading tag is, or go further in the other direction and use the ` ` tags honored by Facebook and Twitter, since the page author has incentive to keep that content up-to-date

But also arguably 'title' is still important maybe even more than before because it shows on the tab, and everybody uses lots of tabs now.

When there was no tabs users could always see the content of the page knowing what they are looking at. But tabs hide other pages so it is important that we know what is in all those other tabs.

Post reply on HN