Live data from Hacker News

The Destruction of the Web

jacquesmattheij.com

31–40 of 98 posts

Re: The Destruction of the Web

#31
What about something like a disavow.txt file that site owners could use to list domains or URLs with unwanted inbound links. It would be similar to the Google Disavow Links tool, but more open and standardized.

We could write a simple spec around it. I'd see it being similar to robots.txt in form and function... easy for a human to write, easy for a search engine to parse, easy to generate programmatically if you need to scale it up.

Also, it avoids the black-hat SEO problem since only folks with access to the site could control the content of the disavow.txt file.

Thoughts?

Re: The Destruction of the Web

#33
post #4

This is a nice description of the hell of modern WWW. But things are worse than that! With few exceptions (Stack Exchange and Wikipedia are notable) most searches will return sites that have been SEOd. My Google search for [spectacles cases] returns these two sites on the first page: ( http://www.spectaclecases.co.uk/ ) ( http://www.aglassescase.co.uk/ ) > Welcome to SpectacleCases.co.uk. You will find a wide selecti…

Interestingly, though much of the content is the same, those domains resolve to different IP addresses, not just the same server tweaking the page for different domains as I thought. Non-authoritative answer: Name: www.spectaclecases.co.uk Address: 83.170.88.134 Non-authoritative answer: Name: www.aglassescase.co.uk Address: 83.170.70.106 Presumably they have two different webservers, perhaps connected to the same or…

The different IPs are for SSL and they are hosted here: https://vps.net

Re: The Destruction of the Web

#34
post #4

This is a nice description of the hell of modern WWW. But things are worse than that! With few exceptions (Stack Exchange and Wikipedia are notable) most searches will return sites that have been SEOd. My Google search for [spectacles cases] returns these two sites on the first page: ( http://www.spectaclecases.co.uk/ ) ( http://www.aglassescase.co.uk/ ) > Welcome to SpectacleCases.co.uk. You will find a wide selecti…

I think Stack Exchange and Wikipedia are examples of SEO'd sites too. For example: Stack Exchange had problems with indexing before they added a sitemap [1]. Wikipedia takes a lot of care to create machine-indexable pages with quality content, that link back to their sources [2] and are linked to internally. Both follow SEO best practices as described in the Google Quality Guidelines [3] Or your definition of "SEO'd"…

Yes, you're right. I tend to use "SEOd" as being a bit sleazy, rather than just following best current practice.

Re: The Destruction of the Web

#35
post #31

What about something like a disavow.txt file that site owners could use to list domains or URLs with unwanted inbound links. It would be similar to the Google Disavow Links tool, but more open and standardized. We could write a simple spec around it. I'd see it being similar to robots.txt in form and function... easy for a human to write, easy for a search engine to parse, easy to generate programmatically if you nee…

Google webmaster tools already verify if the owner is for real. The disavow tool is 'safe' in this sense, and requests to remove links sent via email are definitely not safe in this sense (that's why I never comply with those, for all I know I'm aiding some black hat by killing backlinks of a legitimate site).

Re: The Destruction of the Web

#36
post #31

What about something like a disavow.txt file that site owners could use to list domains or URLs with unwanted inbound links. It would be similar to the Google Disavow Links tool, but more open and standardized. We could write a simple spec around it. I'd see it being similar to robots.txt in form and function... easy for a human to write, easy for a search engine to parse, easy to generate programmatically if you nee…

Google webmaster tools already verify if the owner is for real. The disavow tool is 'safe' in this sense, and requests to remove links sent via email are definitely not safe in this sense (that's why I never comply with those, for all I know I'm aiding some black hat by killing backlinks of a legitimate site).

True, but that still leaves other search engines in the dark about which links should be ignored. I thought disavow.txt might be a better solution to help us avoid The Destruction of the Web.

Re: The Destruction of the Web

#37
post #36

Earlier quoted context omitted.

Google webmaster tools already verify if the owner is for real. The disavow tool is 'safe' in this sense, and requests to remove links sent via email are definitely not safe in this sense (that's why I never comply with those, for all I know I'm aiding some black hat by killing backlinks of a legitimate site).

True, but that still leaves other search engines in the dark about which links should be ignored. I thought disavow.txt might be a better solution to help us avoid The Destruction of the Web.

Good point, of course there are other search engines too and in a way a 'webmaster / search engine' interface that requires webmasters to have direct contact with a search engine when the same thing could be fixed by something the crawler could pick up is less elegant. I had not thought this through when I wrote my reply to you, you are absolutely right.

Re: The Destruction of the Web

#38
post #3

I'm "relatively" new to the web. I started using it around 2004. To me, the only way to browse the web goes through Google. I don't think there's a single day I spent in front of the computer without me hitting Google at one point. I even use Google search when I'm specifically targeting Wikipedia or StackOverflow. For the people who got introduced to the web before me, how was "web browsing" done in the earlier deca…

> For the people who got introduced to the web before me, how was "web browsing" done in the earlier decade of the Web? I'm assuming Google is not the first search engine available, but I'm pretty sure search engines were not the only way to go around.

Yahoo! used to have humans looking at websites and adding them to an index. Yes, someone would send in a link to a porn website, and someone on the porn indexing team would view the site and add it to an index.

Curation efforts like this were important. People built web-rings for similar content; Usenet FAQs listed useful sites.

But people didn't just use WWW. They used Usenet, sometimes Gopher or telnet, ftp, and email. Or they were part of some other online community that had a www gateway. The prices now seem eye-watering.

The electronic landscape in 1988 (https://news.ycombinator.com/item?id=3087928)

(http://i53.tinypic.com/2janfrd.jpg)

Compuserve - $11 per hour.

The Source - $8 per hour

Delphi - $6 per hour

BIX $9 per hour

And this is at a time when people had slow modems and usually paid for the telephone calls too. Thus, offline readers (things like BlueWave for email and fidonet) were popular.

Your last paragraph: I'd have a look through the newsgroup lists for relevant groups. I'd subscribe, fetch headers, look for a faq, retrieve the faq, and read that. I'd lurk the group for a bit, and try to do my own work. Then, after I'd learnt the group for a bit and participated in other stuff I'd try to ask a good question, with links to how far I'd got and an attempt at a correct reply.

Yahoo! was pretty good once you got the hang of rules. HotBot, dogpile, altavista (and astalavista) were also handy. But Google really was revolutionarily good.

Re: The Destruction of the Web

#39
> In the scientific world there are no spammers and there is no direct commercial advantage to creating a lot of nonsense paper that cite your own paper, also there is some oversight in the world of science and the people there have a reasonably high level of integrity.

Um...what? If it were anyone but the OP, who always writes with a lot of thoughtfulness and insight, I would've assumed the graf above is satire. Academic discovery and citation is very much being gamed; the only reason why we don't notice it more is because the academics don't have the same tools and infrastructure that web spammers do and, also, the world of academic research is not something the average person outside of academia closely parses.

Re: The Destruction of the Web

#40
post #5

What Google seems to ultimately be doing with its Pandas and other attempts at stopping search spam is TEACHING businesses that the only safe way forward is great content and white hat practices. Sure, you might temporarily get some advantages with search spam but come next Panda and you might be totally screwed. Better play it safe and play nice with Google.

The problem is that "great content" (aka landing pages) is merely another form of spam.

Can you clarify? Why is "great content" == "landing pages"? How are landing pages spam?
Post reply on HN