Live data from Hacker News

The Destruction of the Web

jacquesmattheij.com

1–10 of 98 posts

Re: The Destruction of the Web

#2
> In the scientific world there are no spammers and there is no direct commercial advantage to creating a lot of nonsense paper that cite your own paper

Yet. Google's on it though.

>> The launch of Google Scholar Citations and Google Scholar Metrics may provoke a revolution in the research evaluation field as it places within every researchers reach tools that allow bibliometric measuring. In order to alert the research community over how easily one can manipulate the data and bibliometric indicators offered by Google s products we present an experiment in which we manipulate the Google Citations profiles of a research group through the creation of false documents that cite their documents, and consequently, the journals in which they have published modifying their H index. For this purpose we created six documents authored by a faked author and we uploaded them to a researcher s personal website under the University of Granadas domain. The result of the experiment meant an increase of 774 citations in 129 papers (six citations per paper) increasing the authors and journals H index. We analyse the malicious effect this type of practices can cause to Google Scholar Citations and Google Scholar Metrics. Finally, we conclude with several deliberations over the effects these malpractices may have and the lack of control tools these tools offer.

http://arxiv.org/abs/1212.0638

Re: The Destruction of the Web

#3
I'm "relatively" new to the web. I started using it around 2004. To me, the only way to browse the web goes through Google. I don't think there's a single day I spent in front of the computer without me hitting Google at one point. I even use Google search when I'm specifically targeting Wikipedia or StackOverflow.

For the people who got introduced to the web before me, how was "web browsing" done in the earlier decade of the Web? I'm assuming Google is not the first search engine available, but I'm pretty sure search engines were not the only way to go around.

I understand the concept behind the web, "a globe spanning network of computers linked by hyperlinks pointing to useful information". But was it as simple as that? You only had access to addresses you knew or links available on these pages? Where did you go to find interesting websites or how would you look for specific information (like, for instance, how would you research the working internals of a car engine for a school project?)

Re: The Destruction of the Web

#4
This is a nice description of the hell of modern WWW.

But things are worse than that! With few exceptions (Stack Exchange and Wikipedia are notable) most searches will return sites that have been SEOd.

My Google search for [spectacles cases] returns these two sites on the first page:

(http://www.spectaclecases.co.uk/)

(http://www.aglassescase.co.uk/)

> Welcome to SpectacleCases.co.uk. You will find a wide selection of Glasses Cases / Spectacle Cases / Sunglass Cases.

> Made from leather, fabric, metal or plastic finished to a very high quality. Hard and soft cases for spectacles, glasses and sunglasses. We also have a good selection of cheap glasses cases which offer great protection for your glasses.

> Welcome to AGlassesCase.co.uk. The one stop shop for Glasses Cases, Spectacle Cases and Sunglass Cases. We also sell a number of Glasses Cloths

> Made from a range of quality materials including leather, fabric, metal and plastic all finished to a very high standard. We sell hard and soft cases for spectacles, glasses and sunglasses. We also have a good selection of cheap glasses cases which offer great protection for your glasses.

These two different sites are the same company.

Maybe they're a great place to buy spectacles cases from, but it's vaguely upsetting that Google can create freakin' awesome stuff (A self driving car! It is actually wonderful and futuristic) yet can't fix this stuff. Obviously, Google are not to blame, and really the problem is with sleazy SEO and odd behaviours by vendors.

Re: The Destruction of the Web

#5
What Google seems to ultimately be doing with its Pandas and other attempts at stopping search spam is TEACHING businesses that the only safe way forward is great content and white hat practices. Sure, you might temporarily get some advantages with search spam but come next Panda and you might be totally screwed. Better play it safe and play nice with Google.

Re: The Destruction of the Web

#6
post #3

I'm "relatively" new to the web. I started using it around 2004. To me, the only way to browse the web goes through Google. I don't think there's a single day I spent in front of the computer without me hitting Google at one point. I even use Google search when I'm specifically targeting Wikipedia or StackOverflow. For the people who got introduced to the web before me, how was "web browsing" done in the earlier deca…

Things were a bit more varied. Netscape Navigator came with a default home page that included a number of different search engines which varied in different and sometimes interesting ways.

Because these search engines were not as powerful or accurate as Google at mining content from the web, you would tend to follow links more, relying on the overall navigational structure of the WWW rather than everything being the two step "Google terms -> Go to website in results page" process it typically is today.

Many people posted fragments called "web rings" on their web pages: these would be a linked "ring" of similar, related sites, grouped by a common interest in the subject matter. It was a useful (though very random and sometimes temperamental) way to discover similar sites and content.

Other sites had "guest books" where people could comment on the site and include a link to their own. There's little functional difference between guest books and today's blog commenting systems, except back then there was less traffic, so site owners tended to actually read the comments, respond to them often, and follow the links of the commenters back to their own sites.

Then SEO spam came along, and eventually most people realised 90% of the links in comments were probably not worth following. Then Google added nofollow, so the chances were the only person who would ever find your site through comments were the owner of the site you posted on - if you were very lucky and it wasn't buried in a sea of spam or one-shot snarks.

Re: The Destruction of the Web

#7
post #3

I'm "relatively" new to the web. I started using it around 2004. To me, the only way to browse the web goes through Google. I don't think there's a single day I spent in front of the computer without me hitting Google at one point. I even use Google search when I'm specifically targeting Wikipedia or StackOverflow. For the people who got introduced to the web before me, how was "web browsing" done in the earlier deca…

It is precisely as you say: Simple as that. You knew of some sites, and you had links from those sites to others. Beyond that, you could take some big names and guess if they had a site. As for your "working internals of a car engine" example, in 1995 I would have instead cracked open a physical encyclopedia at a library. It really was the rise of search engines that helped grant the internet that "Learn/find anything about anything" aura.

All I remember interacting with until about 1997 was a handful of consumer brand's websites, the brands for which I already knew of and could easily guess their domain name.

Re: The Destruction of the Web

#8
post #4

This is a nice description of the hell of modern WWW. But things are worse than that! With few exceptions (Stack Exchange and Wikipedia are notable) most searches will return sites that have been SEOd. My Google search for [spectacles cases] returns these two sites on the first page: ( http://www.spectaclecases.co.uk/ ) ( http://www.aglassescase.co.uk/ ) > Welcome to SpectacleCases.co.uk. You will find a wide selecti…

I think Stack Exchange and Wikipedia are examples of SEO'd sites too. For example: Stack Exchange had problems with indexing before they added a sitemap [1]. Wikipedia takes a lot of care to create machine-indexable pages with quality content, that link back to their sources [2] and are linked to internally. Both follow SEO best practices as described in the Google Quality Guidelines [3]

Or your definition of "SEO'd" is leaning more to manipulative/sleazy.

In your example about the spectacle cases, would the other site being an affiliate make a (huge) difference? If Google punishes companies that have more than one website, people would shift to affiliates or try to hide they own the websites in other ways. In a way I feel multiple McDonalds in one city is perfectly possible. However, there needs to be a balance between the "best" websites (be they from the same owner or not) and diversity of the results (a query deserves diversity).

[1] http://www.codinghorror.com/blog/2008/10/the-importance-of-s... "We knew from the outset that Google would be a big part of our traffic, and I wanted us to rank highly in Google"

[2] http://www.mattcutts.com/blog/pagerank-sculpting/ "In the same way that Google trusts sites less when they link to spammy sites or bad neighborhoods, parts of our system encourage links to good sites."

[3] http://support.google.com/webmasters/bin/answer.py?hl=en&ans...

Re: The Destruction of the Web

#9
post #3

I'm "relatively" new to the web. I started using it around 2004. To me, the only way to browse the web goes through Google. I don't think there's a single day I spent in front of the computer without me hitting Google at one point. I even use Google search when I'm specifically targeting Wikipedia or StackOverflow. For the people who got introduced to the web before me, how was "web browsing" done in the earlier deca…

> how would you look for specific information (like, for instance, how would you research the working internals of a car engine for a school project?)

Maybe I'm just an old fart (I'm 26 btw :P ) but I went to the library.

Re: The Destruction of the Web

#10
Good article, but the author takes a somewhat generous view of academic citations. There are spammers in academia -- e.g. editors and reviewers who block rivals from publishing or who demand citations of their own work in revisions.
Post reply on HN