Live data from Hacker News

Search engines and SEO spam

twitter.com

211–220 of 555 posts

Re: Search engines and SEO spam

#211
Okay, given that we have pretty successful examples of wikipedia as a general crowdsourced information storage and stackoverflow as a specialized domain crowdsourced Q&A site, would it be impossible to build a crowdsourced search engine? Not even scraping the web, but I would just type my search term, if that is already searched and results voted, I would see those. If it wasa completely new search term, I would get no immediate results, but my search would be displayed in "new searches page", which some voluntary people would be following and trying to add relevant results.

Re: Search engines and SEO spam

#212
post #16

Earlier quoted context omitted.

Nerdwallet and wikiHow are both SEO spam content farms. They just happen to have above-average quality content. They don't exist without a search engine.

Not sure how you're differentiating "seo spam content farm" from "website"

That is part of the problem

Re: Search engines and SEO spam

#213
post #135

Earlier quoted context omitted.

The problem isn't that Google doesn't employ these people or invest in their activities. It's that Google has destroyed their own search results in order to continue to expand their revenue opportunities. If Google: - Enabled downvoting on results, like YT videos. (Has its own spam problems, just like YT) - Allowed you to block certain domains from your search results, like YT videos. (If they added some kind of "coo…

The problem with all of this is it would help us greatly, but it would be useless to the 99% that the internet is increasingly being designed for. Modern UI trends are becoming obsessed with removing as many options and features as possible so the dumbest humans bordering on smartest vegetables can still use the service.

And customization breaks caching.

Re: Search engines and SEO spam

#214
post #155
post #115

Earlier quoted context omitted.

> If you want a review for something that came out today, there is no way that work could have been done, so there simply isn't anything to find. That's not strictly true, given that reviewers are often sent pre-release versions of things in order to do that work before release day.

Not sure why you're being downvoted, as you're correct - however to point out there seems to be a trend where reviewers are only given pre-release versions if they practically always give favourable reviews to the products they list, especially if they're provided the product for free; there doesn't have to be an express relationship or contract between a reviewer and a company either, it's the reverse of how Bill Ga…

Yeah, that's been a problem with reviews for a long time. In fact it's what Consumer Reports used initially to differentiate themselves: their "thing" was that they only reviewed products bought anonymously at retail (no free samples or manufacturer-provided review items) and didn't accept any advertising from manufacturers either.

Sites that receive free review samples and are supported by affiliate links are kind of the exact opposite model.

Re: Search engines and SEO spam

#215

Earlier quoted context omitted.

The problem isn't that Google doesn't employ these people or invest in their activities. It's that Google has destroyed their own search results in order to continue to expand their revenue opportunities. If Google: - Enabled downvoting on results, like YT videos. (Has its own spam problems, just like YT) - Allowed you to block certain domains from your search results, like YT videos. (If they added some kind of "coo…

How do you fight brigading, the organization of groups elsewhere to collectively vote on something? Eg white supremacist groups get together and vote down everything by people of color, and vote up their pages about how great they are?

Randomly select votes that are actually recorded. Then add in metavoting that votes on the votes with random sampling. At Google's scale with a sufficiently random sampling you'd be extremely hard pressed to successfully brigade or spam the voting.

Google could easily use its current fingerprinting to constrain (to an extent) multiple votes. Even knowing only a portion of the population will participate in the voting they can use a Wilson confidence interval[0] or similar to properly weight votes.

Random sampling works here since you're not guaranteed one vote per user per page and the outcome in binomial, seen and downvoted or seen and not downvoted.

[0] https://www.mikulskibartosz.name/wilson-score-in-python-exam...

Re: Search engines and SEO spam

#216
Google search results are garbage, at least from a developer's perspective.

Most of the results are poorly formatted content "gathered" from stackoverflow, github, quora, etc.

And from a "person who wants to see an image" perspective, Google is purely a gateway to Pinterest or Gettyimages.

Re: Search engines and SEO spam

#217

Earlier quoted context omitted.

The problem isn't that Google doesn't employ these people or invest in their activities. It's that Google has destroyed their own search results in order to continue to expand their revenue opportunities. If Google: - Enabled downvoting on results, like YT videos. (Has its own spam problems, just like YT) - Allowed you to block certain domains from your search results, like YT videos. (If they added some kind of "coo…

> That would be incredibly valuable. They already have most of the tech. They could even create a subscription service around custom search engines if they really wanted. Plenty of people would find something like that incredibly valuable. Why would they do this? Google's customers are the advertisers, not the end-users. And no one is going to pay for a search engine, it's been tried and has failed.

if you think about it, Google provides advertisers a customized search engine to find customers. So it is not you searching the web, it is web's advertisers searching leads

Re: Search engines and SEO spam

#218
post #183

For medical search the answer is pubmed. Not only is the collection of documents clean (of low-grade scammers, pharma companies have to pay big $ to play) but the NIH has done a large amount of search quality and ontology work -- the system knows "Tylenol" is synonymous with "Paracetamol", "Acetaminophen", etc.

> the system knows "Tylenol" is synonymous with "Paracetamol", "Acetaminophen", etc. This is the exactly the kind of thing that Google cannot fathom manually doing. As if entering facts into a computer were morally wrong somehow. They'd much rather launch the equivalent of a shell script that harnesses face-melting amounts of computational power, processing literally trillions of webpages in bulk, signal and noise to…

Google has long lied about what they do.

I had a chance to debrief people who had left their relevance team and they told me things that were outright contradictory to what rank-and-file Google employees have told me. (What they told me did make sense in terms of my experience as an IR system developer, SEO publisher, etc.)

Microsoft bought a company called PowerSet that had extracted a large database of entities and relationships from Wikipedia and used the technology to make the "Bing" search engine.

Earlier Microsoft engines were a joke, but Bing was so good that Google saw it as a threat so they bought Freebase to get a similar kind of database, then they killed it to incorporate it into the "Google Knowledge Graph".

For all of their hating on semantics note that they hired R. V. Guha as their chief scientist, who worked with Doug Lenat on the notorious

https://en.wikipedia.org/wiki/Cyc

Re: Search engines and SEO spam

#219
post #210

The quote-Tweeted thread mentioned recipes as one of the things that has been SEObliterated. It's a great example of the problem, and also a great example of the problems any solution will encounter. Recipes have become a bellwether Internet problem. In the past, your great-grandmother had a card file with a bunch of 3x5 index cards with the ingredients and instructions on how to make everything, and they pretty much…

>I'm not sure there is any technological solution The technological solution would be to stop rewarding them for these monstrosities. One of the main motivator for turning a short recipes into a 19 page essai about the chef's life is that more words = better ranking.

And funny enough, it's obvious that Google's engineers know this because they're adding more small, self-hosted featured results to the top of the page all the time.

Re: Search engines and SEO spam

#220
post #31

Earlier quoted context omitted.

Also, i'd be very surprised if they didn't have tens of thousands of workers aiding in spam review already. The hard part in all of this isn't finding and stopping spam - it's defining what spam is. Are all the pie recipes where there's a 2000 word essay about their grandma at the top 'spam'? They still have the recipe, and Google Home devices pick up the recipe instructions just fine so people end up not reading it,…

AFAIK the 2000 word essays in recipes are Google's fault - it prioritizes pages with a lot of content, so you have to add that junk to the top in order to rank highly. While I'm sure there's more going on behind the scenes than I'm aware of, it does seem like the rules could be altered on a category-specific basis where a lot of text isn't necessarily a positive.

This reminds me of the page inflation that struck tech books during the late 1990s / early aughts. The Marketing Wisdom was that fat books sold (or took up more shelf space), so texts got padded with weak writing, gratuitous puffery, and other elements, which (much as the recipie essays) simply got in the way of delivering actual informative content.

(The fact that many of these books were rushed out with very poor quality control also didn't help.)

Post reply on HN