Live data from Hacker News

Search engines and SEO spam

twitter.com

61–70 of 555 posts

Re: Search engines and SEO spam

#61
post #5

The funny thing is that if the people who worked on spam at Google were free to talk about it, I'm sure it would become evident that they know more about spam and anti-spam efforts than anybody else in existence. It's a ridiculously hard problem, especially when people are targeting you directly. But they aren't free to talk about it, because if they did it would just give more assistance to the spammers, and make th…

Also, i'd be very surprised if they didn't have tens of thousands of workers aiding in spam review already. The hard part in all of this isn't finding and stopping spam - it's defining what spam is. Are all the pie recipes where there's a 2000 word essay about their grandma at the top 'spam'? They still have the recipe, and Google Home devices pick up the recipe instructions just fine so people end up not reading it,…

> and Google Home devices pick up the recipe instructions just fine so people end up not reading it

I think this isn't entirely related, but that's perhaps the beginning of a bias you might end up having that everyone experiences technology in the same way as it marches on. I've yet to encounter a Google Home in the wild, I imagine far more people are consuming recipes on phones, tablets and PCs.

Re: Search engines and SEO spam

#62
He’s just describing Webrings, except in a reactive tense (“filter out spam sites”) rather than a proactive tense (“associate your site with other worthwhile sites”). Google’s ranking algorithm only works when someone is proactively curating, and only SEO spammers do so these days. Reactive curation is not a viable way to manage information.

The simplest way to compete with Google is to create a DIY Webrings site that disallows harvesting of data by Google. Charge curators to create a webring, and let curators select three hashtags and a description that represent their list of fifty or fewer sites. Use the revenue to pay a human to curate the list of hashtags, and let users tip a webring curator in gratitude with an Apple Pay button.

This is how to make a million dollars, Pinboard-style, out of the ashes of the original curated Yahoo idea and the information structures of hashtagging. It doesn’t work if you allow free-for-all infinite-sized lists, it doesn’t work if you allow free-for-all hashtags, but with clear limits and moderation of tags (instead of webrings), it would thrive. By moderating tags, users can keep the webring they paid for, and SEO rings will be stick out for having no shared network with any other rings, which allows for easier detection and culling of malicious non-participatory actors. Plus, with the curation networks in place, it becomes possible to bubble up rings that have unusual content for positive human moderation activity.

I tried to find some good podcast lists yesterday and each site I visited had a really interesting cross-section, but there were so many duplicates. I wish the ring site existed, so that it could remember what it had shown me already, and I could say “show me rings that intersect with this podcast and have something new I haven’t seen before”.

That’s where the theory of pagerank and the practice of curation and the capabilities of search align, and given that moderation of hashtags scales very cheaply, is a billion dollar opportunity that Google and Amazon cannot compete with if handled properly. It’s not about trying to get a cut of every visit’s revenue potential. It’s about giving human beings a directory that respects their time and remembers what they’ve seen.

Re: Search engines and SEO spam

#63

Earlier quoted context omitted.

This x100000. There is no scenario - none - where thousands of engineers at Google working on search wake up in the morning and say "we sure have made it good enough wr2 SPAM. I think I'll have another Danish."

When the cafes were open, you can bet they said, "I'll have another Danish, and then get back to work on this problem that never seems to go away."

I sure wish I had problems that were totally unsolvable, they are so easy to measure progress on. /sarcasm

I think it’s more likely that because they are just building hundreds of tiny tweak experiments and it’s someone else who desides what to build and if it even worked. Search quality is such a meta-problem that it goes beyond any real hope of simply working on it in anything beyond piecemeal trial and error fashion on their dataset.

Re: Search engines and SEO spam

#64
post #9

>This may not just be a problem with Google but possibly also the recipe for beating Google. A startup usually has to start with a niche market. Why not try writing a search engine specifically for some category dominated by SEO spam? >You might need to do a lot of manual spam fighting initially. That could be both the thing-that-doesn't-scale, and the thing that differentiates you by being alien to Google's DNA. (Th…

And makes me think that StumbleUpon had a similar curation ability, in that the value qualifier is how often [hopefully] real people interact with content - tracked by who's using SU and agreed to allow tracking; can't remember if sharing that was optional or not?

The gamification of the system then would have to come through onboarding fake users, pretending/mimicking real user behaviour to send that signal into the system; not sure if SU ever ran into that problem or was actively paying attention to trying to identify and removing fake or suspicious signals from their output?

I feel a much better system is easily within reach, it's simply getting the right structure to it, the right foundation, and then it will quickly take off due to the quality difference. I've already figured out a design pattern that Twitter and Facebook has indoctrinated us with, making us think it is normal - and keeping us blind to an actual normal way or organizing or communicating, but that isn't conducive to control or ad revenues - and so extending my future plans to include a better search-directory system would fit snugly into my efforts.

Re: Search engines and SEO spam

#65
post #9

>This may not just be a problem with Google but possibly also the recipe for beating Google. A startup usually has to start with a niche market. Why not try writing a search engine specifically for some category dominated by SEO spam? >You might need to do a lot of manual spam fighting initially. That could be both the thing-that-doesn't-scale, and the thing that differentiates you by being alien to Google's DNA. (Th…

Some of the examples used in the Twitter thread Paul was referring to would be better served by a manually curated directory service with a possible addition of a search engine only surfacing content from the sites in the directory.

For health information and recipes in particular there are only a handful of really high quality sites that have quality content for 95% of the information most people need. I bet if you wanted to increase the coverage to 99%, that list would expand to less than a thousand sites. At those numbers manually curating the information would be easily achievable.

How to get people to use your top notch Google replacement instead of Google, however. That's the hard problem.

Re: Search engines and SEO spam

#66
post #9

>This may not just be a problem with Google but possibly also the recipe for beating Google. A startup usually has to start with a niche market. Why not try writing a search engine specifically for some category dominated by SEO spam? >You might need to do a lot of manual spam fighting initially. That could be both the thing-that-doesn't-scale, and the thing that differentiates you by being alien to Google's DNA. (Th…

He's right as often as he's wrong

Re: Search engines and SEO spam

#67
Just a minute ago, I made a small typo in a non-obscure programming-related search term.

    Showing results for searchterm
    No results found for searchterm
Followed by an unending list of random celebrities I don't know nor care about, businesses I've never been that sell items I have absolutely no use of, and random foreign news articles.

Failure to recognise the typo is unexpected but forgivable. But then, rather than helping me with my search, they attempt to distract and lead me away from it - using triggers that you'd think they should have known wouldn't work.

I really don't understand how this is even possible, and it's not a rare occurrence.

Re: Search engines and SEO spam

#68
post #5

The funny thing is that if the people who worked on spam at Google were free to talk about it, I'm sure it would become evident that they know more about spam and anti-spam efforts than anybody else in existence. It's a ridiculously hard problem, especially when people are targeting you directly. But they aren't free to talk about it, because if they did it would just give more assistance to the spammers, and make th…

This x100000. There is no scenario - none - where thousands of engineers at Google working on search wake up in the morning and say "we sure have made it good enough wr2 SPAM. I think I'll have another Danish."

I agree with this and the grandparent comment wholeheartedly. That said, there's a kind of institutional blindness that can build up in companies—especially ones that dominate their sector. It may have roots in intransigent upper management, ossified and inflexible process, wide-scale burnout, a culture of passing the buck, or any number of other pathologies.

I don't claim that Google has any of these and certainly have no insight into their search group. But I've personally been at powerful companies with best-of-the-best talent that were blind to the decay in their own living room, so I would caution against immediate dismissal of PG's take.

Re: Search engines and SEO spam

#69
post #5

The funny thing is that if the people who worked on spam at Google were free to talk about it, I'm sure it would become evident that they know more about spam and anti-spam efforts than anybody else in existence. It's a ridiculously hard problem, especially when people are targeting you directly. But they aren't free to talk about it, because if they did it would just give more assistance to the spammers, and make th…

Also, i'd be very surprised if they didn't have tens of thousands of workers aiding in spam review already. The hard part in all of this isn't finding and stopping spam - it's defining what spam is. Are all the pie recipes where there's a 2000 word essay about their grandma at the top 'spam'? They still have the recipe, and Google Home devices pick up the recipe instructions just fine so people end up not reading it,…

Let's have niches where the content is hand curated by human beings instead of pure statistics by machines.

Hmm why stop there let's actually make the users do the curating and even the content creation by rewarding them with social validation. Let’s have hard working moderators who work on the community full time.

Then we could just build a search engine over it. We could call it Reddit. Or HackerNews.

Maybe the users aren't all as good as professionals at curating the information. Let's hire professionally trained curators pay them well and we could call them newspapers. Then we can come in disrupt them and replace them with an algorithmic marketplace that eventually becomes infested with click bait.

Re: Search engines and SEO spam

#70
post #5

The funny thing is that if the people who worked on spam at Google were free to talk about it, I'm sure it would become evident that they know more about spam and anti-spam efforts than anybody else in existence. It's a ridiculously hard problem, especially when people are targeting you directly. But they aren't free to talk about it, because if they did it would just give more assistance to the spammers, and make th…

> But they aren't free to talk about it, because if they did it would just give more assistance to the spammers, and make the problem worse.

The reality is more that some Google engineer will come up with an algorithm change that makes the result 40% better, but it will come at the expense of making that search 3ms slower so the change won't get merged. Or it will make the results worse for some niche set of queries that the business team really cares about, so again it won't get merged.

There are lots of consumers who would gladly pay $1 a month or whatever in order to use a couple extra milliseconds of compute power per per search in exchange for drastically better results, so there is lots of room for a startup to compete.

Post reply on HN