New ways we're tackling spammy, low-quality content on Search
121–130 of 205 posts
Re: New ways we're tackling spammy, low-quality content on Search
#122The best way to improve Google search results is to stop and use Kagi[0]. Seriously. If you are reading this article and these comments then you must be somewhat interested in search results. Do yourself (and the search world) a favor and subscribe. [0] https://kagi.com
There is so much Kagi references on HN that I’m slowly beginning to feel sus about it. I’m not saying this is happening (I’m not saying anyone is a shill). It’s just making me feel a little “hmm” about it.
However, when I'm trying to find some archaic technical solutions that I know is out there on the web somewhere, Google just can't really do the job any more at all. Exact search doesn't even work! Kagi's filters DO work, and the ability to black list stackoverflow scrapers cuts out a lot of the useless noise. It's at those times when Kagi earns its place.
Re: New ways we're tackling spammy, low-quality content on Search
#123The best way to improve Google search results is to stop and use Kagi[0]. Seriously. If you are reading this article and these comments then you must be somewhat interested in search results. Do yourself (and the search world) a favor and subscribe. [0] https://kagi.com
There is so much Kagi references on HN that I’m slowly beginning to feel sus about it. I’m not saying this is happening (I’m not saying anyone is a shill). It’s just making me feel a little “hmm” about it.
Re: New ways we're tackling spammy, low-quality content on Search
#124The expired domain abuse thing --- that's interesting because finding that signal was one of the few things I got done while I was at Google... in November 2006! And I vaguely remember being told that the fix went into the codebase after I identified that it was happening.
Re: New ways we're tackling spammy, low-quality content on Search
#125I would really love for Google to provide advanced search capabilities or filters to superusers who know EXACTLY what they are searching for. I wouldn't even mind paying for it tbh.
Re: New ways we're tackling spammy, low-quality content on Search
#126Re: New ways we're tackling spammy, low-quality content on Search
#127The expired domain abuse thing --- that's interesting because finding that signal was one of the few things I got done while I was at Google... in November 2006! And I vaguely remember being told that the fix went into the codebase after I identified that it was happening.
Re: New ways we're tackling spammy, low-quality content on Search
#128As AI generated content takes over the web, algorithmic search will become increasingly useless. In fact, for many things - like what to buy - it already is useless. Your best bet is asking your family/friend/neighborhood groups. If I had trillions like google does, I'd invest in hiring a whole bunch of humans to curate a search index that is actually useful. It's something a new player will never be able to scale up…
This honestly seems unrelated to their poor search quality—hell, I'm even open to AI generated content for some queries. I blame catering to their clients and attempting to manipulate the content on the internet in the name of SEO for why I find useful results buried beneath products and ads.
Re: New ways we're tackling spammy, low-quality content on Search
#129Until Pinterest is gone from search results, I won’t believe them
Re: New ways we're tackling spammy, low-quality content on Search
#130As AI generated content takes over the web, algorithmic search will become increasingly useless. In fact, for many things - like what to buy - it already is useless. Your best bet is asking your family/friend/neighborhood groups. If I had trillions like google does, I'd invest in hiring a whole bunch of humans to curate a search index that is actually useful. It's something a new player will never be able to scale up…
Right, LLMs will change the world, we know. Algorithmic search has been useless for 5+ years. LLMs will just generate those spam pages faster, but they're already there.
They obviously have the training data. Even if they needed to manually label a bunch of it.
The problem must be that they make too much money from the web pages that are just "500 words rephrasing the headline plus constantly-refreshing ads between every paragraph"