Live data from Hacker News

How SEO Ruined the Internet

superhighway98.com

91–100 of 296 posts

Re: How SEO Ruined the Internet

#91

The article doesn't present complete facts. Regarding zero-sum game, this perhaps was true in the old pagerank algorithm. But I'd believe Google's ranking algorithm has advanced beyond simple keyword density, passing links. What I've noticed is it now gives much more emphasis to user experience. (With metrics like bounce rate meaning the searcher didnt find what he looked for and went back to search results) We all l…

How does Google determine bounce rate? That I click on another search result after I click on the first one?

Re: How SEO Ruined the Internet

#92
post #40

Earlier quoted context omitted.

> "For years it has been next to impossible to get a result that is faithful to the search you actually typed in." Good lord, yes. If I type two words, I want preference for sites that contain both of them, yet the first results all have either one or the other, because surely I must be more interested in a popular site that uses only one of these, right? Google is sometimes too smart, trying to interpret exact words…

> I don't think AI is in any danger of taking over the world just yet. The scary thing about AI is that, even as the algorithms have greater and greater intelligence, we're still not much closer to teaching them to do what we want them to do. They can game the system better than ever, and then the universe is tiled with surgical masks.

Maybe this is how AC finally reversed entropy.

https://www.multivax.com/last_question.html

Re: How SEO Ruined the Internet

#93

Google with default settings is useless and using more sophisticated queries quickly walls off the user with increasingly annoying captcha. The "world's knowledge under you fingertips" motto is still valid and brilliant though. My personal solution is library of OCR-ed PDFs with most established books from various domains, git repository for each domain. Greppable in miliseconds, locally. Hijack this, SEO experts!

do you have a link to this solution

ImageMagick and Tesseract for OCR-ing each page of a PDF into a separate text file (through TIFF image format, disregard the huge TIFFs afterwards), private git repos for hosting, then ag/grep for searching. Not as easy to find the phrase back in PDF as with eg. Google Books, but then GB with copyright related content restrictions is useless most of the time.

Re: How SEO Ruined the Internet

#94
post #40

Earlier quoted context omitted.

> "For years it has been next to impossible to get a result that is faithful to the search you actually typed in." Good lord, yes. If I type two words, I want preference for sites that contain both of them, yet the first results all have either one or the other, because surely I must be more interested in a popular site that uses only one of these, right? Google is sometimes too smart, trying to interpret exact words…

> I don't think AI is in any danger of taking over the world just yet. The scary thing about AI is that, even as the algorithms have greater and greater intelligence, we're still not much closer to teaching them to do what we want them to do. They can game the system better than ever, and then the universe is tiled with surgical masks.

So if AI ever takes over the world and kills us all, it will probably be because it failed to understand what we actually wanted.

Re: How SEO Ruined the Internet

#95

Google with default settings is useless and using more sophisticated queries quickly walls off the user with increasingly annoying captcha. The "world's knowledge under you fingertips" motto is still valid and brilliant though. My personal solution is library of OCR-ed PDFs with most established books from various domains, git repository for each domain. Greppable in miliseconds, locally. Hijack this, SEO experts!

How did you acquire the books? Are they in the public domain, or did you have to buy them? In either case is there a place to acquire/buy these books massively or did you do it one by one manually?

Re: How SEO Ruined the Internet

#96

I think most spam websites today could be filtered out with very simple algorithms. But that would lead to fewer people ending up on these websites and clicking ads. So if your search engine is also an ad network, filtering out spam websites is not in your interest.

This is the real problem.

SEO will always be a game of cat and mouse. The original algorithms were designed to surface useful content relevant to the search query with the limitations of the technology at the time (so they could be gamed).

Nowadays technology has improved and processing power is much cheaper so it should be possible to use machine learning to recognise what’s “good” and what’s SEO spam and thus get ahead of the SEO crowd again.

The problem here is that the spam sites are also the ones with ads (often Google ads), so there is no financial incentive for Google to actually do anything about those.

Re: How SEO Ruined the Internet

#97

The article doesn't present complete facts. Regarding zero-sum game, this perhaps was true in the old pagerank algorithm. But I'd believe Google's ranking algorithm has advanced beyond simple keyword density, passing links. What I've noticed is it now gives much more emphasis to user experience. (With metrics like bounce rate meaning the searcher didnt find what he looked for and went back to search results) We all l…

How does Google determine bounce rate? That I click on another search result after I click on the first one?

Google search result links don’t link to the site directly but go through a Google-provided redirect that presumably has a reference to the original search query.

If you were to go back to the same search result page and click on another result within a short timeframe they will assume you “bounced”.

They also have Google Analytics littering the majority of the web, so I’m assuming that gives them a signal as well.

Re: How SEO Ruined the Internet

#99

Google with default settings is useless and using more sophisticated queries quickly walls off the user with increasingly annoying captcha. The "world's knowledge under you fingertips" motto is still valid and brilliant though. My personal solution is library of OCR-ed PDFs with most established books from various domains, git repository for each domain. Greppable in miliseconds, locally. Hijack this, SEO experts!

How did you acquire the books? Are they in the public domain, or did you have to buy them? In either case is there a place to acquire/buy these books massively or did you do it one by one manually?

Whichever most convenient way to obtain a full restriction-free PDF of a book. Fetched one by one through various channels in my case. Very few are in the public domain, if any. BTW one can do the same with academic publications, device manuals, or whatever else content available in PDF.

Re: How SEO Ruined the Internet

#100
post #5

I totally agree. Google has become useless for about half my searches. It gives me only the biggest, most commercial or most popular results. Anything obscure is impossible to find. I'd like to have a search engine where you get only the most obscure, hard-to-find content. One where you can tweak the kind of content you're looking for, or even switch between different modes: am I just looking for the definition of a…

I wish I would have seen this coming years ago. I would have built a Google Custom Search Engine, and every time I ran into a good website, added it to the whitelist. By now, it would probably be alright.
Post reply on HN