Live data from Hacker News

Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

news.ycombinator.com

221–230 of 492 posts

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#221
post #13

Earlier quoted context omitted.

But shouldn't all the blogspam be so hyperoptimized for Google's algorithm that is should be straightforward to detect and ignore/downrank it?

No because google's algorithm is not well known publicly. Also, if it was straightforward to detect then google could downrank it as well.

I wonder if you could evaluate a page using your own algorithm, which is probably not gamed as much as Google's (because who cares about your search engine?)

Then, check Google's ranking of the page. If it is much higher than it seems the page should be, assume the page is being SEO hyper-optimized and penalize the page proportionately.

Basically, using the variance between Google's model and your model as an indicator of an SEO spam page.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#222
post #78

Earlier quoted context omitted.

It is. The alternative is scooping everything and using algos to curate. That seems worse imo.

Perhaps vote on results like on Reddit posts? Gets the junk sites down (and out of the index eventually).

That just means that you have to curate the people allowed to vote. Otherwise, it would be rule by the obsessed and the search engine optimizers, and the junk sites will dominate the index.

I'm not convinced that Google's recursive AI algos aren't a functional equivalent. They let you vote by tracking your clicks.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#223
post #180

Everyone runs in the other direction anytime a search engine is mentioned. The thought of competing with Google turns people off. Even in 2021, despite how bad it's become, it's still miles ahead of other competitors.

I disagree. A lot of people I know already switched to Duckduckgo. Google’s ability to get relevant results is dropping like a brick, while the quality of DDG has been improving slowly but steadily.

I wish I could agree but from my experience, DDG's search results aren't really that great. Often even worse than Google's.

And another private company is not the answer I believe. We need something more drastic, an open-source search engine organized as a genuine non-profit organization. Something like that. Otherwise, whatever replaces Google will just turn into another Google as soon as it gets any momentum.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#224
post #195

Earlier quoted context omitted.

Yes and no. A lot of those sites are small local businesses trying to get found. A front page listing can be the difference between surviving and going under. Much of the time the blog spam is what floats hours, contact info, and services provided to the first page.

Be that as it may, search ranking is a zero sum game. The unfair advantage SEO gives this particular struggling business means another goes under. I'd rather punish the guy trying to game the system than the one with enough principles not to.

The difference is far more likely to be in capability or expertise than principles.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#225

Earlier quoted context omitted.

Be that as it may, search ranking is a zero sum game. The unfair advantage SEO gives this particular struggling business means another goes under. I'd rather punish the guy trying to game the system than the one with enough principles not to.

The difference is far more likely to be in capability or expertise than principles.

Either way, capability for fuckery is not something I'd want to encourage.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#226
post #12

Ha, yes, I've done that at https://gigablast.com/ . The biggest problems now are the following: 1) Too hard to spider the web. Gatekeeper companies like Cloudflare (owned in part by Google) and Cloudfront make it really difficult for upstart search engines to download web pages. 2) Hardware costs are too high. It's much more expensive now to build a large index (50B+ pages) to be competitive. I believe my algorithms…

maybe just add small webpages into your index, dont bother yo execute JS and dont download any images. The content quality will be higher and it's a lot cheaper.

Out of curiosity, why would not executing JavaScript or not downloading images equal higher content quality?

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#227
post #12

Ha, yes, I've done that at https://gigablast.com/ . The biggest problems now are the following: 1) Too hard to spider the web. Gatekeeper companies like Cloudflare (owned in part by Google) and Cloudfront make it really difficult for upstart search engines to download web pages. 2) Hardware costs are too high. It's much more expensive now to build a large index (50B+ pages) to be competitive. I believe my algorithms…

I tried out four search words with your search engine, and I am not convinced that it is mainly the index size and not the algorithm that is to blame for bad search results. There are way too much high ranking false positives. Here is what I tried:

  a) "Berlin": 

  1. The movie festival "Berlinale"
  2. The Wikipedia entry about Berlin
  3. Something about a venue "Little Berlin", but the link resolves to an online gaming site from Singapure
  4. "Visit Berlin", the official tourism site of Berlin
  5. The hash tag "#Berlin" on Twitter
  6. "1011 Now" a local news site for Lincoln, Nebraska
  7. "Freie Universität Berlin"
  8. Some random "Berlin" videos on Youtube
  9. The Berlin Declaration of the Open Access Initiative
  10. Some random "Berlin" entries on IMDb
  11. A "Berlin" Nightclub from Chicago
  12. Some random "Berlin" books on Amazon
  13. The town of Berlin, Maryland
  14. Some random "Berlin" entries on Facebook
  15. The BMW Berlin Marathon
  
  b) "philosophy"

  1. The Wikipedia entry about philosophy
  2. "Skin Care, Fragrances, and Bath & Body Gifts" from philosophy.com
  3. "Unconditional Love Shampoo, Bath & Shower Gel" from philosophy.com
  4. Definition of Philosophy at Dictionary.com
  5. The Stanford Encyclopedia of Philosophy
  6. PhilPapers, an index and bibliography of philosophy
  7. The University of Science and Philosophy, a rather insignificant institution that happens to use the domain philosophy.org
  8. "What Can I Do With This Major?" section about philosophy
  9. Pages on "philosophy" from "Psychology Today". I looked at the first and found it to be too short and eclectic to be useful.  
  10. The Department of philosophy of Tufts University
  
  c) "history"

  1. Some random pages from history.com
  2. "Watch Full Episodes of Your Favorite Shows" from history.com
  3. Some random pages from history.org
  4. "Battle of Bunker Hill begins" from history.com
  5. Some random "History" pages from bbc.co.uk
  6. Some random pages from historyplace.com
  7. The hash tag "#history" on Twitter
  8. The Missouri Historical Society (mohistory.com)
  9. Some random pages from History Channel
  10. Some random pages from the U.S. Census Bureau (www.census.gov/history/)
  
  d) "Caesar"

  1. The Wikipedia entry about Caesar
  2. Little Caesars Pizza 
  3. "CAESAR", a source for body measurement data. But the link is dead and resolves to SAE International, a professional association for engineering
  4. The Caesar Stiftung, a neuroethology institute
  5. Some random "Caesar" books on Amazon
  6. Hotels and Casinos of a Caesars group
  7. A very short bio of Julius Ceasar on livius.org 
  8. Texts on and from Caesar provided by a University of Chicago scholar
  9. (Extremely short) articles related to Caesar from britannica.com
  10. "Syria: Stories Behind Photos of Killed Detainees | Human Rights Watch". The photos were by an organization called the Caesar Files Group
  
So what I can see are some high ranked false positives that are somehow using the search term, but not in its basic meaning (a3, a11, b2, b3, d2, d3, d4, d6) or not even that (a6). Some results are ranking prominently although they are of minor importance for the (general) search term (a9, a13, b7, b8 -- perhaps a15 and d10). Then there are the links to the usual suspects such as Wikipedia, Twitter, Amazon, etc. (a2, a5, a8, a10, a12, a14, b7, c5, d1, d5); I understand that Wikipedia articles are featuring prominently, but for the others I would rather go directly to eg. Amazon when I am interested in finding a book (or use a search term like "Caesar amazon" or "Caesar books"). Well, and then there are the search results that are not completely off, but either contain almost no information, at least compared to the corresponding Wikipedia article and its summary (b4, b9, d7, d9), or that are too specific for the general search term (c1, c2, c3, c4, c6, c9, c10).

That leaves me with the following more or less high quality results (outside of the Wikipedia pages): a1, a4, a7, b5, b6, b10, and d8. The a15 and d10 results I could tolerate if there had been more high quality results in front of them; but as a fourth and second, respectively, good result they seem to me to be too prominent. Also in the case of "Berlin" a4 should have been more prominent than a1, and a7 is somewhat arbitrary, because Humbolt University and the Technical University of Berlin are likewise important; what is completely missing is the official Website of the city of Berlin (English version at www.berlin.de/en/).

All in all, I would say that your ranking algorithm lacks semantic context. It seems the prominence of an entry is mainly determined by either just being from the big players like Twitter, Youtube, Amazon, Facebook, etc. or by the search term appearing in the domain name or the path of the resource, regardless of the quality of the content.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#228
post #83

Earlier quoted context omitted.

I think some NLP is strictly beneficial for a search engine. You may think "grep for the web" sounds like a good idea, but let me tell you, having tried this, manually going through every permutation of plural forms of words and manually iterating the order of words to find a result is a chore and a half. Like, instead of trying PDP11 emulator PDP-11 emulator "PDP 11" emulator PDP11 emulators PDP-11 emulators "PDP 11…

I get that for general-purpose searches this is a good idea, but it would be nice if there was an easy way to disable this when you know you don't want it - for example, for most programming searches, if I type SomeAPINameHere the most relevant results will always be those that include my search term verbatim. I don't need Google to helpfully suggest "Did you mean Some API Name Here?", which will virtually always ret…

I feel your pain. Two workarounds when Google gets it wrong are to put the term in quotation marks, or to enable Verbatim mode in the toolbelt. (I know various people have come up with ways to add "Google Verbatim" as a search engine option in their browser, or use a browser extension to make Verbatim enabled by default.)

Disclaimer: I work on Google search.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#229
These guys [0] have built something really close to 2005-Google, and possibly slightly better.

The parent company, Tiscali, was a huge hit in the 1990s, as it provided internet access to millions of Italians. It went through some struggle for several years, but lately the original founder, Renato Soru, came back to run the company.

The company is based in Cagliari, the capital of Sardinia, Italy.

[0]: https://www.istella.it/en/

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#230
post #218

Earlier quoted context omitted.

> 1) Google is better at AI, for example let's take this sloppy search: "some joke where you can't tell if it is serious or joke" > It is called Poe's law, and Google returned it at #4. Bing or Duckduckgo don't have a clue... Interesting, I was looking for a good benchmark like this. For me Google returned it at #5 with an image/related terms carousel before it which places it physically more around #7 on the page. B…

The Brave results though seem to contain “good sites” whereas the Google results are content mill blogspam. The exact placement of Poe’s Law is somewhat less important.

I agree. I switched to Brave Search after running this test.
Post reply on HN