The consistent theme every time this comes up is that dealing with the sheer weight of the internet is almost impossible today. SEO spam is hard to fight and the index gets too heavy. However, I wonder if this is a sign that we're looking at the problem wrong. What if instead of even trying to index the entire web, we moved one step back towards the curated directories of the early web? Give users a search engine and…
what you are suggesting would make the problem of echo chamber (bubble) worse than it is today!
Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?
311–320 of 492 posts
Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?
#312Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?
#313The consistent theme every time this comes up is that dealing with the sheer weight of the internet is almost impossible today. SEO spam is hard to fight and the index gets too heavy. However, I wonder if this is a sign that we're looking at the problem wrong. What if instead of even trying to index the entire web, we moved one step back towards the curated directories of the early web? Give users a search engine and…
That's basically what I'm doing with my search "site:reddit.com" I wonder if anyone at Google is aware of this trend and taking notes.
Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?
#314The consistent theme every time this comes up is that dealing with the sheer weight of the internet is almost impossible today. SEO spam is hard to fight and the index gets too heavy. However, I wonder if this is a sign that we're looking at the problem wrong. What if instead of even trying to index the entire web, we moved one step back towards the curated directories of the early web? Give users a search engine and…
That's basically what I'm doing with my search "site:reddit.com" I wonder if anyone at Google is aware of this trend and taking notes.
Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?
#315The consistent theme every time this comes up is that dealing with the sheer weight of the internet is almost impossible today. SEO spam is hard to fight and the index gets too heavy. However, I wonder if this is a sign that we're looking at the problem wrong. What if instead of even trying to index the entire web, we moved one step back towards the curated directories of the early web? Give users a search engine and…
This might work well in some situations (e.g. research, development), however it would also increase the effect of echo chambers I think.
Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?
#316Earlier quoted context omitted.
Have you ever looked at the Amazon file? I'll see if I can track down the link but I remember somebody sharing a dump with me from Amazon that apparently was a recent scrape. Edit: https://registry.opendata.aws/commoncrawl/
That's Common Crawl, they do the spidering of some billions of webpages but that's still a tiny percentage of the web versus Google or Bing.
Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?
#317The consistent theme every time this comes up is that dealing with the sheer weight of the internet is almost impossible today. SEO spam is hard to fight and the index gets too heavy. However, I wonder if this is a sign that we're looking at the problem wrong. What if instead of even trying to index the entire web, we moved one step back towards the curated directories of the early web? Give users a search engine and…
Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?
#318Ha, yes, I've done that at https://gigablast.com/ . The biggest problems now are the following: 1) Too hard to spider the web. Gatekeeper companies like Cloudflare (owned in part by Google) and Cloudfront make it really difficult for upstart search engines to download web pages. 2) Hardware costs are too high. It's much more expensive now to build a large index (50B+ pages) to be competitive. I believe my algorithms…
If you're serious about this, add a paid tier. Until it's free, I don't trust you will not ever sell my data to make bank.
Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?
#319Ha, yes, I've done that at https://gigablast.com/ . The biggest problems now are the following: 1) Too hard to spider the web. Gatekeeper companies like Cloudflare (owned in part by Google) and Cloudfront make it really difficult for upstart search engines to download web pages. 2) Hardware costs are too high. It's much more expensive now to build a large index (50B+ pages) to be competitive. I believe my algorithms…
Regarding the Gatekeeper companies like Cloudflare, it sounds like anti-competitive behavior that could potentially be targeted with anti-trust legislation, correct?
Cloudflare is giving it's customers what they want. They don't want all kinds of bots claiming to be search engines crawling their sites. Cloudflare isn't hurting cloudflare competitors by doing this. Cloudflare isn't hurting their customers by doing this. To repeat - most websites don't want lots and lots of crawlers. They want the 2 or 3 which matter and no more, because at some point it's difficult to tell what the crawler is doing... (is it a search engine???). They aren't obliged to help search engines. Even if Cloudflare wasn't offering this, bigger customers would roll their own and do.. more or less the same thing.
Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?
#320The consistent theme every time this comes up is that dealing with the sheer weight of the internet is almost impossible today. SEO spam is hard to fight and the index gets too heavy. However, I wonder if this is a sign that we're looking at the problem wrong. What if instead of even trying to index the entire web, we moved one step back towards the curated directories of the early web? Give users a search engine and…