Ha, yes, I've done that at https://gigablast.com/ . The biggest problems now are the following: 1) Too hard to spider the web. Gatekeeper companies like Cloudflare (owned in part by Google) and Cloudfront make it really difficult for upstart search engines to download web pages. 2) Hardware costs are too high. It's much more expensive now to build a large index (50B+ pages) to be competitive. I believe my algorithms…
If you're serious about this, add a paid tier. Until it's free, I don't trust you will not ever sell my data to make bank.
Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?
281–290 of 492 posts
Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?
#282I think DuckDuckGo is closer to what you want. Same results for everyone, better privacy, and they're proactive about improving their results. https://duckduckgo.com/ Part of the problem is that there's a lot more low-quality content to wade through now than there was in 2005. I think the Google of 2005 would have trouble delivering quality results today also.
DuckDuckGo isn’t really a search engine, it’s a website that uses bing’s api.
Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?
#283Earlier quoted context omitted.
If you're serious about this, add a paid tier. Until it's free, I don't trust you will not ever sell my data to make bank.
You are going to pay for a generalized web search when DDG/Google/Bing/etc are free?
A clear pricing transaction sounds much nicer to me. Should generate better results too.
Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?
#284Ha, yes, I've done that at https://gigablast.com/ . The biggest problems now are the following: 1) Too hard to spider the web. Gatekeeper companies like Cloudflare (owned in part by Google) and Cloudfront make it really difficult for upstart search engines to download web pages. 2) Hardware costs are too high. It's much more expensive now to build a large index (50B+ pages) to be competitive. I believe my algorithms…
I wrote a "meta" search utility for myself that can query multiple search engines from the command line.^1 It mixes the results into a simplified SERP ("metaSERP"), optimised for a text-only browser, with indicators to show which search engine each result came from. The key feature is that it allows for what I might call "continuation searches". Each metaSERPs contains timestamps in its page source to indicate when searches were executed, as well as preformatted HTTP paths. The next search can thus pick up where the previous one left off. Thus I can, if desired, build a maximum-sized metaSERP for each query.
The reason I wrote this is because search engines (not GigaBlast) funded by ads are increasingly trying to keep users on page one, where the "top ads" are, and they want to keep the number of results small. That's one change from 2005 and earlier. With AltaVista I used to dig deep into SERPs and there was a feeling of comprehensiveness; leave no stone unturned. Google has gradually ruined the ability to perform this type of searching with their now secretive and obviously biased behind-the-scenes ranking procedures.
Why is there no way to re-order results according to objective criteria, e.g., alphabetical; the user must accept the search engines' ordering, giving them the ability to "hide" results on pages the user will never view or simply not return them. That design is more favorable to advertising and less favorable to intellectual curiosity.
Each metaSERP, OTOH, is a file and is saved in a search directory for future reference; I will often go back to previous queries. I can later add more results to a metaSERP if desired. I actually like that GigaBlast's results are different than other search engines. The variety of results I get from different sources arguably improves the quality of the metaSERP. And, of course, metSERPs can be sorted according to objective criteria.
This is, AFAIK, a different way of searching. The "meta-search engines" of yesteryear did not do "continuations", probably because it was not necessary. Nor was there en expectation that user would want to save meta-searches to local files. Users were not trying to minimise their usage of a website, they were not trying to "un-google".
Today's world of web search is different, IMO. There seems to be a belief that the operator of a search engine can guess what a user is searching for, that a user who sends a query is only searching for one specific thing, and that the website has an ad to match with that query. At least, those are the only searches that really matter for advertising purposes. Serendipitous discovery while perusing results is not contemplated in the design. By serendipitous discovery I do not mean sending a random query, e.g., adding an "I'm feeling lucky" button, which to me always seemed like a bad joke.
The only downside so far is I ocassionally have to prune "one-off" searches that I do not want to save from the search directory. I am going to add an indicator at search time that a search is to be considered "ephemeral" and not meant to be saved. Periodically these ephemeral searches can then be pruned from the search directory automatically.
1. Of course this is not limited to web search engines. I also include various individual site search engines, e.g., Github.
Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?
#285Earlier quoted context omitted.
If you're serious about this, add a paid tier. Until it's free, I don't trust you will not ever sell my data to make bank.
You are going to pay for a generalized web search when DDG/Google/Bing/etc are free?
If you don't pay, you are the product. Simple as that.
Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?
#286Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?
#287Earlier quoted context omitted.
> You do a search on Gigablast and say, well, why didn't it get this result that Google got. And that's because the index isn't big enough I wionder how much this is true, and how much (despite all our rhetoric to the contrary) it's because we have actually come to expect Google's modern proprietary page ranking, which counts more than just inbound links but all sorts of other signals (freshness, relevance to our pre…
I think people also have an inflated recollection of how good Google actually was back in 2005. Back then Google was only going up against indexes and link-rings, not 2021 Google/Bing/DDG/etc.
There was also indexes, which Yahoo, AOL (remember them!) had but there was, what was it called, dmoz?, the open web directory. When Google started, being in the right web directory gave you a boost in SERPs as it was used as a domain trust indicator, and the categories were used for keywords. Of course it got gamed hard.
Google was good, but I used it as an alt for maybe 6 months before it won over my main SE at the time. I've tried but can't remember what SE that was, Omni-something??
One of the main things Google had was all the extra operators like link: inurl:, etc., but they had Boolean logic operators too at one point I think.
Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?
#288Search engine isn’t singular, it’s plural.
(1) Search engine for something I know exists.
(2) Search engine for finding something new.
There’s a market for both, but you don’t have to solve both problems with the same product.
Sometimes I switch to Google for the former, but the latter works well enough for me that I don’t care what else Google would’ve shown me.
More often than not, my feeling is Google would only have shown me more ads in addition to whatever I could already find elsewhere.
Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?
#289Rather than being told "No, there are only eight pages of results on anything in the goddamned world. Really. Would I lie to you?"
Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?
#290What if instead of even trying to index the entire web, we moved one step back towards the curated directories of the early web? Give users a search engine and indexer that they control and host. Allow them to "follow" domains (or any partial URLs, like subreddits) that they trust.
Make it so that you can configure how many hops it is allowed to take from those trusted sources, similar to LinkedIn's levels of connections. If I'm hosting on my laptop, I might set it at 1 step removed, but if I've got an S3 bucket for my index I might go as far as 3 or 4 steps removed.
There are further optimizations that you could do, such as having your instance not index Wikipedia or Stack Overflow or whatever (instead using the built-in search and aggregating results).
I'm sure there are technical challenges I'm not thinking of, and this would absolutely be a tool that would best serve power users and programmers rather than average internet users. Such an engine wouldn't ever replace Google, but I'd think it would go a long way to making a better search engine for a single user's (or a certain subset of users') everyday web experience.