Live data from Hacker News

Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

news.ycombinator.com

351–360 of 492 posts

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#352
post #87

Earlier quoted context omitted.

Crypto’s biggest achievement is being the financial equivalent of the gulf war oil fires. Just massive pollution. Think of all the good things that computing could be used for… we used to have all kinds of interesting collaboration projects. Instead we are setting those CPU cycles on fire for short term profit.

Imagine if all that processing power was used for Folding@Home. The problem is that cryptocurrencies do not inherently need tons of processing power to operate. You could theoretically run the entire Bitcoin network on a Raspberry Pi. But the PoW algorithm was designed to always produce a block every 10 minutes, no matter how much hashing power was dedicated to the network. Everyone wanted a piece of the block reward…

> Everyone wanted a piece of the block reward pie, so the arms race was created.

And that's intentional – getting people pursue the goal for their own egoistic reasons, because that's bound to succeed. As a result, they all increase the security and stability of the network whether they want it or not, the only way to not do this is to not participate. If the network were running on a single Raspberry, someone bringing two Raspberries could effectively outcompete the other person on block rewards.

I'm not sure how this can be avoided without fundamental changes in society with regards to competition and adversity.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#353

Earlier quoted context omitted.

I tried out four search words with your search engine, and I am not convinced that it is mainly the index size and not the algorithm that is to blame for bad search results. There are way too much high ranking false positives. Here is what I tried: a) "Berlin": 1. The movie festival "Berlinale" 2. The Wikipedia entry about Berlin 3. Something about a venue "Little Berlin", but the link resolves to an online gaming si…

I think you have some great feedback here but for me it also highlights how subjective search results can be for individuals - for example, these false positives that you mention (b2, b3) appear as the top result on Google for me for that query. It makes me think there must be some fairly large segment of the population that want that domain returned as a result for their query, no?

I would not deny that a large part of subjectivity is involved. This is why I used several markers of subjectivity in my evaluation ("what I can see", "that leaves me", "they seem to me", "I would say", etc.). And related to that: I also agree with other responses that a search often needs to be refined. So my four examples where in no way an exhaustive evaluation, but an explorative experiment, where I just used two proper names, one for a city and one for a historical person, and two general disciplines as search words, in order to see what happens and what is noteworthy (to me). So much to the subjective side.

But what can be said about ideal search results for these terms beyond subjectivity? I do not think that we can arrive at an objective search result, but are nevertheless allowed to criticise search results with respect to their (hidden or obvious) agenda.

Let me give an example of the good old days: When I was searching for my surename on Google in the early 2000s the search results contained a lot of university papers or personal Web-sites (then called "homepages") from other people of that name. But suddenly, I can't remember when exactly this was, the search results contained almost exlusively companies that contained that surename in its company name. The shift was not gradually, as if it were representing a slow shift in the contents of the Internet itself, but abrupt. It was apparently due to an intentional modification of the ranking algorithm that put business far above anything else on the Internet.

My explanation for this is the following: The objective metrics for Google search results is the stream of revenue they generate for Google. But not only for Google. The fundamental monetary incentive to Bing (and its derivative Ecosia) is more or less the same. And how different the impact of the somewhat different business model of Duck Duck Go is, is open for debate.

If maximum revenue is the goal, the aim is to provide the best search results according to the business model (advertisment, market research and whatever else) without driving the users away. But the best search results according to the business model are not necessarily the optimal search results for the typical user. And as long as all relevant competitors are following the same economic pressure of maximizing revenue the basic situation an thus the qualitiy of the search results for the user will not improve above a certain level. If we want this situation to change, we need competitors with a different, non-commerical agenda. Either from the public sector (an analogon to the excellent information services about physical books provided by libraries) or from non-profit organizations (an analogon to Wikipedia or Open Street Map).

To answer your question about b2 and b3: I checked with other search engines; besides Google they appear for me also on Bing (as #8, same product but on a different Web-site) and Duck Duck Go (as #10); Bing also has a reference to them in the right margin as a suggestion for a refined search (this time exactly b2 and b3). Although I do not think that the results from those search engines should be considered as a general benchmark for good search results for the reasons given above, we may speculate why they appear on the first page of search results. I would guess that it is a combination of gaming the search engines by using a generic term as a product and domain name to get free advertising, and search engine algorithms making this possible by generally ranking products and companies high in their search results.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#355
post #12

Ha, yes, I've done that at https://gigablast.com/ . The biggest problems now are the following: 1) Too hard to spider the web. Gatekeeper companies like Cloudflare (owned in part by Google) and Cloudfront make it really difficult for upstart search engines to download web pages. 2) Hardware costs are too high. It's much more expensive now to build a large index (50B+ pages) to be competitive. I believe my algorithms…

Dude, I use your engine regularly, it is spectacular. The amount of work you put into this takes some dedication.

I was curious if you ever intend to implement OpenSearch API so that we could use it as default in browser or embed it in applications?

Also how can people contribute to help you maintain a larger index and/or keep the service going?

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#356

Earlier quoted context omitted.

$9/mo is not going to cut it. Google's domestic annual revenue per user in 2019 was $256. [0] That's $21.33 per month. Not all of Google's revenue is from Ads, of course, but the vast majority is. (Let's ignore for now the valid counterpoint that Ads are increasingly served on other Google properties than Search.) But even charging users $21.33/mo for an ad-free search experience most likely wouldn't be enough. By pr…

Let’s say ads will always make more money (I have no reason to believe they won’t), and that’s required to be the dominant search engine because the web is big and expensive to organize. I’d bet there’s some way to characterize what I and others liked about the earlier web and create a search engine that just worries about that stuff. I’d pay $9/mo for whatever 1/3 of Google’s spend per user would get me. That’s not…

I doubt it, because 1/3 of Google's spend per user isn't enough when you can't attract many paying users in the first place, because you would charge much more than $9/mo, because almost no one wants to pay for a search engine so your revenue will have to make up for those people too, and then even fewer people are willing to pay more than $9/mo for 1/3 of the quality.

And then I'd guess the 20 remaining users will still complain because 1999 Google is a nostalgic memory impossible to recreate without a 1999 internet for a 1999 self to live in and has little to do with raw search quality.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#357

Earlier quoted context omitted.

> 2) Hardware costs are too high. Which is why the next big search engine should be distributed: https://yacy.net .

"distributed" doesn't make things more hardware efficient... It literally always make them less efficient. If e.g : mastodon had the same number of users as Twitter it would use 10x the ressources for the same traffic.

Sure, but it does spread the costs among users and makes them more manageable. One guy shouldering the cost of a search index is less viable than letting users shoulder the costs. Some charge customers as a solution to this, and that works, but then they need a minimum revenue to continue, or have to monetize with investors which usually means changing direction and goals. The other option, letting people host portions of the index, spreads the cost out, and the product gets about as good (best case scenario) as it's utility to people.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#358
> I seem to recall that Google consistently produced relevant results and strictly respected search operators in 2005 (?), unlike the modern Google.

You recall wrong.

Probably because you were a child and searching for "reddit".

Now as an adult, Google can't just hand you adult results by magic.

Search operators have changed, but that's because the internet is 1000's of times bigger since 2005. Where as the number of people went from ~1 Billion to ~5 Billion.

> I think search results were the same for everyone, rather than being customized for each user.

You are not a baby, turn off the customisation. The same issue existed ~2005, Google customised and we had to work to turn it off. Also my idea we were become one world was totally wrong. Google customising for my location was more correct than my idealism. It also helped local businesses get online. That Google is 'evil' by default is a shitty assumption.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#359
post #12

Ha, yes, I've done that at https://gigablast.com/ . The biggest problems now are the following: 1) Too hard to spider the web. Gatekeeper companies like Cloudflare (owned in part by Google) and Cloudfront make it really difficult for upstart search engines to download web pages. 2) Hardware costs are too high. It's much more expensive now to build a large index (50B+ pages) to be competitive. I believe my algorithms…

I wanted to add my site to Gigablast, but it said it would cost 25 cents. How is this a good thing?

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#360
post #12

Ha, yes, I've done that at https://gigablast.com/ . The biggest problems now are the following: 1) Too hard to spider the web. Gatekeeper companies like Cloudflare (owned in part by Google) and Cloudfront make it really difficult for upstart search engines to download web pages. 2) Hardware costs are too high. It's much more expensive now to build a large index (50B+ pages) to be competitive. I believe my algorithms…

I really like GigaBlast. I wrote a "meta" search utility for myself that can query multiple search engines from the command line.^1 It mixes the results into a simplified SERP ("metaSERP"), optimised for a text-only browser, with indicators to show which search engine each result came from. The key feature is that it allows for what I might call "continuation searches". Each metaSERPs contains timestamps in its page…

Wow, do you happen to have published your utility so that other people can play with it?
Post reply on HN