Live data from Hacker News

Two upstart search engines are teaming up to take on Google

wired.com

81–90 of 299 posts

Re: Two upstart search engines are teaming up to take on Google

#81
post #51
post #4

Every new search engine I've seen was a Bing wrapper with sometimes light reranking. I understand that competing with Google was borderline impossible a decade ago. But in 2024, we have cheap compute, great OSS distributed DBs, powerful new vector search tech. Amateur search engines like Marginalia even run on consumer hardware. CommonCrawl text-only is ~100TB, and can fit on my home server. Why is no company buildin…

>competing with Google was borderline impossible a decade ago. But in 2024, we have cheap compute, great OSS distributed DBs, powerful new vector search tech. [...] CommonCrawl text-only is ~100TB, Those example 3 bullet points of today's improved 2024 computing power you list isn't even enough to process Google's scale 14 years ago in 2010 when the search index was 100+ petabytes : https://googleblog.blogspot.com/20…

You don't have to match Google's technical prowess if that capability is being superceded by MBAs doing aggressive enshittification.

Re: Two upstart search engines are teaming up to take on Google

#82
post #4

Every new search engine I've seen was a Bing wrapper with sometimes light reranking. I understand that competing with Google was borderline impossible a decade ago. But in 2024, we have cheap compute, great OSS distributed DBs, powerful new vector search tech. Amateur search engines like Marginalia even run on consumer hardware. CommonCrawl text-only is ~100TB, and can fit on my home server. Why is no company buildin…

Mostly because Google bought, developed, acquired or effectively control all the major distribution points with default placement deals: eg Apple, Samsung, Chrome, Android, Firefox. In time remedies are coming though, in the antitrust case lost versus DoJ. Another major factor is that building a search index and algorithms that searches across billions of pages with good enough latency is very hard. Easy enough for 1…

Remember that MS lost the anti-trust too, and after the presidential transfer in 2000, they pretty much dropped it with few consequences for MS.

Re: Two upstart search engines are teaming up to take on Google

#83
post #51

Earlier quoted context omitted.

>competing with Google was borderline impossible a decade ago. But in 2024, we have cheap compute, great OSS distributed DBs, powerful new vector search tech. [...] CommonCrawl text-only is ~100TB, Those example 3 bullet points of today's improved 2024 computing power you list isn't even enough to process Google's scale 14 years ago in 2010 when the search index was 100+ petabytes : https://googleblog.blogspot.com/20…

I was talking about text-only, filtered and deduped content. Most of Google's 100PB is picture and video. Filtering the spam and deduping the content helped Google reduce the ~50B page index in 2012 to ~<10B today.

Where are these figures coming from?

Re: Two upstart search engines are teaming up to take on Google

#84
post #48

Earlier quoted context omitted.

A cursory glance at their market share in the search space clearly says that’s not true. For a big site I help run, we’re getting about 8.2x the impressions on Google compared to Bing.

To interpret GP charitably I think they mean that Google is there not because they are a good search engine these days but because of inertia. For you as a site owner, Google is the best: it delivers the impressions. For a user who wants to search, Google has gone downhill since around 2009 and the only thing that confuse me is why DDG - who initially felt better - chose to run after Google down the path of insisting…

> the only thing that confuse me is why DDG - who initially felt better - chose to run after Google down the path of insisting to give me results for things I didn't ask for.

I have the same frustration. The killer feature that got me to switch from Google to DDG was that DDG would reliably return results for the search query I entered, long after Google had stopped doing so. Now that they've taken the same path the benefit is much less. Although I suspect this is more due to a change on the part of Bing than a conscious decision from DDG.

Re: Two upstart search engines are teaming up to take on Google

#85
post #4

Every new search engine I've seen was a Bing wrapper with sometimes light reranking. I understand that competing with Google was borderline impossible a decade ago. But in 2024, we have cheap compute, great OSS distributed DBs, powerful new vector search tech. Amateur search engines like Marginalia even run on consumer hardware. CommonCrawl text-only is ~100TB, and can fit on my home server. Why is no company buildin…

Because nowdays more than ever content you need is in silos. Your facebooks/twiters/instagram/stack overflow/reddit ... And they all have limited expensive api's, and have bulk scrapping detection. Sure you can clobber together something that will work for a while, but you can't runn a buissness on that. Aditionaly most paywalled sites (like news) explicitly whitlist google and bing, and if someone cretes new site, t…

And JavaScript/dynamic content. Entrenched search engines have had a long time to optimize scraping for complex sites

Re: Two upstart search engines are teaming up to take on Google

#86

Why isn’t there a distributed, decentralized or open index that all of these startups can utilize? I understand that these startups are all are focusing in on different problem areas, but doesn’t it make sense to have something like open street maps so that all of these companies can share their compute resources in order to maintain something competitive with the big guys? Or even if it’s not fully decentralized the…

Yacy is still around. While I wouldn't want to disrupt it's decentralized/p2p nature, I think there's a case to be made for a community-managed central aggregation server to help seed the index at various snapshots. I might even be interested in helping run such a thing.

Re: Two upstart search engines are teaming up to take on Google

#87
post #81
post #51

Earlier quoted context omitted.

>competing with Google was borderline impossible a decade ago. But in 2024, we have cheap compute, great OSS distributed DBs, powerful new vector search tech. [...] CommonCrawl text-only is ~100TB, Those example 3 bullet points of today's improved 2024 computing power you list isn't even enough to process Google's scale 14 years ago in 2010 when the search index was 100+ petabytes : https://googleblog.blogspot.com/20…

You don't have to match Google's technical prowess if that capability is being superceded by MBAs doing aggressive enshittification.

You'd be surprised how long it takes to enshittify a piece of tech as well established as Google. The MBAs may be trying but there are still a lot of dedicated folks deep in the org holding out.

Re: Two upstart search engines are teaming up to take on Google

#88
post #4

Every new search engine I've seen was a Bing wrapper with sometimes light reranking. I understand that competing with Google was borderline impossible a decade ago. But in 2024, we have cheap compute, great OSS distributed DBs, powerful new vector search tech. Amateur search engines like Marginalia even run on consumer hardware. CommonCrawl text-only is ~100TB, and can fit on my home server. Why is no company buildin…

I'm guessing there's no money in it unless you glue an ad machine to the side and use search to drive advertising.

Re: Two upstart search engines are teaming up to take on Google

#89
I really hope they can make it work.

We all hate how shitty the big sites have gotten in the name of maximizing profit, but I think it's more insidious than that. It would be one thing if running (say) a news site with decency and integrity was merely less profitable; there would still be people doing it. But I fear that it's actually become impossible to sustain a business like that. The ones that try either die or sell out to survive. (Or are so limited in scope that one person's unpaid part-time labor can sustain them.)

Re: Two upstart search engines are teaming up to take on Google

#90
post #44
post #22

It shouldn't be too hard to achieve what Google were good at before. Their recent search results for me (last 6-12 months) have been so far removed from what I'm searching it felt like a meme. Even after rephrasing things, more details, special quotations etc that everyone knows as the 'search tricks' the results are terrible.

I moved to DDG a couple of years ago, and initially, I found myself often using the `!g` switch, but I honestly can't recall the last time I needed to do that. Only when I'm shopping do I find the goog to be a slightly better tool for finding products sold by niche suppliers. I honestly think google's monopoly on search at this time is 100% powered by momentum, there is almost no other reason to use it over something…

Basically exactly the same thing here. It used to be that google had the better results, but DDG got a little better and google a lot more shit, so here we are I'm pre-filtering what I look for.
Post reply on HN