Every new search engine I've seen was a Bing wrapper with sometimes light reranking. I understand that competing with Google was borderline impossible a decade ago. But in 2024, we have cheap compute, great OSS distributed DBs, powerful new vector search tech. Amateur search engines like Marginalia even run on consumer hardware. CommonCrawl text-only is ~100TB, and can fit on my home server. Why is no company buildin…
>competing with Google was borderline impossible a decade ago. But in 2024, we have cheap compute, great OSS distributed DBs, powerful new vector search tech. [...] CommonCrawl text-only is ~100TB, Those example 3 bullet points of today's improved 2024 computing power you list isn't even enough to process Google's scale 14 years ago in 2010 when the search index was 100+ petabytes : https://googleblog.blogspot.com/20…
Two upstart search engines are teaming up to take on Google
81–90 of 299 posts
Re: Two upstart search engines are teaming up to take on Google
#82Every new search engine I've seen was a Bing wrapper with sometimes light reranking. I understand that competing with Google was borderline impossible a decade ago. But in 2024, we have cheap compute, great OSS distributed DBs, powerful new vector search tech. Amateur search engines like Marginalia even run on consumer hardware. CommonCrawl text-only is ~100TB, and can fit on my home server. Why is no company buildin…
Mostly because Google bought, developed, acquired or effectively control all the major distribution points with default placement deals: eg Apple, Samsung, Chrome, Android, Firefox. In time remedies are coming though, in the antitrust case lost versus DoJ. Another major factor is that building a search index and algorithms that searches across billions of pages with good enough latency is very hard. Easy enough for 1…
Re: Two upstart search engines are teaming up to take on Google
#83Earlier quoted context omitted.
>competing with Google was borderline impossible a decade ago. But in 2024, we have cheap compute, great OSS distributed DBs, powerful new vector search tech. [...] CommonCrawl text-only is ~100TB, Those example 3 bullet points of today's improved 2024 computing power you list isn't even enough to process Google's scale 14 years ago in 2010 when the search index was 100+ petabytes : https://googleblog.blogspot.com/20…
I was talking about text-only, filtered and deduped content. Most of Google's 100PB is picture and video. Filtering the spam and deduping the content helped Google reduce the ~50B page index in 2012 to ~<10B today.
Re: Two upstart search engines are teaming up to take on Google
#84Earlier quoted context omitted.
A cursory glance at their market share in the search space clearly says that’s not true. For a big site I help run, we’re getting about 8.2x the impressions on Google compared to Bing.
To interpret GP charitably I think they mean that Google is there not because they are a good search engine these days but because of inertia. For you as a site owner, Google is the best: it delivers the impressions. For a user who wants to search, Google has gone downhill since around 2009 and the only thing that confuse me is why DDG - who initially felt better - chose to run after Google down the path of insisting…
I have the same frustration. The killer feature that got me to switch from Google to DDG was that DDG would reliably return results for the search query I entered, long after Google had stopped doing so. Now that they've taken the same path the benefit is much less. Although I suspect this is more due to a change on the part of Bing than a conscious decision from DDG.
Re: Two upstart search engines are teaming up to take on Google
#85Every new search engine I've seen was a Bing wrapper with sometimes light reranking. I understand that competing with Google was borderline impossible a decade ago. But in 2024, we have cheap compute, great OSS distributed DBs, powerful new vector search tech. Amateur search engines like Marginalia even run on consumer hardware. CommonCrawl text-only is ~100TB, and can fit on my home server. Why is no company buildin…
Because nowdays more than ever content you need is in silos. Your facebooks/twiters/instagram/stack overflow/reddit ... And they all have limited expensive api's, and have bulk scrapping detection. Sure you can clobber together something that will work for a while, but you can't runn a buissness on that. Aditionaly most paywalled sites (like news) explicitly whitlist google and bing, and if someone cretes new site, t…
Re: Two upstart search engines are teaming up to take on Google
#86Why isn’t there a distributed, decentralized or open index that all of these startups can utilize? I understand that these startups are all are focusing in on different problem areas, but doesn’t it make sense to have something like open street maps so that all of these companies can share their compute resources in order to maintain something competitive with the big guys? Or even if it’s not fully decentralized the…
Re: Two upstart search engines are teaming up to take on Google
#87Earlier quoted context omitted.
>competing with Google was borderline impossible a decade ago. But in 2024, we have cheap compute, great OSS distributed DBs, powerful new vector search tech. [...] CommonCrawl text-only is ~100TB, Those example 3 bullet points of today's improved 2024 computing power you list isn't even enough to process Google's scale 14 years ago in 2010 when the search index was 100+ petabytes : https://googleblog.blogspot.com/20…
You don't have to match Google's technical prowess if that capability is being superceded by MBAs doing aggressive enshittification.
Re: Two upstart search engines are teaming up to take on Google
#88Every new search engine I've seen was a Bing wrapper with sometimes light reranking. I understand that competing with Google was borderline impossible a decade ago. But in 2024, we have cheap compute, great OSS distributed DBs, powerful new vector search tech. Amateur search engines like Marginalia even run on consumer hardware. CommonCrawl text-only is ~100TB, and can fit on my home server. Why is no company buildin…
Re: Two upstart search engines are teaming up to take on Google
#89We all hate how shitty the big sites have gotten in the name of maximizing profit, but I think it's more insidious than that. It would be one thing if running (say) a news site with decency and integrity was merely less profitable; there would still be people doing it. But I fear that it's actually become impossible to sustain a business like that. The ones that try either die or sell out to survive. (Or are so limited in scope that one person's unpaid part-time labor can sustain them.)
Re: Two upstart search engines are teaming up to take on Google
#90It shouldn't be too hard to achieve what Google were good at before. Their recent search results for me (last 6-12 months) have been so far removed from what I'm searching it felt like a meme. Even after rephrasing things, more details, special quotations etc that everyone knows as the 'search tricks' the results are terrible.
I moved to DDG a couple of years ago, and initially, I found myself often using the `!g` switch, but I honestly can't recall the last time I needed to do that. Only when I'm shopping do I find the goog to be a slightly better tool for finding products sold by niche suppliers. I honestly think google's monopoly on search at this time is 100% powered by momentum, there is almost no other reason to use it over something…