Live data from Hacker News

Two upstart search engines are teaming up to take on Google

wired.com

61–70 of 299 posts

Re: Two upstart search engines are teaming up to take on Google

#62
post #22

It shouldn't be too hard to achieve what Google were good at before. Their recent search results for me (last 6-12 months) have been so far removed from what I'm searching it felt like a meme. Even after rephrasing things, more details, special quotations etc that everyone knows as the 'search tricks' the results are terrible.

I'm using Kagi for quite some time. It's invisible to me. I search, get 30ish high quality results per search, and I'm a happy camper. No ads, no seo grafting, nothing. Moreover, I can block sites and customize my own search results. This feels good. When I first started using Kagi, it felt like leaving a closed building and stepping out to open air.

I’ve heard good things about it actually I must go check it out now. Thanks for reminding me. I don’t mind paying for things that save me time.

Re: Two upstart search engines are teaming up to take on Google

#63
post #51
post #4

Every new search engine I've seen was a Bing wrapper with sometimes light reranking. I understand that competing with Google was borderline impossible a decade ago. But in 2024, we have cheap compute, great OSS distributed DBs, powerful new vector search tech. Amateur search engines like Marginalia even run on consumer hardware. CommonCrawl text-only is ~100TB, and can fit on my home server. Why is no company buildin…

>competing with Google was borderline impossible a decade ago. But in 2024, we have cheap compute, great OSS distributed DBs, powerful new vector search tech. [...] CommonCrawl text-only is ~100TB, Those example 3 bullet points of today's improved 2024 computing power you list isn't even enough to process Google's scale 14 years ago in 2010 when the search index was 100+ petabytes : https://googleblog.blogspot.com/20…

I was talking about text-only, filtered and deduped content.

Most of Google's 100PB is picture and video. Filtering the spam and deduping the content helped Google reduce the ~50B page index in 2012 to ~<10B today.

Re: Two upstart search engines are teaming up to take on Google

#64
post #29

every new search engine i have tried so far, failed to do regional searches properly. for some queries i need regional index without adding my country to the query.

Bing and Google do have the benefit of websites defining their market via their webmaster tools sections. And a critical mass of click through metrics, etc.

Re: Two upstart search engines are teaming up to take on Google

#65
post #9

Earlier quoted context omitted.

Because it is super expensive and difficult to keep an index up to date. People expect to be able to get current events, and expect search results to be updated in minutes/seconds.

Some sources update faster than others, you could index news sources hourly and low velocity sites weekly. Google does that. CommonCrawl gets 7TB/month, indexing and vectorizing that is quite manageable.

For news-only there's https://littleberg.com

Re: Two upstart search engines are teaming up to take on Google

#67
Shoutout for Mojeek. Highlights:

- Has its own index

- Had a pro-privacy privacy policy well before DDG existed

- No conflated marketing (DDG wrt it passing on 3 octets of an IP to Bing for local searches, Ecosia for not including Bing's carbon footprint as a meta search engine)

- Has an API

It's a smaller index and has more limited resources, but pretty much the best genuine alternative, and great for finding older resources that have long since been buried by G/B.

Re: Two upstart search engines are teaming up to take on Google

#68
post #51

Earlier quoted context omitted.

>competing with Google was borderline impossible a decade ago. But in 2024, we have cheap compute, great OSS distributed DBs, powerful new vector search tech. [...] CommonCrawl text-only is ~100TB, Those example 3 bullet points of today's improved 2024 computing power you list isn't even enough to process Google's scale 14 years ago in 2010 when the search index was 100+ petabytes : https://googleblog.blogspot.com/20…

I was talking about text-only, filtered and deduped content. Most of Google's 100PB is picture and video. Filtering the spam and deduping the content helped Google reduce the ~50B page index in 2012 to ~<10B today.

[deleted]

Re: Two upstart search engines are teaming up to take on Google

#69
post #4

Every new search engine I've seen was a Bing wrapper with sometimes light reranking. I understand that competing with Google was borderline impossible a decade ago. But in 2024, we have cheap compute, great OSS distributed DBs, powerful new vector search tech. Amateur search engines like Marginalia even run on consumer hardware. CommonCrawl text-only is ~100TB, and can fit on my home server. Why is no company buildin…

"CommonCrawl [being] text-only is only ~100TB, and can fit on my home server."

Are any individual users downloading CC for potential use in the future?

It may seem like a non-trivial task to process ~100TB at home today but in the future processing this amount of data will likely seem trivial. CC data is available for download to anyone but, to me, it appears only so-called "tech" companies and "researchers" are grabbing a copy.

Many years ago I began storing copies of the publicly-available com. and net. zonefiles from Verisign. At the time it was infeasible for me to try to serve multi-GB zonefiles on the local network at home. Today, it's feasible. And in the future it will be even easier.

NB. I am not employed by a so-called "tech" company. I store this data for personal, non-commercial use.

Re: Two upstart search engines are teaming up to take on Google

#70

Earlier quoted context omitted.

Because nowdays more than ever content you need is in silos. Your facebooks/twiters/instagram/stack overflow/reddit ... And they all have limited expensive api's, and have bulk scrapping detection. Sure you can clobber together something that will work for a while, but you can't runn a buissness on that. Aditionaly most paywalled sites (like news) explicitly whitlist google and bing, and if someone cretes new site, t…

This is the best (and saddest) answer. LLMs break the social contract of the internet, we're in a feudalisation process. The decentralized nature of the internet was amazing for businesses, and monopolization could ruin the space and slow innovation down significantly.

While LLMs have accelerated, it, it was already the case that silos were blocking non-Google and non-Bing results before LLMs. LLMs have only made existing problems of the web worse, but they were problems before LLMs too and banning LLMs won't fix the core issues of silos and misinformation.
Post reply on HN