Live data from Hacker News

Two upstart search engines are teaming up to take on Google

wired.com

1–10 of 299 posts

Re: Two upstart search engines are teaming up to take on Google

#4
Every new search engine I've seen was a Bing wrapper with sometimes light reranking.

I understand that competing with Google was borderline impossible a decade ago. But in 2024, we have cheap compute, great OSS distributed DBs, powerful new vector search tech. Amateur search engines like Marginalia even run on consumer hardware. CommonCrawl text-only is ~100TB, and can fit on my home server.

Why is no company building their own search engine from scratch?

Re: Two upstart search engines are teaming up to take on Google

#6
post #4

Every new search engine I've seen was a Bing wrapper with sometimes light reranking. I understand that competing with Google was borderline impossible a decade ago. But in 2024, we have cheap compute, great OSS distributed DBs, powerful new vector search tech. Amateur search engines like Marginalia even run on consumer hardware. CommonCrawl text-only is ~100TB, and can fit on my home server. Why is no company buildin…

Because it is super expensive and difficult to keep an index up to date. People expect to be able to get current events, and expect search results to be updated in minutes/seconds.

Re: Two upstart search engines are teaming up to take on Google

#7
How things change... I recall having a subscription (had one of those friends who always seemed to know what was coming out right before it came out) right when they started publishing, what was that, 1994? It was so cool.

I just viewed the front page, looks like the definition of "internet chum".

But seriously, to upend Google, you are going to need to be the default on what people use, which I think for now is phones.

Another barrier, maybe they need to get away from this: "will require succeeding at home and growing revenue, which largely comes from running ads." So what do you do when everyone runs some ublock-origin thing? We need to figure out a search monetization beyond "feed me weird things I do not want" on the sidebar. Should websites pay? No, wait, then only those with real money are on the web. Should we pay? Now, wait, we have had it for "free" for too long (could be wrong, I might at this stage in the game pay to have an actual search engine, like say google circa-internet 2004, of course the net was a different place, but still).

Re: Two upstart search engines are teaming up to take on Google

#8
post #3

Google isn't even a competitor in the search space anymore. They've been completely unusable for a decade.

It remains a competitor as long as it continues to capture attention (eye-balls), even if its usability has diminished.

It remains a conmpetitor as long as it continues to be the default search engine in at least two of the most important mainsteam web browsers.

Re: Two upstart search engines are teaming up to take on Google

#9
post #4

Every new search engine I've seen was a Bing wrapper with sometimes light reranking. I understand that competing with Google was borderline impossible a decade ago. But in 2024, we have cheap compute, great OSS distributed DBs, powerful new vector search tech. Amateur search engines like Marginalia even run on consumer hardware. CommonCrawl text-only is ~100TB, and can fit on my home server. Why is no company buildin…

Because it is super expensive and difficult to keep an index up to date. People expect to be able to get current events, and expect search results to be updated in minutes/seconds.

Some sources update faster than others, you could index news sources hourly and low velocity sites weekly. Google does that. CommonCrawl gets 7TB/month, indexing and vectorizing that is quite manageable.

Re: Two upstart search engines are teaming up to take on Google

#10
post #4

Every new search engine I've seen was a Bing wrapper with sometimes light reranking. I understand that competing with Google was borderline impossible a decade ago. But in 2024, we have cheap compute, great OSS distributed DBs, powerful new vector search tech. Amateur search engines like Marginalia even run on consumer hardware. CommonCrawl text-only is ~100TB, and can fit on my home server. Why is no company buildin…

I’m my second year into Kagi and loving it. I actually just upgraded to get Kagi Assistant (basically, cloud access to every LLM out there). But the search alone is worth every penny, and it’s built/operated fully in-house as far as I know.

https://kagi.com

Post reply on HN