Live data from Hacker News

Two upstart search engines are teaming up to take on Google

wired.com

161–170 of 299 posts

Re: Two upstart search engines are teaming up to take on Google

#161

Earlier quoted context omitted.

To properly compete with Google search you have to: 1. Create your own index/crawl and true search engine on top of it, rather than delegating that back to bing or Google. In doing so, solve or address all the related problems like SEO/spam, and cover the “long tail” (expensive part) 2. Monetize it somehow to afford all of the personnel/hardware/software costs. Probably with your own version of Google’s ad products.…

I often wonder if banning (or very strongly penalizing) websites with affiliate links would largely get rid of spam results. Sure, some very useful websites have affiliate links, but perhaps it would be worth it to at least let users hide them?

I don't think it's a solution. The average person seems to already have a strong dislike of ads, or paying for something that would otherwise be free (likely funded by ads or aff links)

A solution could be more than one algorithm being used to rank results, i.e. other engines, other rules. They'll likely use many of the signals available to them that Google uses for quality and relevance, but highly unlikely a genuine alternative search would rank them exactly the same- and much more unlikely an SEO could rank well in multiple engines.

The aff links aren't the problem, it's the proliferation of pages that are created solely to rank and get the links clicked on. Sometimes the content is useful, sometimes it's padded nonsense.

Re: Two upstart search engines are teaming up to take on Google

#162

Earlier quoted context omitted.

It's another feature of Kagi. They know they have thousands of results, but they provide you a single page of most relevant results . If you want to see more because you exhausted the page, there's a "more results" button at the bottom. Kagi reduces mental load by default, and this is a good thing.

That's an interesting positive spin on what is clearly a cost-cutting feature. I'm on the unlimited search tier, so it hasn't really been a big deal, but it's worth noting that clicking the "more results" button charges your account for an additional search too. Or at least it used to.

I'm also on the unlimited search tier, and honestly never needed to reach for the more results button even once. Kagi always delivered.

Having a low number of results has a net benefit of lower cognitive load for me, so I like how Kagi returns less results, not more.

But "more results" is counted towards a new search, and I didn't know that. Thanks for pointing out.

Re: Two upstart search engines are teaming up to take on Google

#163

Earlier quoted context omitted.

I was talking about text-only, filtered and deduped content. Most of Google's 100PB is picture and video. Filtering the spam and deduping the content helped Google reduce the ~50B page index in 2012 to ~<10B today.

But what if I don't want to search Reddit, stack overflow, and blogs from the early 2000s and all the content you just threw away as irrelevant actually contains information I am looking for. There is an entire working generation that never heard a modem sound and has never even made a consideration for making sure their content is accessible in plaintext. I'm sure all the LLM providers are already considering this,…

There is still large opportunity. Most of my searches are for plain text information.

> But what if I don't want to search Reddit, stack overflow, and blogs from the early 2000s

That is a strawman. There are huge numbers of websites (including authoritative ones like governments and universities) and a lot of content.

> There is an entire working generation that never heard a modem sound and has never even made a consideration for making sure their content is accessible in plaintext.

If they want video they will do the same as everyone else and search Youtube. Different niche.

> I'm sure all the LLM providers are already considering this, but there's so much important information that is locked away in videos and pictures that isn't even obvious from a transcript or description.

That is true, but if you are getting bad search results (and the market for other search engines are people who are not happy with Google and Bing results) that does not help much are you are not seeing the information you want anyway.

Re: Two upstart search engines are teaming up to take on Google

#164
post #153

Don't believe a single second the marketing speech about these 2 engines, they are both total crap trying to gain users with whatever hype subject is around. Qwant used to pretend being a champion of privacy, a "french made" tech, but with a search engine mostly based on Bing, lobbying with Microsoft against our interests, and with a boss sucking as much public funding as possible to finance luxury HQ locations and l…

I used Qwant a lot in its early phase, but then they decided to become yahoo-esque and also removed the lite version, hence I went back to ddg and brave.

What does yahoo-esque mean in this context?

Re: Two upstart search engines are teaming up to take on Google

#165
Since I don't want to pay for searching, I think that an alternative is to allow search engines to use my computer (one per cent of computer power). The problem is how can I be sure that the use of my computer is not reading my private information?

Another idea is to provide rating for some local products in exchange for searching.

Re: Two upstart search engines are teaming up to take on Google

#166
post #4

Every new search engine I've seen was a Bing wrapper with sometimes light reranking. I understand that competing with Google was borderline impossible a decade ago. But in 2024, we have cheap compute, great OSS distributed DBs, powerful new vector search tech. Amateur search engines like Marginalia even run on consumer hardware. CommonCrawl text-only is ~100TB, and can fit on my home server. Why is no company buildin…

>Why is no company building their own search engine from scratch?

Here's one list:

>A look at search engines with their own indexes

https://seirdy.one/posts/2021/03/10/search-engines-with-own-...

HN discussions:

https://news.ycombinator.com/item?id=26429942

https://news.ycombinator.com/item?id=31820149

https://news.ycombinator.com/item?id=40626011

Re: Two upstart search engines are teaming up to take on Google

#168
post #4

Every new search engine I've seen was a Bing wrapper with sometimes light reranking. I understand that competing with Google was borderline impossible a decade ago. But in 2024, we have cheap compute, great OSS distributed DBs, powerful new vector search tech. Amateur search engines like Marginalia even run on consumer hardware. CommonCrawl text-only is ~100TB, and can fit on my home server. Why is no company buildin…

Kagi is not one of these and I love it enough to pay for it. It has kept the semi-"advanced" features of respecting the negative sign and emphasizing results that match quoted terms in the search query.

Re: Two upstart search engines are teaming up to take on Google

#169
post #14

Earlier quoted context omitted.

I’m my second year into Kagi and loving it. I actually just upgraded to get Kagi Assistant (basically, cloud access to every LLM out there). But the search alone is worth every penny, and it’s built/operated fully in-house as far as I know. https://kagi.com

As much as I like kagi and wish it success, it's not a search engine from scratch. Kagi uses other search engines (Google and Bing) wraps them and does a light reranking

Key differentiator: Kagi still properly responds to the negation sign and quotes in your search terms. This is why I pay for it. The signal-to-noise ratio is way higher than with other engines.

Re: Two upstart search engines are teaming up to take on Google

#170

Since I don't want to pay for searching, I think that an alternative is to allow search engines to use my computer (one per cent of computer power). The problem is how can I be sure that the use of my computer is not reading my private information? Another idea is to provide rating for some local products in exchange for searching.

Since I can never know that, I pay for search (just offering the alternate practical view. I know many will not wish to do so)
Post reply on HN