Live data from Hacker News

Two upstart search engines are teaming up to take on Google

wired.com

151–160 of 299 posts

Re: Two upstart search engines are teaming up to take on Google

#151
post #51

Earlier quoted context omitted.

>competing with Google was borderline impossible a decade ago. But in 2024, we have cheap compute, great OSS distributed DBs, powerful new vector search tech. [...] CommonCrawl text-only is ~100TB, Those example 3 bullet points of today's improved 2024 computing power you list isn't even enough to process Google's scale 14 years ago in 2010 when the search index was 100+ petabytes : https://googleblog.blogspot.com/20…

Just serving up content from Reddit and HN and a few other websites would be enough to beat Google for most of us. Sprinkle in the top 100 websites and you have a legitimate contender. There is no open web anymore. Google killed it. There are probably fewer than 100k useful websites in the world now. Which is good for startups, because the problem is entirely tractable.

I liken Google Search and YouTube to how Blockbusters video rental stores used to operate.

If you went into Blockbusters then there was actually a small subset of the available videos to rent. Films that had been around for decades were not on the shelves yet garbage released very recently would be there in abundance. If you had an interest in film and, say, wanted to watch everything by Alfred Hitchcock, there would not be a copy of 'The Birds' there for you.

Or another analogy would be a big toy shop. If you grew up in a small town then the toy shop would not stock every LEGO set. You would expect the big toy shop in the big city would have the whole range, but, if you went there, you would just find what the small toy shop had but piled high, the full range still not available.

Record shops were the worst for this. The promise of Virgin Megastore and their like was always a bit of a let down with the local, independently owned record shop somehow having more product than the massive record shop.

Google is a bit like this with information. Youtube is even worse. I cottoned on to this with some testing on other people's devices. Not having Apple products, I wanted to test on old iPads, Macbooks and phones. For this I needed a little bit of help from neighbours and relatives. I already knew I had a bug to workaround, and that there was a tutorial on Youtube I needed to do a quick fix so I could test everything else. So this meant I had to open Youtube on different devices owned by very different people, with their logged in account.

I was very surprised to see that we all had very similar recommendations to what I could expect. I thought the elderly lady downstairs and my sister would have very different recommendations to myself, but they did not. I am sure the adverts would have been different, but I was only there to find a particular tutorial and not be nosy.

I am sure that Google have all this stuff cached at the 'edge', wherever the local copper meets the fibre optic. It is a model a bit like Blockbusters, but where you can get anything on special request, much like how you can order a book from a library for them to get it out of storage for you.

The logical conclusion of this is to have Google text search becoming more like an encyclopedia and dictionary of old, where 90% of what you want can be looked up in a relatively small body of words. I am okay with this, but I still want the special requests. There was merit in old-school Alta Vista searches where you could do what amounts to database queries with logical 'and's 'or's and the like.

The web was written in a very unstructured way, with WYSIWYG being the starting point, with nobody using content sectioning elements to scope headings to words. This mess suits Google as they can gatekeep search, since you need them to navigate a 'sea of divs'.

Really a nation such as France with a language to keep need to make information a public good with content structured and information indexed as a government priority. This immediately screams 'big brother', but it does not have to be like that. Google are not there to serve the customer, they only care about profits. They are not the defenders of democracy and free speech.

If a country such as France or even a country such as Sweden gets their act together and indexes their stuff in their language as a public good, they can export that knowhow to other language groups. It is ludicrous that we are leaving this up to the free market.

Re: Two upstart search engines are teaming up to take on Google

#152

Don't believe a single second the marketing speech about these 2 engines, they are both total crap trying to gain users with whatever hype subject is around. Qwant used to pretend being a champion of privacy, a "french made" tech, but with a search engine mostly based on Bing, lobbying with Microsoft against our interests, and with a boss sucking as much public funding as possible to finance luxury HQ locations and l…

Qwant is just burning French tax payer money in an effort to "compete".

Is it somehow related to the Quaero search engine project that burnt 400 million € of taxpayer money in the early 2000s?

https://en.wikipedia.org/wiki/Quaero

https://www.heise.de/news/400-Millionen-Euro-fuer-europaeisc...

Re: Two upstart search engines are teaming up to take on Google

#153

Don't believe a single second the marketing speech about these 2 engines, they are both total crap trying to gain users with whatever hype subject is around. Qwant used to pretend being a champion of privacy, a "french made" tech, but with a search engine mostly based on Bing, lobbying with Microsoft against our interests, and with a boss sucking as much public funding as possible to finance luxury HQ locations and l…

I used Qwant a lot in its early phase, but then they decided to become yahoo-esque and also removed the lite version, hence I went back to ddg and brave.

Re: Two upstart search engines are teaming up to take on Google

#154

Earlier quoted context omitted.

I’m my second year into Kagi and loving it. I actually just upgraded to get Kagi Assistant (basically, cloud access to every LLM out there). But the search alone is worth every penny, and it’s built/operated fully in-house as far as I know. https://kagi.com

I loved kagi, but their cost is prohibitive for me. The only way there is a chance for me to afford Kagi might be to buy "search credit" without a subscription and without minimum consumption. And then it would only be good if they allowed more than 1000 domain rules and showed more results (when available)

The search credit model sounds great. I would probably also pay Youtube credit. I use it too rarely that paying their monthly rate makes sense. The user experience with ads sucks, so I further try to reduce my usage. At least that's good for the environment and I can do more useful things.

Re: Two upstart search engines are teaming up to take on Google

#155
post #4

Every new search engine I've seen was a Bing wrapper with sometimes light reranking. I understand that competing with Google was borderline impossible a decade ago. But in 2024, we have cheap compute, great OSS distributed DBs, powerful new vector search tech. Amateur search engines like Marginalia even run on consumer hardware. CommonCrawl text-only is ~100TB, and can fit on my home server. Why is no company buildin…

To properly compete with Google search you have to: 1. Create your own index/crawl and true search engine on top of it, rather than delegating that back to bing or Google. In doing so, solve or address all the related problems like SEO/spam, and cover the “long tail” (expensive part) 2. Monetize it somehow to afford all of the personnel/hardware/software costs. Probably with your own version of Google’s ad products.…

I've heard more or less the same thing said about IBM, Walmart, Microsoft, etc.

There is one way to compete with Google search. Google search is a general search engine. But what if you only wanted medical information? information about cars? information about physics? electronics? history? etc.

Specializing makes for a much smaller scope of the problem, and being specialized means it could deliver more useful results.

For example, imdb.com. I don't ask google about movie stuff, I go to imdb.com because it specializes in movie info.

Re: Two upstart search engines are teaming up to take on Google

#156

Earlier quoted context omitted.

Just serving up content from Reddit and HN and a few other websites would be enough to beat Google for most of us. Sprinkle in the top 100 websites and you have a legitimate contender. There is no open web anymore. Google killed it. There are probably fewer than 100k useful websites in the world now. Which is good for startups, because the problem is entirely tractable.

I liken Google Search and YouTube to how Blockbusters video rental stores used to operate. If you went into Blockbusters then there was actually a small subset of the available videos to rent. Films that had been around for decades were not on the shelves yet garbage released very recently would be there in abundance. If you had an interest in film and, say, wanted to watch everything by Alfred Hitchcock, there would…

> It is ludicrous that we are leaving this up to the free market.

If you leave it up to the government, inevitably you're going to get only information approved by the people in power in that government.

You could call that search engine "Pravda".

Re: Two upstart search engines are teaming up to take on Google

#157
post #4

Every new search engine I've seen was a Bing wrapper with sometimes light reranking. I understand that competing with Google was borderline impossible a decade ago. But in 2024, we have cheap compute, great OSS distributed DBs, powerful new vector search tech. Amateur search engines like Marginalia even run on consumer hardware. CommonCrawl text-only is ~100TB, and can fit on my home server. Why is no company buildin…

To properly compete with Google search you have to: 1. Create your own index/crawl and true search engine on top of it, rather than delegating that back to bing or Google. In doing so, solve or address all the related problems like SEO/spam, and cover the “long tail” (expensive part) 2. Monetize it somehow to afford all of the personnel/hardware/software costs. Probably with your own version of Google’s ad products.…

I often wonder if banning (or very strongly penalizing) websites with affiliate links would largely get rid of spam results. Sure, some very useful websites have affiliate links, but perhaps it would be worth it to at least let users hide them?

Re: Two upstart search engines are teaming up to take on Google

#158

Earlier quoted context omitted.

To properly compete with Google search you have to: 1. Create your own index/crawl and true search engine on top of it, rather than delegating that back to bing or Google. In doing so, solve or address all the related problems like SEO/spam, and cover the “long tail” (expensive part) 2. Monetize it somehow to afford all of the personnel/hardware/software costs. Probably with your own version of Google’s ad products.…

I often wonder if banning (or very strongly penalizing) websites with affiliate links would largely get rid of spam results. Sure, some very useful websites have affiliate links, but perhaps it would be worth it to at least let users hide them?

The solution is to force feed all affiliate linked sites into a living archive LLM that digests them into a summary stripped of all links.

You run this bloated mass as a co-engine to your search. If you stumble upon any of the digested sites or articles you just have the site-eater blob regurgitate a html re-creation locally. They get no traffic, you get whatever they wrote purged of all links and with optional formatting to remove the fluff BS copywrite that many of these sites pad the tiny core of usefulness with.

Re: Two upstart search engines are teaming up to take on Google

#159
post #4

Every new search engine I've seen was a Bing wrapper with sometimes light reranking. I understand that competing with Google was borderline impossible a decade ago. But in 2024, we have cheap compute, great OSS distributed DBs, powerful new vector search tech. Amateur search engines like Marginalia even run on consumer hardware. CommonCrawl text-only is ~100TB, and can fit on my home server. Why is no company buildin…

Google ingests almost the whole public web almost every day. I don't see any startup competing with them they might come out with a great algo or something but will need the infrastructure and huge investments to compete.

Even then after using Google Gemini subscription the last few months I think the problem is not Google search rather the web ecosystem as if google gives me most of the answers without having to click any link you and I might be happy but billion people living off those links directly or indirectly won't be happy.

Re: Two upstart search engines are teaming up to take on Google

#160
post #152

Earlier quoted context omitted.

Qwant is just burning French tax payer money in an effort to "compete".

Is it somehow related to the Quaero search engine project that burnt 400 million € of taxpayer money in the early 2000s? https://en.wikipedia.org/wiki/Quaero https://www.heise.de/news/400-Millionen-Euro-fuer-europaeisc...

It is not directly related except by the fact that the dynamic is the same and that it is a repeat of the previous attempt.

There are 2 clear things that is so common in Europe:

Politics injecting a shit load of public money thinking that if you give the money you will be able to reproduce American company success and co.

In the end, the money is wasted for their own interest by big groups, intermediaries, and opportunists. When the thing fails, it is the fault of no one, it was just "too hard" and "maybe the money budget was not enough for such a subject"

This is totally different to what leads innovation like Google, where you have doers that create something first and when they are able to show or convince that they have breakthrough, money will flow in by itself.

And at the beginning cash is used for brain and development instead of giving big salaries to top management and political friends.

The second thing that is usual is the pattern with this kind of projects:

- corporate sucks all the money

- responsibility is shared between multiple actors to spread the blame in case of problem.

- project fail and corporate give up. "Not their fault"

- one year later the initial hype subject is back on the table (European search engine sovereignty for ex) and politics announce that they will spend that much more money to resolve it

- same corporate vampires starts again from zero...

- and it fails the same in a loop

Post reply on HN