Live data from Hacker News

Two upstart search engines are teaming up to take on Google

wired.com

261–270 of 299 posts

Re: Two upstart search engines are teaming up to take on Google

#261

Earlier quoted context omitted.

But what if I don't want to search Reddit, stack overflow, and blogs from the early 2000s and all the content you just threw away as irrelevant actually contains information I am looking for. There is an entire working generation that never heard a modem sound and has never even made a consideration for making sure their content is accessible in plaintext. I'm sure all the LLM providers are already considering this,…

There is still large opportunity. Most of my searches are for plain text information. > But what if I don't want to search Reddit, stack overflow, and blogs from the early 2000s That is a strawman. There are huge numbers of websites (including authoritative ones like governments and universities) and a lot of content. > There is an entire working generation that never heard a modem sound and has never even made a con…

> That is a strawman. There are huge numbers of websites (including authoritative ones like governments and universities) and a lot of content.

Ya know... a search engine that was limited to *.gov, *.edu and country equivalents (*.ac.uk, etc) would actually be pretty useful. Ok, I know you can do something like it with site: modifiers in the search, but if you know from the beginning you're never going to search the commercial internet you can bake that assumption into the design of your search engine in interesting ways.

And the spam problem goes away.

Hmm.

Re: Two upstart search engines are teaming up to take on Google

#262
post #51

Earlier quoted context omitted.

>competing with Google was borderline impossible a decade ago. But in 2024, we have cheap compute, great OSS distributed DBs, powerful new vector search tech. [...] CommonCrawl text-only is ~100TB, Those example 3 bullet points of today's improved 2024 computing power you list isn't even enough to process Google's scale 14 years ago in 2010 when the search index was 100+ petabytes : https://googleblog.blogspot.com/20…

Just serving up content from Reddit and HN and a few other websites would be enough to beat Google for most of us. Sprinkle in the top 100 websites and you have a legitimate contender. There is no open web anymore. Google killed it. There are probably fewer than 100k useful websites in the world now. Which is good for startups, because the problem is entirely tractable.

I dunno. Everybody's got their own personal long tail, and everybody's got a different long tail.

(I don't think the open web is dead, but it's looking awfully unwell).

Re: Two upstart search engines are teaming up to take on Google

#263

I was hoping this will be about Kagi. A lot of companies that make bold promises of taking on Google end up in the drain 2 years later.

I'm a paying customer of Kagi. Kagi itself runs on top of Google. It's certainly worth paying for (IMO), but let's not kid ourselves -- it is *not* a competitor (as in, can replace Google). Without Google, there is no Kagi.

I used Kagi for half a year - I loved it!

I was hoping Kagi got some media coverage and this was an article covering how it's so much better than Google.

My problem is with companies who start by declaring they are going to take on Google. These kind of companies are nowhere to be found a few years later.

Re: Two upstart search engines are teaming up to take on Google

#264
post #4

Every new search engine I've seen was a Bing wrapper with sometimes light reranking. I understand that competing with Google was borderline impossible a decade ago. But in 2024, we have cheap compute, great OSS distributed DBs, powerful new vector search tech. Amateur search engines like Marginalia even run on consumer hardware. CommonCrawl text-only is ~100TB, and can fit on my home server. Why is no company buildin…

It occurs to me that competing with Google on their home turf and at their scale might be impossible. But doing what Google was good at, when they just a search engine, might not be all that much harder today than it was back then. And it may fly "under the radar" of Google's current business priorities.

I'm thinking about how Etsy took "the good bit" of Ebay and ran with it. Or how tyre-and-exhaust shops take "the profitable bit" of being a mechanic and run with it. I know there's a name for this, but it eludes me.

Re: Two upstart search engines are teaming up to take on Google

#265

Earlier quoted context omitted.

You should use the right tool for the right purpose. If you're searching for a physical place and want opening hours, reviews and photos, it's Google Maps. If you're looking for updates and photos for a physical place, it's Instagram. If it's detailed articles it's Wikipedia. If it's to define or translate a word, right click and "Look up". Or if it's general search, you have Kagi.

Wikipedia? Are you kidding?

It's a good starting point.

Re: Two upstart search engines are teaming up to take on Google

#266
post #127

Earlier quoted context omitted.

There is probably room for one or five lifestyle businesses but convincing venture capital to drop the megabux to go big would be a feat and eventually land at some sub-optimal state anyway. Finding some hack to democratize&decentralize the indexing and expensive processes like JavaScript interpretation, image interpretation, OCR, etc is an open angle and even an avenue for "Web3" to offload the cost. But you will ul…

I would want to use a search engine that does not perform JavaScript interpretation, image interpretation, OCR, etc. (This is not the same as excluding web pages with JavaScripts from the search results. They would still be included but only indexed by whatever text is available without JavaScripts; if there isn't any such text, then they should be excluded. This would also apply if it is only pictures, video, etc an…

Luckily none of these things are mutually exclusive.

Thanks to all the SPA idiocy you will miss enough content to matter if you have zero JS interpretation, so you would want to let the user choose which indexes they want for a query because sometimes you need these other resource types to answer the query.

Re: Two upstart search engines are teaming up to take on Google

#267

Earlier quoted context omitted.

Not to mention that the cost per search in terms of compute and energy is so much smaller for web search than for running an LLM. I forget the exact numbers now, but it was orders of magnitude as I recall. Search engines are just cheaper to run. I don't know that there's a good, long term model for a free LLM-based search replacement because of how much higher the operating costs are, ad supported or not.

On top of that, search usually uses CPU instead of GPU. A large infrastructure with CPU’s is easier to reuse for jobs other than search.

These are great reasons why this business will be hard, but given how ChatGPT and Perplexity are making inroads into search traffic, you can't deny it's an experience consumers prefer.

Re: Two upstart search engines are teaming up to take on Google

#268
post #102

My hot take on this new era of search engines is that "search is a bug" and even trying to be a search engine is a fool's errand. Search solved a problem of the legacy internet where you wanted information and that information would be on one of a million websites. If someone is going to disrupt Google, it's because they've cut out the middleman that is search results and simply give you what you're asking for . Chat…

Search is still better for getting to specific, existing documents you need. Even the RAG people have been finding that out with hybrid models becoming more popular over time. I also think you can update search indexes more cheaply than further pretraining LLM’s.

This is a great point, but I wonder how much of that kind of search intent is part of google's traffic. If that becomes the only reason people use Google I wonder if they'll go the way of Yahoo. Maybe that's hyperbolic, but there was a time when Yahoo's dominance seemed unquestionable (I'm old).

To be clear, I'm not arguing for everything should be part of a pretrained LLM, but the experience of knowledge searching that ChatGPT and Perplexity provide are pretty superior to Google today (when they work).

Re: Two upstart search engines are teaming up to take on Google

#269
post #123

Earlier quoted context omitted.

Remember that MS lost the anti-trust too, and after the presidential transfer in 2000, they pretty much dropped it with few consequences for MS.

If I had to guess I'd say history is going to repeat itself and Google will also escape any significant consequences.

Maybe, maybe not. Remember that Google is "woke" - an enemy of the faction that now finds itself in power. They might continue the lawsuit to set an example.

Re: Two upstart search engines are teaming up to take on Google

#270
post #182

Earlier quoted context omitted.

It says right on the subtitle (emphasis mine): > Generative AI and new rules targeting tech giants are giving Ecosia and Qwant fuel to challenge Google and Microsoft and develop a web index for Europe. “New rules” means regulation. “Market forces” kept the incumbents at the top, it is regulation that is giving others a chance.

What are the rules?

You can literally do a web search for “new rules for big tech”.
Post reply on HN