Live data from Hacker News

Two upstart search engines are teaming up to take on Google

wired.com

211–220 of 299 posts

Re: Two upstart search engines are teaming up to take on Google

#211

Earlier quoted context omitted.

You should use the right tool for the right purpose. If you're searching for a physical place and want opening hours, reviews and photos, it's Google Maps. If you're looking for updates and photos for a physical place, it's Instagram. If it's detailed articles it's Wikipedia. If it's to define or translate a word, right click and "Look up". Or if it's general search, you have Kagi.

Wikipedia? Are you kidding?

wiki is adequate most of the time for proper nouns.

Re: Two upstart search engines are teaming up to take on Google

#212
post #4

Every new search engine I've seen was a Bing wrapper with sometimes light reranking. I understand that competing with Google was borderline impossible a decade ago. But in 2024, we have cheap compute, great OSS distributed DBs, powerful new vector search tech. Amateur search engines like Marginalia even run on consumer hardware. CommonCrawl text-only is ~100TB, and can fit on my home server. Why is no company buildin…

A single CommonCrawl dump might be 100TB but that represents less than 1% of the Internet. CommonCrawl crawls news parts of the Internet every month / trimester and there is little overlap between crawl dumps.

Re: Two upstart search engines are teaming up to take on Google

#213

I'm beginning to suspect LLMs' viability for search purposes is dependent on your existing search habits. 40% of my search queries are just copy-pasted error messages. Another 10% are business names, for the sole purpose of finding their hours or phone number. Less than 10% are complete clauses or sentences. I just don't see how an LLM fits my search habits. I tried using ChatGPT for search purposes, and it was dread…

LLMs open a whole new 'search' up for me, when you are not sure what you are searching for. For example, if you can't remember the name of a lib but you know roughly what it does (or if you have misremembered the name of it). This is often very time consuming.

It's also far better than Google for recommending things (really anything) as you can use much more "precise" requirements for the recommendations and correct them, instead of getting at best some SEO spam.

It's obviously not good for stuff that requires real time updates like opening times, though I suspect with time OpenAI et al will combine their own search index + RAG to solve a lot of this.

The elephant in the room though is this is long term going to kill the incentive for a lot of content to be published. AI overviews have definitely reduced the amount of clicking through I do, even though I try and subconsciously 'support' sites by trying to find a link. But often it's right there.

Re: Two upstart search engines are teaming up to take on Google

#214

I'm beginning to suspect LLMs' viability for search purposes is dependent on your existing search habits. 40% of my search queries are just copy-pasted error messages. Another 10% are business names, for the sole purpose of finding their hours or phone number. Less than 10% are complete clauses or sentences. I just don't see how an LLM fits my search habits. I tried using ChatGPT for search purposes, and it was dread…

There are some tasks that I would have used a search engine for in the past, because that was more or less the only option. Say, searching for "ffmpeg transcode syntax" and then spending 15 minutes comparing examples from Stack Overflow with the official documentation to try to make sense of them. Now I can tell Claude exactly what I'm trying to accomplish and it will give me an answer in 30 seconds that's either correct, or close enough for me to quickly make up the difference.

I'm still going to turn to Google to find out what a store's opening hours or phone number is, as well as a lot of other tasks. But there are types of queries that are better suited for an LLM, that previously could only be done in a search engine.

There's also a non-technical reason for LLM search. Google built its business on free search, paid for by advertising, which seemed like a good idea in the early 2000's. A few decades later and we have a better appreciation for the value of the ad-driven business model. Right now, there's a whole lot of money being thrown at online LLMs, so for the most part they're not really doing ads yet. It's refreshing to make a query and not have sponsored results at the top of the list. Obviously, the free online LLM business model isn't going to last indefinitely. In the pretty near future, we'll either need to start paying a usage fee, or parse through advertisements delivered by LLMs as well. But it's nice while it lasts.

Re: Two upstart search engines are teaming up to take on Google

#215
post #189

Earlier quoted context omitted.

It is not directly related except by the fact that the dynamic is the same and that it is a repeat of the previous attempt. There are 2 clear things that is so common in Europe: Politics injecting a shit load of public money thinking that if you give the money you will be able to reproduce American company success and co. In the end, the money is wasted for their own interest by big groups, intermediaries, and opport…

Tbf, it's not like this approach can't work at all. Airbus was born out of a similar dynamic, and it's giving Boeing a run for its money now. Afaik France has a couple of other giants in technology-heavy industries such as shipping or mining, but I couldn't speak to their success. What seems clear by now, though, is that the approach isn't suitable for "tech" (in the typical SV sense of the word), especially consumer…

>Tbf, it's not like this approach can't work at all. Airbus was born out of a similar dynamic, and it's giving Boeing a run for its money now.

It literally can't work at all. When was the last time you went and bought an Airbus? Airbus doesn't make consumer products. Passengers are the consumers flying inside them but they're not the ones buying them, it's the airlines who only have a monopoly of 2 global players to choose from in a highly regulated industry with expensive moats to enter meaning Airbus and Boeing don't really need to compete cut-throat.

Governments excel at building large infrastructure and defense companies like Airbus, Boeing, what have you, not at building consumer products at scale like Google, Apple, etc sine their success is dictated by the consumer spending preferences, not by requirements a government makes up.

Communist regimes did not make the best consumer products, the free market did.

Re: Two upstart search engines are teaming up to take on Google

#216

Earlier quoted context omitted.

To properly compete with Google search you have to: 1. Create your own index/crawl and true search engine on top of it, rather than delegating that back to bing or Google. In doing so, solve or address all the related problems like SEO/spam, and cover the “long tail” (expensive part) 2. Monetize it somehow to afford all of the personnel/hardware/software costs. Probably with your own version of Google’s ad products.…

> You will probably at the very minimum need to implement your own platform (device, OS, browser alone probably won’t cut it). What do you mean? Google gained dominance just being a dot com URL. This in a time when competitors were already very well established.

Google got big when being a dot-com URL meant something and the primary way the average person accessed the Internet was through a desktop PC. Neither of those things is true anymore.

Re: Two upstart search engines are teaming up to take on Google

#217

Earlier quoted context omitted.

You should use the right tool for the right purpose. If you're searching for a physical place and want opening hours, reviews and photos, it's Google Maps. If you're looking for updates and photos for a physical place, it's Instagram. If it's detailed articles it's Wikipedia. If it's to define or translate a word, right click and "Look up". Or if it's general search, you have Kagi.

Wikipedia? Are you kidding?

What would you suggest instead?

Re: Two upstart search engines are teaming up to take on Google

#218
post #205

Earlier quoted context omitted.

Compared to what?

Compared to Google, X years ago, for example. Unless I'm mixing threads up, that's what we're talking about anyway: the degradation of Google's search results.

Yeah, the problem is that Google today is still generally better than its competitors today, even though Google today is worse than Google yesterday.

Re: Two upstart search engines are teaming up to take on Google

#219
post #4

Every new search engine I've seen was a Bing wrapper with sometimes light reranking. I understand that competing with Google was borderline impossible a decade ago. But in 2024, we have cheap compute, great OSS distributed DBs, powerful new vector search tech. Amateur search engines like Marginalia even run on consumer hardware. CommonCrawl text-only is ~100TB, and can fit on my home server. Why is no company buildin…

Even ChatGPT Search was revealed to be a Bing wrapper, this meme summed it up - https://ibb.co/8csc3gv

Re: Two upstart search engines are teaming up to take on Google

#220

Earlier quoted context omitted.

To properly compete with Google search you have to: 1. Create your own index/crawl and true search engine on top of it, rather than delegating that back to bing or Google. In doing so, solve or address all the related problems like SEO/spam, and cover the “long tail” (expensive part) 2. Monetize it somehow to afford all of the personnel/hardware/software costs. Probably with your own version of Google’s ad products.…

I'd like to add to your #5. If Google deems you a legitimate threat, then they can just de-crapify their own search for a bit by going back to their old algo. It's extremely easy for them to fight back.

That would be a victory for users.
Post reply on HN