Live data from Hacker News

Two upstart search engines are teaming up to take on Google

wired.com

271–280 of 299 posts

Re: Two upstart search engines are teaming up to take on Google

#271

The best way to stop SEO is to build a bot which detects commercial activity on the page (and/or availability of cart/payment controls), and commercial language in the text of the page (does the content look like an ad or product/sales page). Also combine with number of hops to a commercial page (with cart, prices, advertising material, etc), and create a distance metric. Offer the user a commercial activity slider.…

> The best way to stop SEO is to build a bot which detects commercial activity on the page (and/or availability of cart/payment controls), and commercial language in the text of the page (does the content look like an ad or product/sales page). Thanks for describing the detection algorithm, I will now design my spam sites to defeat it. This is how SEO works. There are no silver bullets.

I guess you could make a spam site that didn’t have affiliate links, advertisements, or marketing text… but why bother with the spam?

Re: Two upstart search engines are teaming up to take on Google

#272
post #4

Every new search engine I've seen was a Bing wrapper with sometimes light reranking. I understand that competing with Google was borderline impossible a decade ago. But in 2024, we have cheap compute, great OSS distributed DBs, powerful new vector search tech. Amateur search engines like Marginalia even run on consumer hardware. CommonCrawl text-only is ~100TB, and can fit on my home server. Why is no company buildin…

It's worth noting that building a solid search index needs more than just text. The full Common Crawl archive includes metadata, links, page structure, and other signals that are needed for relevance and ranking, and to date is over 7 PiB, so that's rather more than tends to fit on home servers

Re: Two upstart search engines are teaming up to take on Google

#273
post #238

Earlier quoted context omitted.

That would be a victory for users.

A temporary reprieve, more like. Google will just wait for the competitor to fail and then go back to their old ways.

Yeah, that is what I meant. :(

Re: Two upstart search engines are teaming up to take on Google

#275

Earlier quoted context omitted.

The presence of affiliate links should be the biggest indicator that the information provided is going to be heavily financially incentivized (ie. bad). Would be much simpler to calculate and detect. I would be super surprised if Google's algorithm doesn't already know how to detect affiliate links.

Many high quality youtube videos on e.g. wristwatch repair have lots (sometimes dozens) affiliate links to every tool used, which is just an additional way for the author to make some money. If your search engine filters this out, it's worthless.

While you might like and trust these youtube creators, incentives drive everything.

It's downright silly to believe these creators won't behave rationally, given those incentives (eg. lean more and more toward pushing products with the highest affiliate payouts).

I've observed this myself in pretty much any niche I follow. Every youtuber/creator eventually devolves into an affiliate shill, no matter how honest and high quality their content is at the beginning.

If you're getting paid when you say something is good, you're not trustworthy. Period. What's funny is we all used to understand this! The internet broke everyone's brain.

Re: Two upstart search engines are teaming up to take on Google

#276
post #267

Earlier quoted context omitted.

On top of that, search usually uses CPU instead of GPU. A large infrastructure with CPU’s is easier to reuse for jobs other than search.

These are great reasons why this business will be hard, but given how ChatGPT and Perplexity are making inroads into search traffic, you can't deny it's an experience consumers prefer.

I agree that there’s interest in it. I found ChatGPT and AI search very convenient in some situations where I used them. Other times they hallucinated. I have no idea, though, what customers prefer until I see large-scale surveys by companies not pushing A.I..

It could also become a differentiator allowing multiple suppliers. On one hand, you have people doing search for quality results. Other search engines include the AI results. The user could choose between them on a job by job basis or the search provider might, like !G in DDG, allow 3rd-party AI search as an option.

The bigger problem I have is with scale for the dollar. Search companies with their own indexes already mostly failed. There’s a few using Bing. It’s down to just three or four with their own index. Massive consolidation of power. If GPU’s and AI search cost massively more, wouldn’t that problem further increase?

Re: Two upstart search engines are teaming up to take on Google

#277
post #268

Earlier quoted context omitted.

Search is still better for getting to specific, existing documents you need. Even the RAG people have been finding that out with hybrid models becoming more popular over time. I also think you can update search indexes more cheaply than further pretraining LLM’s.

This is a great point, but I wonder how much of that kind of search intent is part of google's traffic. If that becomes the only reason people use Google I wonder if they'll go the way of Yahoo. Maybe that's hyperbolic, but there was a time when Yahoo's dominance seemed unquestionable (I'm old). To be clear, I'm not arguing for everything should be part of a pretrained LLM, but the experience of knowledge searching t…

I’ll add that I used to love Yahoo Directory. I couldn’t imagine it doing anything but grow. Sadly, it wasn’t to be.

Re: Two upstart search engines are teaming up to take on Google

#278

Earlier quoted context omitted.

See also: IndexNow [1], a protocol used by Bing, Naver, Yandex, Seznam, and Yep where sites can ping one of these search engines when a page is updated and all others will be immediately notified. Unfortunately it does seem somewhat closed as to requirements for joining as a search engine. [1]: https://www.indexnow.org/

What year was this thing created, because the /.well-known URI scheme[1] has existed for a long time and is designed for this kind of junk https://www.indexnow.org/documentation#:~:text=Hosting%20a%2... 1: https://www.iana.org/assignments/well-known-uris/well-known-...

It was announced October 18, 2021: https://blogs.bing.com/webmaster/october-2021/IndexNow-Insta...

Re: Two upstart search engines are teaming up to take on Google

#279

Earlier quoted context omitted.

To properly compete with Google search you have to: 1. Create your own index/crawl and true search engine on top of it, rather than delegating that back to bing or Google. In doing so, solve or address all the related problems like SEO/spam, and cover the “long tail” (expensive part) 2. Monetize it somehow to afford all of the personnel/hardware/software costs. Probably with your own version of Google’s ad products.…

Imo you need to do all of this, plus be compelling different. Nobody is going to beat Google playing Google's game.

The Chinese did it. Nationalism is a powerful force.

Re: Two upstart search engines are teaming up to take on Google

#280
post #264

Earlier quoted context omitted.

It occurs to me that competing with Google on their home turf and at their scale might be impossible. But doing what Google was good at, when they just a search engine, might not be all that much harder today than it was back then. And it may fly "under the radar" of Google's current business priorities.

I'm thinking about how Etsy took "the good bit" of Ebay and ran with it. Or how tyre-and-exhaust shops take "the profitable bit" of being a mechanic and run with it. I know there's a name for this, but it eludes me.

It could be they also took the bit that eBay abandoned. There was a point when eBay announced to the business press that they were going to de-emphasize the "America's yard sale" aspect, and shift towards being a regular e-commerce site like Amazon.
Post reply on HN