Live data from Hacker News

Two upstart search engines are teaming up to take on Google

wired.com

251–260 of 299 posts

Re: Two upstart search engines are teaming up to take on Google

#251

Earlier quoted context omitted.

No search engine is refreshing every website every minute. Most websites don't update frequently, and if you poll them more than once every month, your crawler will get blocked incredibly fast. The problem of being able to provide fresh results is best solved by having different tiers of indices, one for frequently updating content, and one for slowly updating content with a weekly or monthly cadence. You can get a l…

See also: IndexNow [1], a protocol used by Bing, Naver, Yandex, Seznam, and Yep where sites can ping one of these search engines when a page is updated and all others will be immediately notified. Unfortunately it does seem somewhat closed as to requirements for joining as a search engine. [1]: https://www.indexnow.org/

What year was this thing created, because the /.well-known URI scheme[1] has existed for a long time and is designed for this kind of junk https://www.indexnow.org/documentation#:~:text=Hosting%20a%2...

1: https://www.iana.org/assignments/well-known-uris/well-known-...

Re: Two upstart search engines are teaming up to take on Google

#252

Earlier quoted context omitted.

> You will probably at the very minimum need to implement your own platform (device, OS, browser alone probably won’t cut it). What do you mean? Google gained dominance just being a dot com URL. This in a time when competitors were already very well established.

Google got big when being a dot-com URL meant something and the primary way the average person accessed the Internet was through a desktop PC. Neither of those things is true anymore.

That's why you got to have an app. People install more apps than ever before, it's completely normal.

Re: Two upstart search engines are teaming up to take on Google

#253

The best way to stop SEO is to build a bot which detects commercial activity on the page (and/or availability of cart/payment controls), and commercial language in the text of the page (does the content look like an ad or product/sales page). Also combine with number of hops to a commercial page (with cart, prices, advertising material, etc), and create a distance metric. Offer the user a commercial activity slider.…

This seems unrealistic. Sliding the commercial slider to zero is unlikely what anyone wants. Even Wikipedia would be hidden due to its incessant thirst for donations. Visiting a nice open source repo that happens to have Patreon or GitHub sponsor? Gone. Visiting the help page for a hardware product I've already purchased? That page probably upsells you on a higher tier product and is gone. Trying to get news? All news sites have ads now and they are gone; the only ones without ads are those paid for by Big Oil or Big Pharma or similar with a hidden agenda to influence you.

Instead of bending the web into something it is not, have you tried simply visiting a library and getting all information from there?

Re: Two upstart search engines are teaming up to take on Google

#254

The best way to stop SEO is to build a bot which detects commercial activity on the page (and/or availability of cart/payment controls), and commercial language in the text of the page (does the content look like an ad or product/sales page). Also combine with number of hops to a commercial page (with cart, prices, advertising material, etc), and create a distance metric. Offer the user a commercial activity slider.…

> The best way to stop SEO is to build a bot which detects commercial activity on the page (and/or availability of cart/payment controls), and commercial language in the text of the page (does the content look like an ad or product/sales page).

Thanks for describing the detection algorithm, I will now design my spam sites to defeat it.

This is how SEO works. There are no silver bullets.

Re: Two upstart search engines are teaming up to take on Google

#255

For me, it seems the direction for search is going towards AI sites. (Gemini, ChatGPT) Trying to reinvent Google/ Search in 2024 seems a bit like jumping the shark

Completely agree. Search engines are dead to me.

For general information like recipes, tech stuff, how-to's, etc LLMs blow search engine results out of the water. Even the local/contextual content is better. The new maps feature in ChatGPT is amazing.

The only case where a search engine beats an LLM for me is when I'm looking for a site where the site itself provides a bespoke experience like shopping, where the engine provides a quicker path to the url I don't yet have in my browser history, like " Tour" -> somebandofficial.com

Re: Two upstart search engines are teaming up to take on Google

#256

The best way to stop SEO is to build a bot which detects commercial activity on the page (and/or availability of cart/payment controls), and commercial language in the text of the page (does the content look like an ad or product/sales page). Also combine with number of hops to a commercial page (with cart, prices, advertising material, etc), and create a distance metric. Offer the user a commercial activity slider.…

> bad behavior stems from the drive to make money

“The only moral income is my income”

It’s definitely an interesting POV. A lot of the way Google works comes from the fact that people use Google at the moment they’ve already decided they want to buy something. Versus say Instagram where you might see ads but you haven’t decided to buy something and then use Instagram. It comes down to what people other than you are using search engines for.

Re: Two upstart search engines are teaming up to take on Google

#257
post #253

The best way to stop SEO is to build a bot which detects commercial activity on the page (and/or availability of cart/payment controls), and commercial language in the text of the page (does the content look like an ad or product/sales page). Also combine with number of hops to a commercial page (with cart, prices, advertising material, etc), and create a distance metric. Offer the user a commercial activity slider.…

This seems unrealistic. Sliding the commercial slider to zero is unlikely what anyone wants. Even Wikipedia would be hidden due to its incessant thirst for donations. Visiting a nice open source repo that happens to have Patreon or GitHub sponsor? Gone. Visiting the help page for a hardware product I've already purchased? That page probably upsells you on a higher tier product and is gone. Trying to get news? All new…

God i detest that attitude: all software sucks because it makes assumptions instead of as many options for the user as possible.

Re: Two upstart search engines are teaming up to take on Google

#258

The best way to stop SEO is to build a bot which detects commercial activity on the page (and/or availability of cart/payment controls), and commercial language in the text of the page (does the content look like an ad or product/sales page). Also combine with number of hops to a commercial page (with cart, prices, advertising material, etc), and create a distance metric. Offer the user a commercial activity slider.…

The presence of affiliate links should be the biggest indicator that the information provided is going to be heavily financially incentivized (ie. bad). Would be much simpler to calculate and detect. I would be super surprised if Google's algorithm doesn't already know how to detect affiliate links.

Many high quality youtube videos on e.g. wristwatch repair have lots (sometimes dozens) affiliate links to every tool used, which is just an additional way for the author to make some money. If your search engine filters this out, it's worthless.

Re: Two upstart search engines are teaming up to take on Google

#259

The best way to stop SEO is to build a bot which detects commercial activity on the page (and/or availability of cart/payment controls), and commercial language in the text of the page (does the content look like an ad or product/sales page). Also combine with number of hops to a commercial page (with cart, prices, advertising material, etc), and create a distance metric. Offer the user a commercial activity slider.…

> bad behavior stems from the drive to make money “The only moral income is my income” It’s definitely an interesting POV. A lot of the way Google works comes from the fact that people use Google at the moment they’ve already decided they want to buy something. Versus say Instagram where you might see ads but you haven’t decided to buy something and then use Instagram. It comes down to what people other than you are…

> A lot of the way Google works comes from the fact that people use Google at the moment they’ve already decided they want to buy something

You missed a key part of the comment:

> Offer the user a commercial activity slider. […] if a Commercial slider is slid to low or zero

Re: Two upstart search engines are teaming up to take on Google

#260

The best way to stop SEO is to build a bot which detects commercial activity on the page (and/or availability of cart/payment controls), and commercial language in the text of the page (does the content look like an ad or product/sales page). Also combine with number of hops to a commercial page (with cart, prices, advertising material, etc), and create a distance metric. Offer the user a commercial activity slider.…

> bad behavior stems from the drive to make money “The only moral income is my income” It’s definitely an interesting POV. A lot of the way Google works comes from the fact that people use Google at the moment they’ve already decided they want to buy something. Versus say Instagram where you might see ads but you haven’t decided to buy something and then use Instagram. It comes down to what people other than you are…

> A lot of the way Google works comes from the fact that people use Google at the moment they’ve already decided they want to buy something.

It is exactly the point at which returning sponsored links maximally screws the user over.

Post reply on HN