Live data from Hacker News

Two upstart search engines are teaming up to take on Google

wired.com

281–290 of 299 posts

Re: Two upstart search engines are teaming up to take on Google

#281

Earlier quoted context omitted.

To properly compete with Google search you have to: 1. Create your own index/crawl and true search engine on top of it, rather than delegating that back to bing or Google. In doing so, solve or address all the related problems like SEO/spam, and cover the “long tail” (expensive part) 2. Monetize it somehow to afford all of the personnel/hardware/software costs. Probably with your own version of Google’s ad products.…

I'd like to add to your #5. If Google deems you a legitimate threat, then they can just de-crapify their own search for a bit by going back to their old algo. It's extremely easy for them to fight back.

[deleted]

Re: Two upstart search engines are teaming up to take on Google

#282

Earlier quoted context omitted.

What year was this thing created, because the /.well-known URI scheme[1] has existed for a long time and is designed for this kind of junk https://www.indexnow.org/documentation#:~:text=Hosting%20a%2... 1: https://www.iana.org/assignments/well-known-uris/well-known-...

It was announced October 18, 2021: https://blogs.bing.com/webmaster/october-2021/IndexNow-Insta...

The irony of a website from 2 major search engines looking like it was made in the early 2000s doesn't escape me. But, to my original point, there's absolutely no way they were ignorant of well-known URIs

Re: Two upstart search engines are teaming up to take on Google

#284

The best way to stop SEO is to build a bot which detects commercial activity on the page (and/or availability of cart/payment controls), and commercial language in the text of the page (does the content look like an ad or product/sales page). Also combine with number of hops to a commercial page (with cart, prices, advertising material, etc), and create a distance metric. Offer the user a commercial activity slider.…

> The best way to stop SEO is to build a bot which detects commercial activity on the page (and/or availability of cart/payment controls), and commercial language in the text of the page (does the content look like an ad or product/sales page). Thanks for describing the detection algorithm, I will now design my spam sites to defeat it. This is how SEO works. There are no silver bullets.

All you can do is poison my results with garbage which won't make you money, which means you are paying out to do that.

Re: Two upstart search engines are teaming up to take on Google

#285
post #253

Earlier quoted context omitted.

This seems unrealistic. Sliding the commercial slider to zero is unlikely what anyone wants. Even Wikipedia would be hidden due to its incessant thirst for donations. Visiting a nice open source repo that happens to have Patreon or GitHub sponsor? Gone. Visiting the help page for a hardware product I've already purchased? That page probably upsells you on a higher tier product and is gone. Trying to get news? All new…

God i detest that attitude: all software sucks because it makes assumptions instead of as many options for the user as possible.

Yes. IMO ideally a good program isn't basic and easy to use for McNormie as it is functionally composable, like groups and algebras.

Re: Two upstart search engines are teaming up to take on Google

#286
post #4

Every new search engine I've seen was a Bing wrapper with sometimes light reranking. I understand that competing with Google was borderline impossible a decade ago. But in 2024, we have cheap compute, great OSS distributed DBs, powerful new vector search tech. Amateur search engines like Marginalia even run on consumer hardware. CommonCrawl text-only is ~100TB, and can fit on my home server. Why is no company buildin…

Isn’t it obvious? Look around. Hackers have been stamped out even on hacker news. Now it’s all FAANG lifers and MBA or VC types trying their hand at grifting. Nothing good comes of that. Whoever is making the next thing “for real” just gets acquired and shut down. Moreover, get out of your echo chamber and you’ll see that for a majority of humanity Google is the internet. You have to supplant the utility, not just th…

> Whoever is making the next thing “for real” just gets acquired and shut down.

why do you think amazon's trying to lay people off? they want them to create new startups that they can acquire.

Re: Two upstart search engines are teaming up to take on Google

#287

Earlier quoted context omitted.

Because nowdays more than ever content you need is in silos. Your facebooks/twiters/instagram/stack overflow/reddit ... And they all have limited expensive api's, and have bulk scrapping detection. Sure you can clobber together something that will work for a while, but you can't runn a buissness on that. Aditionaly most paywalled sites (like news) explicitly whitlist google and bing, and if someone cretes new site, t…

You're thinking too much by the rules. You can absolutely scrape them anyway. Probably the biggest relevant factor is CGNAT and other technologies that make you blend in with a crowd. If I run a scraper on my cellphone hotspot, the site can't block me without blocking a quarter of all cellphones in the country. If the site is less aggressively blocking but only has a per-IP rate limit, buy a subscription to one of th…

> You're thinking too much by the rules. You can absolutely scrape them anyway. Probably the biggest relevant factor is CGNAT and other technologies that make you blend in with a crowd. If I run a scraper on my cellphone hotspot, the site can't block me without blocking a quarter of all cellphones in the country.

I am familiar with most of that, and there is a BIG difference between trying to find a workaround for one site, that you scrape ocasionaly, than to to find workaround for all of the sites.

Big sites will definitely put entire ISP's behind annoying capachas that are designed to stop exactly this (if you ever wonder why you sometimes get capatchas that seem slow to load, have long animations, or other annoying slow things, that is why etc.)

And once you start making enough money to employ all the people you need for doing that consistently, they will find a jurisdiction or 3 where they can sue you.

Also good luck finding residential/mobile ISP's that will stand by, and not try to throttle you after a while.

You definitively can get away with doing all of that for a while, but you absolutely can't build sustainable businesses on that.

Re: Two upstart search engines are teaming up to take on Google

#288
post #270

Earlier quoted context omitted.

What are the rules?

You can literally do a web search for “new rules for big tech”.

Ok, but which rule specifically allows these two search companies to compete?

Re: Two upstart search engines are teaming up to take on Google

#289

Earlier quoted context omitted.

To properly compete with Google search you have to: 1. Create your own index/crawl and true search engine on top of it, rather than delegating that back to bing or Google. In doing so, solve or address all the related problems like SEO/spam, and cover the “long tail” (expensive part) 2. Monetize it somehow to afford all of the personnel/hardware/software costs. Probably with your own version of Google’s ad products.…

I've heard more or less the same thing said about IBM, Walmart, Microsoft, etc. There is one way to compete with Google search. Google search is a general search engine. But what if you only wanted medical information? information about cars? information about physics? electronics? history? etc. Specializing makes for a much smaller scope of the problem, and being specialized means it could deliver more useful result…

I never go to IMDb directly as both it and its search function are noticeably slower than Google’s. So I just tend to Google “ IMDb” and click the first result.

Re: Two upstart search engines are teaming up to take on Google

#290
post #270

Earlier quoted context omitted.

You can literally do a web search for “new rules for big tech”.

Ok, but which rule specifically allows these two search companies to compete?

It’s in the article:

> But Kroll believes tech advances have made affordable indexing more possible, and new EU regulations limiting the power of gatekeepers such as Google are making it a worthwhile pursuit.

Which links to another Wired story:

https://www.wired.com/story/europe-dma-breaking-open-big-tec...

They’re talking about the Digital Markets Act.

https://en.wikipedia.org/wiki/Digital_Markets_Act

Which is also what you’d get inundated by when doing the search I mentioned above. None of this is hard to find. If you want to discuss the article, it’s expected you make a minimum of effort to check what it says.

Post reply on HN