Earlier quoted context omitted.
To properly compete with Google search you have to: 1. Create your own index/crawl and true search engine on top of it, rather than delegating that back to bing or Google. In doing so, solve or address all the related problems like SEO/spam, and cover the “long tail” (expensive part) 2. Monetize it somehow to afford all of the personnel/hardware/software costs. Probably with your own version of Google’s ad products.…
I'd like to add to your #5. If Google deems you a legitimate threat, then they can just de-crapify their own search for a bit by going back to their old algo. It's extremely easy for them to fight back.
Two upstart search engines are teaming up to take on Google
281–290 of 299 posts
Re: Two upstart search engines are teaming up to take on Google
#282Earlier quoted context omitted.
What year was this thing created, because the /.well-known URI scheme[1] has existed for a long time and is designed for this kind of junk https://www.indexnow.org/documentation#:~:text=Hosting%20a%2... 1: https://www.iana.org/assignments/well-known-uris/well-known-...
It was announced October 18, 2021: https://blogs.bing.com/webmaster/october-2021/IndexNow-Insta...
Re: Two upstart search engines are teaming up to take on Google
#283Re: Two upstart search engines are teaming up to take on Google
#284The best way to stop SEO is to build a bot which detects commercial activity on the page (and/or availability of cart/payment controls), and commercial language in the text of the page (does the content look like an ad or product/sales page). Also combine with number of hops to a commercial page (with cart, prices, advertising material, etc), and create a distance metric. Offer the user a commercial activity slider.…
> The best way to stop SEO is to build a bot which detects commercial activity on the page (and/or availability of cart/payment controls), and commercial language in the text of the page (does the content look like an ad or product/sales page). Thanks for describing the detection algorithm, I will now design my spam sites to defeat it. This is how SEO works. There are no silver bullets.
Re: Two upstart search engines are teaming up to take on Google
#285Earlier quoted context omitted.
This seems unrealistic. Sliding the commercial slider to zero is unlikely what anyone wants. Even Wikipedia would be hidden due to its incessant thirst for donations. Visiting a nice open source repo that happens to have Patreon or GitHub sponsor? Gone. Visiting the help page for a hardware product I've already purchased? That page probably upsells you on a higher tier product and is gone. Trying to get news? All new…
God i detest that attitude: all software sucks because it makes assumptions instead of as many options for the user as possible.
Re: Two upstart search engines are teaming up to take on Google
#286Every new search engine I've seen was a Bing wrapper with sometimes light reranking. I understand that competing with Google was borderline impossible a decade ago. But in 2024, we have cheap compute, great OSS distributed DBs, powerful new vector search tech. Amateur search engines like Marginalia even run on consumer hardware. CommonCrawl text-only is ~100TB, and can fit on my home server. Why is no company buildin…
Isn’t it obvious? Look around. Hackers have been stamped out even on hacker news. Now it’s all FAANG lifers and MBA or VC types trying their hand at grifting. Nothing good comes of that. Whoever is making the next thing “for real” just gets acquired and shut down. Moreover, get out of your echo chamber and you’ll see that for a majority of humanity Google is the internet. You have to supplant the utility, not just th…
why do you think amazon's trying to lay people off? they want them to create new startups that they can acquire.
Re: Two upstart search engines are teaming up to take on Google
#287Earlier quoted context omitted.
Because nowdays more than ever content you need is in silos. Your facebooks/twiters/instagram/stack overflow/reddit ... And they all have limited expensive api's, and have bulk scrapping detection. Sure you can clobber together something that will work for a while, but you can't runn a buissness on that. Aditionaly most paywalled sites (like news) explicitly whitlist google and bing, and if someone cretes new site, t…
You're thinking too much by the rules. You can absolutely scrape them anyway. Probably the biggest relevant factor is CGNAT and other technologies that make you blend in with a crowd. If I run a scraper on my cellphone hotspot, the site can't block me without blocking a quarter of all cellphones in the country. If the site is less aggressively blocking but only has a per-IP rate limit, buy a subscription to one of th…
I am familiar with most of that, and there is a BIG difference between trying to find a workaround for one site, that you scrape ocasionaly, than to to find workaround for all of the sites.
Big sites will definitely put entire ISP's behind annoying capachas that are designed to stop exactly this (if you ever wonder why you sometimes get capatchas that seem slow to load, have long animations, or other annoying slow things, that is why etc.)
And once you start making enough money to employ all the people you need for doing that consistently, they will find a jurisdiction or 3 where they can sue you.
Also good luck finding residential/mobile ISP's that will stand by, and not try to throttle you after a while.
You definitively can get away with doing all of that for a while, but you absolutely can't build sustainable businesses on that.
Re: Two upstart search engines are teaming up to take on Google
#288Re: Two upstart search engines are teaming up to take on Google
#289Earlier quoted context omitted.
To properly compete with Google search you have to: 1. Create your own index/crawl and true search engine on top of it, rather than delegating that back to bing or Google. In doing so, solve or address all the related problems like SEO/spam, and cover the “long tail” (expensive part) 2. Monetize it somehow to afford all of the personnel/hardware/software costs. Probably with your own version of Google’s ad products.…
I've heard more or less the same thing said about IBM, Walmart, Microsoft, etc. There is one way to compete with Google search. Google search is a general search engine. But what if you only wanted medical information? information about cars? information about physics? electronics? history? etc. Specializing makes for a much smaller scope of the problem, and being specialized means it could deliver more useful result…
Re: Two upstart search engines are teaming up to take on Google
#290Earlier quoted context omitted.
You can literally do a web search for “new rules for big tech”.
Ok, but which rule specifically allows these two search companies to compete?
> But Kroll believes tech advances have made affordable indexing more possible, and new EU regulations limiting the power of gatekeepers such as Google are making it a worthwhile pursuit.
Which links to another Wired story:
https://www.wired.com/story/europe-dma-breaking-open-big-tec...
They’re talking about the Digital Markets Act.
https://en.wikipedia.org/wiki/Digital_Markets_Act
Which is also what you’d get inundated by when doing the search I mentioned above. None of this is hard to find. If you want to discuss the article, it’s expected you make a minimum of effort to check what it says.