While a lot of people are concerned with local model performance, I wonder how feasible is it now to run a local indexed web search? Surely running an old school Google is possible with the beefy AI rigs today. I know the problem will be crawling which would be bottlenecked by the ISP but I use Google to search SO, Wikipedia, programming language docs, Github issues, and AWS docs. I think a feasible workflow would be…
google.com/goto: Google's anti-scraping update
521–530 of 545 posts
Re: google.com/goto: Google's anti-scraping update
#522Earlier quoted context omitted.
Hamas claims that the journalists are Hamas, not Israel. And I'm not declaring any journalist "good" or "bad". I'm stating that a Hamas member is a valid target. Him moonlighting as a journalist does not change that.
Every claim need to be proven and everyone deserves a process. In a democracy of course, it may be different in a fascist state committing a holocaust.
Re: google.com/goto: Google's anti-scraping update
#523Earlier quoted context omitted.
Everything uses energy. AI is uniquely bad due to its scale. Comparing AI inference and training with posting a one-line comment on a website is like comparing wildfires with candles because they both produce heat. It's true but useless as an argument.
Complaining about AI is just a meme. There are tons of things that use way more energy for arguably way less utility to humans.
Re: google.com/goto: Google's anti-scraping update
#524As much as I am sad that Google died like 15 years ago, I am past the mourning phase. That was when they announced they were shifting from returning websites to "returning answers" and it has been a long slide into shittification I do enjoy using their free AI. For actual web search I actually like using Yandex. It reminds me of old Google, returning reasonable results and much less "shaping results to please our cor…
Re: google.com/goto: Google's anti-scraping update
#525Another "move" is suing companies like Autom, e.g., SerpApi
Google's Amended Complaint from their suit against SerpApi
https://ia801008.us.archive.org/25/items/gov.uscourts.cand.4...
"30. Copyright holders have authorized Google to implement access controls like SearchGuard for the content they license to Google, and in some cases insisted that Google do so. Googles authorization takes many forms. For example, Google has an agreement with a prominent licensing partner that holds copyrights to millions of works that it licenses Google to use in its Search results. Under the parties agreement, versions of which date back to 2017, Google is not only authorized, it is obligated to use commercially reasonable efforts to safeguard the licensed content against unauthorized third-party access. Other license agreements contain similar obligations. For example, another major content provider requires that Google ensure the content it licenses will not be available for download by third parties, thereby authorizing the implementation of technical access controls."
"31. In other cases, Googles authorization to implement access control measures like SearchGuard is part and parcel of the grant of licenses themselves, as Google and its licensors recognize that the value of the licensed rights would be undermined if others were free to access, take and resell the licensed content without restriction. For example, Google has a licensing agreement with Reddit, under which Reddit licenses Google to use the copyrighted content of both Reddit and its users in Search Services."
"32. Googles licensing partners have also expressly requested that Google prevent unauthorized access to licensed content. For example, when Reddit suspected that scrapers like SerpApi were accessing, taking, and reselling the content that Reddit had licensed to Google, it specifically asked Google to employ technical measures to prevent such unauthorized appropriation."
But this does not account for material that is not covered by the "license with a prominent licensing partner", its license with "another major content provider" or its agreement with Reddit
Google not only uses SearchGuard on SERPs containing links to the content covered by these licenses, it uses SearchGuard on _all_ SERPs
Google needs more than a "goto" update. It needs to update its terms to require _all_ copyright holders for the materials it has indexed and cached to give Google authorisation to use "technological protection measures" to deny access to certain members of the public, e.g., Google's perceived competitors including any Google user who "searches too fast"
Re: google.com/goto: Google's anti-scraping update
#526Re: google.com/goto: Google's anti-scraping update
#527Re: google.com/goto: Google's anti-scraping update
#528Earlier quoted context omitted.
FWIW Kagi has worked hard on their pricing over the last few years and has been trying different models. They're not "big search" the price is what it is so it can exist as a business. 100% fine to not be a customer obviously, but then you forfeit your license to complain about Google spying on you and ruining the web. Like they're trying to solve the problem. IMHO speaking just for myself I feel a moral duty to supp…
You can not find the value of 10 searches a day for 5 dollars and complain about Google. Yandex,DDG and Bing are major search engine choices. Kagi has limitations like not showing sites with ads which many people with ad blockers don't mind.
Re: google.com/goto: Google's anti-scraping update
#529Earlier quoted context omitted.
Kagi don't scrape, they pay other engines for API access: https://help.kagi.com/kagi/search-details/search-sources.htm... >Our search results also include anonymized API calls to all major search result providers worldwide
They pay scraping sites for results, including SERP API. When I pointed this out, last time Kagi was discussed on Hacker News, an employee of Kagi said that they're trying to build their own internal index, but he didn't provide details.
I believe they are building their own index but targeted towards useful results not included in the others
Re: google.com/goto: Google's anti-scraping update
#530Earlier quoted context omitted.
Impossible. The majority of websites firewall automated crawler traffic (because of the rise of the bots), only making exceptions for the largest search engines. There is no possibility of starting a new crawler.
It would be interesting to see a decentralised, residential collective that builds and publishes an index. There are surely enough interested people on HN alone that would be willing to run software at home to scrape a small slice of the internet.