Live data from Hacker News

Waiting for dawn in search: Search index, Google rulings and impact on Kagi

blog.kagi.com

81–90 of 266 posts

Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi

#81
post #62
post #24

> Building a comparable one from scratch is like building a parallel national railroad.. Not too be pedantic here but I do have a noob question or two here: 1. One is building the index, which is a lot harder without a google offering its own API to boot. If other tech companies really wanted to break this monopoly, why can't they just do it — like they did with LLM training for base models with the infamous "pile" d…

A huge amount of the web is only crawlable with a googlebot user-agent and specific source IPs.

> And given you-know-what, the battle to establish a new search crawler will be harder than ever. Crawlers are now presumed guilty of scraping for AI services until proven innocent.

I have always wondered but how does wayback machine work, is there no way that we can use wayback archive and then run a index on top of every wayback archive somehow?

Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi

#82
post #77
post #37

Kagi should start building an index of sites that are trying to escape the current slop internet. It’s know they have the Small Web thing. But I’d like to see an index of a “neo internet” that blocks Google et al.

I've been tossing around the very early idea of seeing what we can do to elevate alcoves of the web such as Gemini[1] through Kagi. I am slightly conscious of that some people might not like us operating in that space, it's been on my TODO to poll people about it and take a quick pulse. I love the tech and think we could give it meaningful exposure. Is this along the lines of what you have in mind - any other active…

Relevant https://github.com/kagisearch/smallweb/pull/425

Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi

#83
post #24

> Building a comparable one from scratch is like building a parallel national railroad.. Not too be pedantic here but I do have a noob question or two here: 1. One is building the index, which is a lot harder without a google offering its own API to boot. If other tech companies really wanted to break this monopoly, why can't they just do it — like they did with LLM training for base models with the infamous "pile" d…

> 1. One is building the index, which is a lot harder without a google offering its own API to boot. If other tech companies really wanted to break this monopoly, why can't they just do it? FTA: > Context matters: Google built its index by crawling the open web before robots.txt was a widespread norm, often over publishers’ objections. Today, publishers “consent” to Google’s crawling because the alternative - being i…

A classic case of climbing the wall, and pulling the ladder up afterward. Others try to build their own ladder, and Google uses their deep pockets and political influence to knock the ladder over before it reaches the top.

Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi

#84
post #62
post #24

> Building a comparable one from scratch is like building a parallel national railroad.. Not too be pedantic here but I do have a noob question or two here: 1. One is building the index, which is a lot harder without a google offering its own API to boot. If other tech companies really wanted to break this monopoly, why can't they just do it — like they did with LLM training for base models with the infamous "pile" d…

A huge amount of the web is only crawlable with a googlebot user-agent and specific source IPs.

I do not know a lot about this subject, but couldn’t you make a pretty decent index off of common crawl? It seems to me the bar is so low you wouldn’t have to have everything. Especially if your goal was not monetization with ads.

Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi

#85
post #3

> Because direct licensing isn’t available to us on compatible terms, we - like many others - use third-party API providers for SERP-style results Crazy for a company to admit: "Google won't let us whitelabel their core product so we steal it and resell it."

Is it much different than what Google AI Summaries do?

Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi

#86
post #24

> Building a comparable one from scratch is like building a parallel national railroad.. Not too be pedantic here but I do have a noob question or two here: 1. One is building the index, which is a lot harder without a google offering its own API to boot. If other tech companies really wanted to break this monopoly, why can't they just do it — like they did with LLM training for base models with the infamous "pile" d…

> 1. One is building the index, which is a lot harder without a google offering its own API to boot. If other tech companies really wanted to break this monopoly, why can't they just do it? FTA: > Context matters: Google built its index by crawling the open web before robots.txt was a widespread norm, often over publishers’ objections. Today, publishers “consent” to Google’s crawling because the alternative - being i…

True. But the thing is if one says "We will make sure your site is in a world wide freely availabled index" which is kept fresh, google's monopoly ship already begins to take on water. Here is a appropriate line from a completely different domain of rare earth metals from The Economist on the chinese govt's weaponization of rare earths[1]:

> Reducing its share from 90% to 80% may not sound like much, but it would imply a doubling in size of alternative sources of supply, giving China’s customers far more room for manoeuvre.

[1] https://archive.ph/POkHZ#selection-1233.117-1233.302

Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi

#87
I am rooting for Kagi here, and I applaud their transparency on such matters. It is quite enlightening for someone like me who understands technology but knows little about the inner workings of search.

It remains to be seen how or if the remedies will be enforced, and, of course, how Google will choose to comply with them. I am not optimistic, but at least there is some hope.

As an aside: The 1998 white paper by Brin and Page is remarkable to read knowing what Google has become.

Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi

#88
post #22

The statistics in this article sound like garbage to me. Google used by 90% or the world? ~20% of the human population lives in countries where Google is blocked. OTOH, Baidu is the #1 search engine in China, which has over 15% of the world’s population… but doesn’t reach 1%? These stats are made measuring US-based traffic, rather than “worldwide” as they claim.

Google is only blocked in places where it would already be hard for a company with morals to work in, if not outright blocked as well. This probably represents traffic globally, excluding those places. Instead of downvoting blindly, please state which countries are currently blocking Google that would willingly allow Kagi, a AI/Privacy focused search engine company to exist in their domain? The results may surprise y…

Google is not blocked in the USA.

Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi

#89

If Google provides a Search Index it will be the censored version therefore still politically acceptable. The “Layer 1” idea will not happen.

That's why Kagi combines results from multiple sources, just as it does with Yandex.

Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi

#90

Earlier quoted context omitted.

I wouldn't trust a nationalized search engine company. That said, there are projects like Common Crawl and in Europe, Ecosia + Qwant. I personally would like to see a search enginge PaaS and a music streaming library PaaS that would let others hook up and pay direct usage fees.

An interoperable search index access standard might work. We've done something similar for peering and the backbone of the IP-layer interconnects themselves.

You have to make it economically preferable, and there's No known solution to this. Large networks are still using their positions to bully smaller ones off the IP-layer internet backbone.
Post reply on HN