Live data from Hacker News

Waiting for dawn in search: Search index, Google rulings and impact on Kagi

blog.kagi.com

111–120 of 266 posts

Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi

#111
post #95

Earlier quoted context omitted.

> 1. One is building the index, which is a lot harder without a google offering its own API to boot. If other tech companies really wanted to break this monopoly, why can't they just do it? FTA: > Context matters: Google built its index by crawling the open web before robots.txt was a widespread norm, often over publishers’ objections. Today, publishers “consent” to Google’s crawling because the alternative - being i…

robots.txt was being enforced in court before google even existed, let alone before google got so huge: > The robots.txt played a role in the 1999 legal case of eBay v. Bidder's Edge,[12] where eBay attempted to block a bot that did not comply with robots.txt, and in May 2000 a court ordered the company operating the bot to stop crawling eBay's servers using any automatic means, by legal injunction on the basis of tr…

[flagged]

Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi

#112

Earlier quoted context omitted.

> 1. One is building the index, which is a lot harder without a google offering its own API to boot. If other tech companies really wanted to break this monopoly, why can't they just do it? FTA: > Context matters: Google built its index by crawling the open web before robots.txt was a widespread norm, often over publishers’ objections. Today, publishers “consent” to Google’s crawling because the alternative - being i…

A classic case of climbing the wall, and pulling the ladder up afterward. Others try to build their own ladder, and Google uses their deep pockets and political influence to knock the ladder over before it reaches the top.

Why does Google even need to know about your ladder? Build the bot, scale it up, save all the data, then release. You can now remove the ladder and obey robots.txt just like G. Just like G, once you have the data, you have the data.

Why would you tell G that you are doing something? Why tell a competitor your plans at all? Just launch your product when the product is ready. I know that's anathema to SV startup logic, but in this case it's good business

Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi

#113
post #44

Kagi's "waiting for dawn" is just waiting for Google to legitimize their reseller business Meanwhile, users pay a premium to pretend they're not using Google Fascinating delusion

> Meanwhile, users pay a premium to pretend they're not using Google My searches can’t be tied to me by Google for their ad targeting: this is worth paying a premium for, and I am glad Kagi are providing this service. You seem to have a very limited understanding of the value Kagi provides.

I have a limited understanding of the value Christianity provides. That neither means that Christianity provides no value, nor does it mean that God exists.

Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi

#114
post #50

Earlier quoted context omitted.

> If other tech companies really wanted to break this monopoly, why can't they just do it Google is a verb, nobody can compete with that level of mindshare.

So were AOL, and Skype

I don't ever recall anyone using AOL as a verb. How would you do that?

Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi

#115
post #92

Earlier quoted context omitted.

Yeah but no one uses it. I am not even sure people that are forced to use it like using it because it was productized it pretty poorly. After all who wants another google? They invested 100 Billion dollars, which is a lot of wasted money TBH. Search indexes are hard, surely, but if you were to strip it to just a good index on the browser, made it free, kept it fresh, it cannot be 100 billion dollars to build. Then yo…

> Yeah but no one uses it. I am not even sure people like using it because it was productized it pretty poorly. They invested 100 Billion dollars, which is a lot of wasted money TBH. I mean... Duckduckgo uses bing api iirc and I use duckduckgo and many people use duckduckgo. I also used bing once because bing used to cache websites which weren't available in wayback archive, I don't know how but It was pretty cool so…

I would imagine the users of DDG to be closer to a rounding error than an actual percentage of users. I'd imagine theGoog would love and hate to have 100%. They'd love it because all the data, and hate it as it would prove the monopoly. At the end of the day, the % that is not going to them probably doesn't cause theGoog to lose much sleep

Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi

#116

The statistics in this article sound like garbage to me. Google used by 90% or the world? ~20% of the human population lives in countries where Google is blocked. OTOH, Baidu is the #1 search engine in China, which has over 15% of the world’s population… but doesn’t reach 1%? These stats are made measuring US-based traffic, rather than “worldwide” as they claim.

Yes the stats don't make sense. It appears to be an issue with StatsCounter. The Search Engine wikipedia article [1] has a section on Russia and East Asia market share, which confirms that the roll up used for world wide counts is off, unless the number of people using the Internet is drastically different in some of the countries. Russia * Yandex: 70.7% * Google: 23.3% China: * Baidu: 59.3% * Other domestic engines:…

Maybe it's the same logic that says you can lower the prices of things >100%

Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi

#117
Google's advantage is not just in its index and algorithms, it is that it has built a self-reinforcing flywheel that data mines human attention at massive scale to improve their search results.

This comment (https://news.ycombinator.com/item?id=46709957) points out that Google got its start via PageRank, which essentially ranked sites based on links created by humans. As such, its primary heuristic was what humans thought was good content. Turns out, this is still how they operate.

Basically, as people search and navigate the results, Google harvests their clicks, hovers, dwell-time and other browsing behavior -- i.e. tracking what they pay attention to -- to extract critical signals to "learn" which pages the users actually found useful for the given query. This helps it rank results better and improve search overall, which keeps people coming back, which in turns gives them more queries and data, which improves their results... a never-ending flywheel.

And competitors have no hope of matching this, because if you look at the infrastructure Google has built to harvest this data, it is so much bigger than the massive index! They harvest data through Chrome, ad tracking, Android, Google Analytics, cookies (for which they built Gmail!), YouTube, Maps, and so much more. So to compete with Google Search, you don't need just a massive index, you also need the extensive web infra footprint to harvest user interactions at massive scale, meaning the most popular and widely deployed browser, mobile OS, ad footprint, analytics, email provider, maps...

This also explains why Google spends so many billions in "traffic acquisition costs" (i.e. payments for being the Search default) every year, because that is a direct driver to both, 1) ad revenue, and 2) maintaining its search quality.

This wasn't really a secret, but it turned out to be a major point in the recent Antitrust trial, which is why the proposed remedies (as TFA mentions) include the sharing of search index and "interaction data."

We all knew "if you're not paying for it, you're the product" but the fascinating thing with Google is:

- They charge advertisers to monetize our attention;

- They harvest our attention to better rank results;

- They provide better results, which keeps us coming back, and giving them even more of our attention!

Attention is all you need, indeed.

Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi

#118
post #104

Earlier quoted context omitted.

Google is not blocked in the USA.

Interesting. I'm in the US and use Kagi everyday.

I read it more as "company having morals". Not many US companies have "morals".

Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi

#119
post #104

Earlier quoted context omitted.

Interesting. I'm in the US and use Kagi everyday.

I read it more as "company having morals". Not many US companies have "morals".

Google doesn't, Kagi seems to (hopefully). I meant this more as a jab at countries willing to block Google, as they're generally dictatorships / authoritarian in nature. Oh the irony, as an american saying this in 2026....

Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi

#120
I think one side problem is that part of the web is not even searchable with a search engine.

Here are some examples:

- Discord

- WeChat (is it the web?)

- Rednote

- TikTok (partially)

- X (partially)

- JSTOR (it finds daily, but you find more stuff on the website directly)

- any stuff with a login, obviously.

Post reply on HN