Live data from Hacker News

Waiting for dawn in search: Search index, Google rulings and impact on Kagi

blog.kagi.com

211–220 of 266 posts

Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi

#211
post #186

Earlier quoted context omitted.

Is crawling really solved? Any naive crawler is going to run into the problem that servers can give different responses to different clients which means you can show the crawler something different to what you show real users. That turns crawling into an antagonistic problem where the crawler developers need to continually be on the lookout for new ways of servers doing malicious things that poison/mislead the index.…

I don't mean to say it's trivial. I'm sure there are many hard problems such as the one you mention - though that particular one is more "cleaning the index" part which might work on top of the open common corpus. But my impression is that it's more a question of scale and engineering time than having to invent something new. (disclaimer: I also never worked on a internet-scale search system, maybe I'm very off the b…

Oh, ok. I misunderstood - I think we agree.

Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi

#212

> Layer 3: Paid, subscription-based search Should actually be - Layer 3: Paid, ad-free , subscription-based search. (It's a subtle omission that indicates the direction Kagi search will eventually take).

I don't see how it's conductive to the underlying big stakes discussion if we start condemning Kagi over this omission instead of assuming good faith. It's just distracting from the real issue at hand that is Google's monopoly and the details of the rulings and their enforcement. Call me naive, but I imagine Kagi would be VERY hesitant to force ads on their users, given that such a step would risk alienating a major…

Assuming good-faith requires reciprocal actions to reinforce it from the other side too, which I don't see happening. As I have pointed out elsewhere, they've stopped offering offline installers for their browser, and I suspect a major reason for that is to also collect telemetry / user data - a clever way to get around their advertised claim of "no telemetry" browser. After being burnt many times by Google and Apple ("trust us, we care about your privacy"), others and now streaming services ("no ads if you pay us, promise"), I just can't help being cynical as another for-profit company appears to be using the same tactics ... like I said, trust is earned, not demanded.

Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi

#213

> Layer 3: Paid, subscription-based search Should actually be - Layer 3: Paid, ad-free , subscription-based search. (It's a subtle omission that indicates the direction Kagi search will eventually take).

I don't see how it's conductive to the underlying big stakes discussion if we start condemning Kagi over this omission instead of assuming good faith. It's just distracting from the real issue at hand that is Google's monopoly and the details of the rulings and their enforcement. Call me naive, but I imagine Kagi would be VERY hesitant to force ads on their users, given that such a step would risk alienating a major…

Assuming good faith is a mistake now. Users need to be asking, demanding, proof, not just assuming anyone or anything else has their best interests in mind. I'm tired of seeing this; I'll assume good faith only once I'm not universally treated as an enemy to society.

With that said, Kagi has appeared friendly so far.

Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi

#214
post #23

Earlier quoted context omitted.

Seems like an open question as to whether that violates any laws. Another way to look at it is that if you publish a service on the web, you have limited rights to restrict what people do with it. Isn't that the logic Google search relies on in the first place? I didn't give permission for Google to crawl and index and deep link to my site (let alone summarize and train LLMs on it). They just did it anyway, because i…

Google's stance is "I can copy you and you can't stop me" as well as "You can't copy me, I'll sue you"

Google at least claims that noindex will keep your site from getting crawled [1]. Do people think this is false?

[1] https://developers.google.com/search/docs/crawling-indexing/...

Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi

#215
post #3

> Because direct licensing isn’t available to us on compatible terms, we - like many others - use third-party API providers for SERP-style results Crazy for a company to admit: "Google won't let us whitelabel their core product so we steal it and resell it."

It is basic antitrust practice. If a company starts to control a vertical so much so that they start to exclude others, they get broken up into components and ordered to offer the basic infrastructure service to others. This is how it worked for 100 years (read up on telecoms/fiber; train companies/railroads; heck, even roads used to belong to people in the UK). This is why we have net neutrality - I recommend Tubes by Andrew Blum to go the heart of the matter. Imagine Internet if Google was able to throttle other services if you are not using their own? Here the author is arguing the search index is like infra that needs to be shared for public good. The state will not confiscate it - Google will break it into an independent company, will start paying for it, and let others to pay as well. It's not whitelabeling, stealing and reselling. Gosh - just read a bit people.

Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi

#216
post #68
post #57

Earlier quoted context omitted.

> "Our own small-web index" Has Kagi ever said what this is? I wouldn't be at all surprised if it is just kagi.com pages or a download of Wikipedia.

https://github.com/kagisearch/smallweb

From that:

——

Criteria for posts to show on the website

If the blog is included in small web feed list (which means it has content in English, it is informational/educational by nature and it is not trying to sell anything) we check for these two things to show it on the site:

- Blog has recent posts (

- The website can appear in an iframe

——

Emphasis mine. Restricting visibility to blogs that post at least every week doesn’t feel very ‘small web’ to me.

Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi

#217

Earlier quoted context omitted.

Why does Google even need to know about your ladder? Build the bot, scale it up, save all the data, then release. You can now remove the ladder and obey robots.txt just like G. Just like G, once you have the data, you have the data. Why would you tell G that you are doing something? Why tell a competitor your plans at all? Just launch your product when the product is ready. I know that's anathema to SV startup logic,…

Cost, presumably. From the article: > Microsoft spent roughly $100 billion over 20 years on Bing and still holds single-digit share. If Microsoft cannot close the gap, no startup can do it alone.

Now that you mention it...

It's odd that Microsoft hasn't aggressively pushed for "openness". That's in the usual playbook for attacking a market leader.

(And then pull up the ladder once you've become king of hill.)

Microsoft will probably never topple Google, absent anti-monopolistic enforcement. But they can certainly attack Google's profits.

Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi

#218

Earlier quoted context omitted.

I've been using Kagi for the past few years, but I try to use a brand-agnostic language talking about web search; e.g. "I'm gonna search [the web] for it"; "Use your favorite search engine to look it up".

Likewise and if people say, "why don't you google that?" I usually reply (obviously to everyone's annoyance:-) "I don't use Google". The general response is a blank, uncomprehending look.

To a young enough audience, you will sound like you exclusively use ChatGPT instead.

Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi

#219

The statistics in this article sound like garbage to me. Google used by 90% or the world? ~20% of the human population lives in countries where Google is blocked. OTOH, Baidu is the #1 search engine in China, which has over 15% of the world’s population… but doesn’t reach 1%? These stats are made measuring US-based traffic, rather than “worldwide” as they claim.

To be fair, Kagi won't be used in China either.

I have used it from China, actually. Not big enough to be blocked.

Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi

#220
post #216
post #68

Earlier quoted context omitted.

https://github.com/kagisearch/smallweb

From that: —— Criteria for posts to show on the website If the blog is included in small web feed list (which means it has content in English, it is informational/educational by nature and it is not trying to sell anything) we check for these two things to show it on the site: - Blog has recent posts ( - The website can appear in an iframe —— Emphasis mine. Restricting visibility to blogs that post at least every wee…

[deleted]
Post reply on HN