Live data from Hacker News

Waiting for dawn in search: Search index, Google rulings and impact on Kagi

blog.kagi.com

251–260 of 266 posts

Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi

#251
post #73

Earlier quoted context omitted.

Scraping is hard. Very good scraping is even harder. And today, being a scraping business is veeery difficult; there are some "open"/public indices, but none of these other indices ever took off

Well sure yes, I don't contend with the fact that its hard, but if the top tech companies joined their heads I am sure if for example, Meta, Apple, MS have enough talent between to make an open source index if only to reap gains from the de-monopolization of it all.

They will prefer to band up with Google, and rip us off.

Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi

#252
> A government-backed, ad-free, intermediary-free, taxpayer-funded search service providing baseline, non-discriminatory access to information. Imagine search.org.

There is no way the government provides a search engine that doesn’t become a political football or weapon.

Maybe in a different age.

I completely agree that monopoly remedies, such as fair open paid licensing, are needed. I prefer that to breakups, when this kind of cooperative/competitive leveling works.

Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi

#253
post #71

Earlier quoted context omitted.

There is a practical limit that we can't cache results for too long; Search engine users are particularly sensitive to stale data, especially around current events. Without a holistic and realiable way to know when the cache ought to be invalidated, our caching is mostly focused on mitigating "abuse", e.g., someone / bunch of people spamming the same search in a short timespan; no sense in repeating all those upstrea…

To me, a lot of problems with "building a search engine" don't seem to be problems with "building a search engine," they seem to be problems with "building a Google." Nobody said a search engine needs to have fresh data, for example. Nor has anybody said a search engine needs to index the entire web. Yet these are two things every search engine tries to do, and then they usually fail to compare with Google. To put it…

I'll say that a search engine needs to have fresh data. When I search for a phrase from a reddit thread I saw earlier, I want that exact thread to be in the results.

When I search for a brand new restaurant, I want to see a map entry for that restaurant and a link to a newspaper article, ad, or facebook post announcing the opening of that restaurant (though I probably won't click on the third).

Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi

#254

> A government-backed, ad-free, intermediary-free, taxpayer-funded search service providing baseline, non-discriminatory access to information. Imagine search.org. There is no way the government provides a search engine that doesn’t become a political football or weapon. Maybe in a different age. I completely agree that monopoly remedies, such as fair open paid licensing, are needed. I prefer that to breakups, when t…

> There is no way the government provides a search engine that doesn’t become a political football or weapon.

Maybe it doesn't have to be based in the US? Maybe we could make this a world effort, run by a coalition instead, across border lines, like a library for the modern age.

Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi

#255
post #145

Earlier quoted context omitted.

Not only was eBay v. Bidder's Edge technically after Google existed, not before, more critically the slippery-slope interpretation of California trespass to chattels law the District Court relied on in it was considered and rejected by the California Supreme Court in Intel v. Hamidi (2003), and similar logic applied to other states trespass to chattels laws have been rejected by other courts since; eBay v. Bidder's E…

The point is, robots.txt was definitely a thing that people expected to be respected before and during google's early existence. This Kagi claim seems to be at least partially false: > Google built its index by crawling the open web before robots.txt was a widespread norm, often over publishers’ objections.

> robots.txt was definitely a thing that people expected to be respected before and during google's early existence

As someone who was a web developer at that time, robots.txt wasn't a "widespread norm" by a large margin, even if some individuals "expected it to be respected". Google's use of robots.txt + Google's own growth made robots.txt a "widespread norm" but I don't think many people who were active in the web-dev space at that time, would agree that it was a widespread norm before Google.

Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi

#256

Earlier quoted context omitted.

Let me AOL this for you

said no one ever

You clearly did not live in the world of watching two teens on computers in the same room hold two entirely different conversations out-loud and over AIM.

Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi

#258
post #22

Earlier quoted context omitted.

Google is only blocked in places where it would already be hard for a company with morals to work in, if not outright blocked as well. This probably represents traffic globally, excluding those places. Instead of downvoting blindly, please state which countries are currently blocking Google that would willingly allow Kagi, a AI/Privacy focused search engine company to exist in their domain? The results may surprise y…

Google and Facebook would be very happy to operate in China, but they're too closely tied to the US intelligence apparatus to agree to the terms that China requires.

For years Facebook wanted to get into Chinese market, so much that Zuckerberg asked Xi Jinping to name his child: https://www.the-independent.com/news/people/china-s-presiden...

No I didn't make this up.

And there was reporting like this: https://www.msn.com/en-us/news/world/zuckerberg-s-meta-was-w...

Although a few years they seemed to completely abandon the effort and started to criticize China, although I can't find the article.

You'll be amazed at how quickly Zuckerberg "adapts" to things. Which is why I never trust a single word that comes out of his mouth.

Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi

#259

If google is serving 90% traffic & others are unable to enter - Doesn't that mean google is doing something right for the customer and others are unable to outcompete it? Isn't this how life works?

Competition only works when we have an even field.

Like shooting your opponents in the leg before a Marathon will surely improve your chances, but it doesn't mean you are the best out of them. This is like the very tenet of markets, reaching as far back as Adam Smith.

"Funnily" enough this requires some external system that upholds the rules of the competition, e.g. governments. That's why busting monopolies make sense.

Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi

#260
post #60

Earlier quoted context omitted.

> "Could a meaningfully better search engine realistically displace Google today?” ChatGPT clearly demonstrated that displacing Google is possible. All previous monopoly arguments seemed even more flimsy after that.

ChatGPT did not build a search engine though. They built something else (equally impressive) and then were able to use their weight to enter the web search business where most sites now have to allow them in. While it's good that building other products is possible, it doesn't detract from the point that search engines are a de-facto monopoly.

(Not quite the same as a search engine, but to create a base model in an LLM, they pretty much did "download the whole internet" for it)
Post reply on HN