Live data from Hacker News

Waiting for dawn in search: Search index, Google rulings and impact on Kagi

blog.kagi.com

71–80 of 266 posts

Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi

#71

I hope they cache search results to further reduce the number of calls to Google. And Marginalia Search was not mentioned? Marginalia Search says they are licensing their index to Kagi. Perhaps it's counted under "Our own small-web index" which is highly misleading if true.

There is a practical limit that we can't cache results for too long; Search engine users are particularly sensitive to stale data, especially around current events. Without a holistic and realiable way to know when the cache ought to be invalidated, our caching is mostly focused on mitigating "abuse", e.g., someone / bunch of people spamming the same search in a short timespan; no sense in repeating all those upstream calls.

Most "cost saving engineering" is involved in finding cases/hueristics where we only need to use a subset of sources and omitting calls in the first place, without compromising quality. For example, we probably don't need to fire all of our sources to service a query like "youtube" or "facebook".

Marginalia data is physically consolidated into the same infra that we use for small web results in our SERP, but also among other small scale sources besides those two. That line is simply referring directly to https://kagi.com/smallweb (https://github.com/kagisearch/smallweb).

Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi

#72
post #24

> Building a comparable one from scratch is like building a parallel national railroad.. Not too be pedantic here but I do have a noob question or two here: 1. One is building the index, which is a lot harder without a google offering its own API to boot. If other tech companies really wanted to break this monopoly, why can't they just do it — like they did with LLM training for base models with the infamous "pile" d…

> 1. One is building the index, which is a lot harder without a google offering its own API to boot. If other tech companies really wanted to break this monopoly, why can't they just do it?

FTA:

> Context matters: Google built its index by crawling the open web before robots.txt was a widespread norm, often over publishers’ objections. Today, publishers “consent” to Google’s crawling because the alternative - being invisible on a platform with 90% market share - is economically unacceptable. Google now enforces ToS and robots.txt against others from a position of monopoly power it accumulated without those constraints. The rules Google enforces today are not the rules it played by when building its dominance.

Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi

#73
post #24

> Building a comparable one from scratch is like building a parallel national railroad.. Not too be pedantic here but I do have a noob question or two here: 1. One is building the index, which is a lot harder without a google offering its own API to boot. If other tech companies really wanted to break this monopoly, why can't they just do it — like they did with LLM training for base models with the infamous "pile" d…

Scraping is hard. Very good scraping is even harder. And today, being a scraping business is veeery difficult; there are some "open"/public indices, but none of these other indices ever took off

Well sure yes, I don't contend with the fact that its hard, but if the top tech companies joined their heads I am sure if for example, Meta, Apple, MS have enough talent between to make an open source index if only to reap gains from the de-monopolization of it all.

Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi

#74

The statistics in this article sound like garbage to me. Google used by 90% or the world? ~20% of the human population lives in countries where Google is blocked. OTOH, Baidu is the #1 search engine in China, which has over 15% of the world’s population… but doesn’t reach 1%? These stats are made measuring US-based traffic, rather than “worldwide” as they claim.

To be fair, Kagi won't be used in China either.

Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi

#75

Earlier quoted context omitted.

If Google Search is shit, why does Kagi want access to it?

They want access to the index. They will perform their own sorting to determine the best results to show from that index.

…without having advertiser interests to cater to.

Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi

#77
post #37

Kagi should start building an index of sites that are trying to escape the current slop internet. It’s know they have the Small Web thing. But I’d like to see an index of a “neo internet” that blocks Google et al.

I've been tossing around the very early idea of seeing what we can do to elevate alcoves of the web such as Gemini[1] through Kagi. I am slightly conscious of that some people might not like us operating in that space, it's been on my TODO to poll people about it and take a quick pulse. I love the tech and think we could give it meaningful exposure.

Is this along the lines of what you have in mind - any other active efforts you're aware of that you think we should look into?

[1] https://en.wikipedia.org/wiki/Gemini_(protocol)

Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi

#78
post #60

Earlier quoted context omitted.

Google is allowed to be big, be better and win users. But happy customers is not the full test of monopolization. The real question is, "Could a meaningfully better search engine realistically displace Google today?” If the answer is no, then competition is broken

> "Could a meaningfully better search engine realistically displace Google today?” ChatGPT clearly demonstrated that displacing Google is possible. All previous monopoly arguments seemed even more flimsy after that.

I think you’re proving the monopoly argument yourself: if they only way to compete with Google is an innovation that generations of scientists have been working towards, it does paint a grim picture of competition in this space. Besides, are we ignoring Gemini?

Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi

#79
post #24

> Building a comparable one from scratch is like building a parallel national railroad.. Not too be pedantic here but I do have a noob question or two here: 1. One is building the index, which is a lot harder without a google offering its own API to boot. If other tech companies really wanted to break this monopoly, why can't they just do it — like they did with LLM training for base models with the infamous "pile" d…

I don’t think it’s comparable to today’s AI race. Google has a monopoly, an entrenched customer base, and stable revenue from a proven business model. Anyone trying to compete would have to pour massive money into infrastructure and then fight Google for users. In that game, Google already won. The current AI landscape is different. Multiple players are competing in an emerging field with an uncertain business model.…

> ... stable revenue from a proven business mode... In that game, Google already won.

But if they were to pour that money strategically to capture market share one of two things would happen if google was replaced/lost share:

1. it would be the start of the commoditization of search. i.e. search engine/index would become a commodity and more specialized and people could buy what they want and compete.

2. A new large tech company takes rein. In which case it would be as bad as this time.

Like what I don't get is that if other big tech companies actually broke apart monopoly on search, several google dominos in mobile devices, browser tech, location capabilities would fall. It would be a massive injection of new competition into the economy, lots of people would spend more dollars across the space(and ad driven buying too) money would not accrue in an offshore tax haven in ireland

To play the devils advocate, I think the only reason its not happening is because meta, apple, microsoft have very different moats/business models to profit off. They all have been stung one time or another is small or big ways for trying to build something that could compete but failed. MS with bing, Meta with facebook search, Foursquare — not big tech but still — with Maurauder's Map.

Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi

#80
post #73

Earlier quoted context omitted.

Scraping is hard. Very good scraping is even harder. And today, being a scraping business is veeery difficult; there are some "open"/public indices, but none of these other indices ever took off

Well sure yes, I don't contend with the fact that its hard, but if the top tech companies joined their heads I am sure if for example, Meta, Apple, MS have enough talent between to make an open source index if only to reap gains from the de-monopolization of it all.

I mean, doesn't microsoft have bing?
Post reply on HN