It is even worse that the Google search become shit in last years. So they gate keep only relevant information for themselves and not using them with intent to improve search quality. As always if you have no competition your innovation goes only towards cost reduction. Not product improvement.
Waiting for dawn in search: Search index, Google rulings and impact on Kagi
61–70 of 266 posts
Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi
#62> Building a comparable one from scratch is like building a parallel national railroad.. Not too be pedantic here but I do have a noob question or two here: 1. One is building the index, which is a lot harder without a google offering its own API to boot. If other tech companies really wanted to break this monopoly, why can't they just do it — like they did with LLM training for base models with the infamous "pile" d…
Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi
#63> Building a comparable one from scratch is like building a parallel national railroad.. Not too be pedantic here but I do have a noob question or two here: 1. One is building the index, which is a lot harder without a google offering its own API to boot. If other tech companies really wanted to break this monopoly, why can't they just do it — like they did with LLM training for base models with the infamous "pile" d…
Microsoft had a chance (well another chance, after they gave up IE's lead) to break up Google's browser monopoly, but they decided to use Chromium for free instead.
Ultimately all these decisions come down to what's more profitable, not what's in the best interests of the public. We have learned this lesson x1000000. Stop relying on corporations to uphold freedoms (software or otherwise), becuase that simply isn't going to happen.
Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi
#64It is even worse that the Google search become shit in last years. So they gate keep only relevant information for themselves and not using them with intent to improve search quality. As always if you have no competition your innovation goes only towards cost reduction. Not product improvement.
If Google Search is shit, why does Kagi want access to it?
Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi
#65It may be impracticable to share the crawled data, but from the stand point of content providers, having a single entity collecting the information (rather than a bunch of people doing) would seem to be better for everyone. Likely need to have some form of robots.txt which would allow the content provider to indicate how their content could be used (i.e research, web search, AI, etc.).
The people accessing the crawled data would end up paying (reasonable) fees to access the level of data they want, and some portion of that fee would go to the content provider (30% to the crawler and 70% to the crawler? :P maybe).
Maybe even go so far as to allow the Paywalled content providers to set a price on accessing their data for the different purposes. Should they be allowed to pick and choose who within those types should be allowed (or have it be based on violations of the terms of access)
It seems in part the content providers have the following complaints:
* Too many crawlers (see note below re crawlers)
* Crawlers not being friendly
* Improper use of the crawled data
* Not getting compensated for their content
Why not the index? The index, to me, is where a bunch of the "magic" happens and where individual companies could differentiate themselves from everyone else.Why can't Microsoft retain Bing traffic when it's the default on stock Windows installs?
* Do they not have enough crawled data?
* Their index isn't very good?
* Their searching their index isn't good
* The way they present the data is bad?
* Google is too entrenched?
* Combination of the above?
There are several entities intending to crawl all / large portions of the Internet: Baidu, Bing, Brave, Google, DuckDuckGo, Gigablast, Mojeek, Sogou and Yandex [1]. That does not include any of the smaller entities, research projects, etc.[1] https://en.wikipedia.org/wiki/Search_engine#2000s–present:_P... (2019)
Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi
#66> Building a comparable one from scratch is like building a parallel national railroad.. Not too be pedantic here but I do have a noob question or two here: 1. One is building the index, which is a lot harder without a google offering its own API to boot. If other tech companies really wanted to break this monopoly, why can't they just do it — like they did with LLM training for base models with the infamous "pile" d…
A huge amount of the web is only crawlable with a googlebot user-agent and specific source IPs.
Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi
#67Does anyone else use the phrase "I'm going to google XYZ" while referring to actually searching it up on Kagi, DDG, or another search engine?
Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi
#68I hope they cache search results to further reduce the number of calls to Google. And Marginalia Search was not mentioned? Marginalia Search says they are licensing their index to Kagi. Perhaps it's counted under "Our own small-web index" which is highly misleading if true.
> "Our own small-web index" Has Kagi ever said what this is? I wouldn't be at all surprised if it is just kagi.com pages or a download of Wikipedia.
Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi
#69> Building a comparable one from scratch is like building a parallel national railroad.. Not too be pedantic here but I do have a noob question or two here: 1. One is building the index, which is a lot harder without a google offering its own API to boot. If other tech companies really wanted to break this monopoly, why can't they just do it — like they did with LLM training for base models with the infamous "pile" d…
Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi
#70The statistics in this article sound like garbage to me. Google used by 90% or the world? ~20% of the human population lives in countries where Google is blocked. OTOH, Baidu is the #1 search engine in China, which has over 15% of the world’s population… but doesn’t reach 1%? These stats are made measuring US-based traffic, rather than “worldwide” as they claim.
The Search Engine wikipedia article [1] has a section on Russia and East Asia market share, which confirms that the roll up used for world wide counts is off, unless the number of people using the Internet is drastically different in some of the countries.
Russia
* Yandex: 70.7%
* Google: 23.3%
China: * Baidu: 59.3%
* Other domestic engines: "smaller shares"
* Bing: 13.6%
South Korea: * Naver: 59.8%
* Google: 35.4%
Japan:
* Google: 76.2%
* Yahoo! Japan: 15.8%[1] https://en.wikipedia.org/wiki/Search_engine#Market_share