> Because direct licensing isn’t available to us on compatible terms, we - like many others - use third-party API providers for SERP-style results Crazy for a company to admit: "Google won't let us whitelabel their core product so we steal it and resell it."
Waiting for dawn in search: Search index, Google rulings and impact on Kagi
91–100 of 266 posts
Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi
#92Earlier quoted context omitted.
Well sure yes, I don't contend with the fact that its hard, but if the top tech companies joined their heads I am sure if for example, Meta, Apple, MS have enough talent between to make an open source index if only to reap gains from the de-monopolization of it all.
I mean, doesn't microsoft have bing?
Search indexes are hard, surely, but if you were to strip it to just a good index on the browser, made it free, kept it fresh, it cannot be 100 billion dollars to build. Then you use this DoJ decision and fight against google to not deny a free index to have equal rights on chrome you can have a massive shot at a win for a LOT less money.
Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi
#93Earlier quoted context omitted.
I mean, doesn't microsoft have bing?
Yeah but no one uses it. I am not even sure people that are forced to use it like using it because it was productized it pretty poorly. After all who wants another google? They invested 100 Billion dollars, which is a lot of wasted money TBH. Search indexes are hard, surely, but if you were to strip it to just a good index on the browser, made it free, kept it fresh, it cannot be 100 billion dollars to build. Then yo…
I mean... Duckduckgo uses bing api iirc and I use duckduckgo and many people use duckduckgo.
I also used bing once because bing used to cache websites which weren't available in wayback archive, I don't know how but It was pretty cool solution for a problem.
I hate bing too and I am kind of interested in ecosia/qwant's future as well (yes there's kagi too and good luck to kagi as well! but I am currently still staying on duckduckgo)
Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi
#94Earlier quoted context omitted.
A huge amount of the web is only crawlable with a googlebot user-agent and specific source IPs.
I do not know a lot about this subject, but couldn’t you make a pretty decent index off of common crawl? It seems to me the bar is so low you wouldn’t have to have everything. Especially if your goal was not monetization with ads.
Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi
#95> Building a comparable one from scratch is like building a parallel national railroad.. Not too be pedantic here but I do have a noob question or two here: 1. One is building the index, which is a lot harder without a google offering its own API to boot. If other tech companies really wanted to break this monopoly, why can't they just do it — like they did with LLM training for base models with the infamous "pile" d…
> 1. One is building the index, which is a lot harder without a google offering its own API to boot. If other tech companies really wanted to break this monopoly, why can't they just do it? FTA: > Context matters: Google built its index by crawling the open web before robots.txt was a widespread norm, often over publishers’ objections. Today, publishers “consent” to Google’s crawling because the alternative - being i…
> The robots.txt played a role in the 1999 legal case of eBay v. Bidder's Edge,[12] where eBay attempted to block a bot that did not comply with robots.txt, and in May 2000 a court ordered the company operating the bot to stop crawling eBay's servers using any automatic means, by legal injunction on the basis of trespassing.[13][14][12] Bidder's Edge appealed the ruling, but agreed in March 2001 to drop the appeal, pay an undisclosed amount to eBay, and stop accessing eBay's auction information.[15][16]
Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi
#96What even is market rate? Kagi themselves admits there's no market, the one competitor quit providing the service.
Obviously Google doesn't want to become an index provider.
Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi
#97With Google's search engine making almost $200 billion a year in revenue, I'm not sure Kagi could afford what market rates would be here. They also spent billions developing the technology to crawl, index, and rank billions of pages, factoring that in, again I don't think a good price can be put on it. What even is market rate? Kagi themselves admits there's no market, the one competitor quit providing the service. O…
> Google must provide Web Search Index data (URLs, crawl metadata, spam scores) at marginal cost.
I'm guessing that the "marginal cost" of a search is small and it's not connected to the how much ad revenue that search is worth.
Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi
#98Earlier quoted context omitted.
A huge amount of the web is only crawlable with a googlebot user-agent and specific source IPs.
> And given you-know-what, the battle to establish a new search crawler will be harder than ever. Crawlers are now presumed guilty of scraping for AI services until proven innocent. I have always wondered but how does wayback machine work, is there no way that we can use wayback archive and then run a index on top of every wayback archive somehow?
1. IIUC depends a lot on "Save Page Now" democratization, which could work, but its not like a crawler.
2. In absence of alexa they depend quite heavily on common crawl, which is quite crazy because there literally is no other place to go. I don't think they can use google's syndicated API, cause they would then start showing ads in their database, which is garbage that would strain their tiny storage budget.
3. Minor from a software engineering perspective but important for survival of the company: since they are an artifact of record storage, to convert that to an index would need a good legal team to battle google to argue. They do that the DoJ's recent ruling in their favor.
Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi
#99> Building a comparable one from scratch is like building a parallel national railroad.. Not too be pedantic here but I do have a noob question or two here: 1. One is building the index, which is a lot harder without a google offering its own API to boot. If other tech companies really wanted to break this monopoly, why can't they just do it — like they did with LLM training for base models with the infamous "pile" d…
> If other tech companies really wanted to break this monopoly, why can't they just do it Google is a verb, nobody can compete with that level of mindshare.
Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi
#100Earlier quoted context omitted.
Yeah but no one uses it. I am not even sure people that are forced to use it like using it because it was productized it pretty poorly. After all who wants another google? They invested 100 Billion dollars, which is a lot of wasted money TBH. Search indexes are hard, surely, but if you were to strip it to just a good index on the browser, made it free, kept it fresh, it cannot be 100 billion dollars to build. Then yo…
> Yeah but no one uses it. I am not even sure people like using it because it was productized it pretty poorly. They invested 100 Billion dollars, which is a lot of wasted money TBH. I mean... Duckduckgo uses bing api iirc and I use duckduckgo and many people use duckduckgo. I also used bing once because bing used to cache websites which weren't available in wayback archive, I don't know how but It was pretty cool so…
The small distributed team grinding it out against the goliath. They are awesome and perhaps the right example of what a path like this would look like. Maybe someone from their team can chime in on the difficulties of building a search engine that works in the face of tremendous odds.