Live data from Hacker News

Waiting for dawn in search: Search index, Google rulings and impact on Kagi

blog.kagi.com

151–160 of 266 posts

Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi

#151
post #52

Earlier quoted context omitted.

Strange to pick on Kagi when there's much bigger companies on that list.

Those companies allegedly have used SerpAPI (probably to check visibility), but not to resell a Google Search knock-off.

> knock-off

Is it though? It feels so better than Google results[1], while being still built partly with Google results.

In the last 3 years as a Kagi customer i have rarely if ever felt the need to use bangs !g and on few occasions i did use them, it was with instant regret.

In the previous decade or so using DDG, using bangs !g Google would be 30-50% of searches, i would have to consciously try the results first instead of starting with !g and then think to myself DDG was at least getting the query data to improve their results.

[1] While the de-cluttered UI is a relief, on just the results list comparison, Google search is so bad that less time saved in not redrafting the queries constantly, filtering out the spam, the AI summaries, sponsored content, all the "cards" , recommended search listicles on is worth more than the $10/month.

Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi

#152

Earlier quoted context omitted.

If a crawler offered enough money they could be allowed too. It's not like Google has exclusive crawling rights.

There is a logistics problem here - even if you had enough money to pay, how would you get in touch with every single site to even let them know you're happy to pay? It's not like site operators routinely scan their error logs to see your failed crawling attempts and your offer in the user-agent. Even if they see it, it's a classic chicken & egg problem: it's not worth the time of the site operator to engage with you…

Realistically you don't need every single site on board before you index becomes valuable. You can get in touch with sites via social media, email, discord, or even visiting them face to face.

Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi

#153
post #23

Earlier quoted context omitted.

Seems like an open question as to whether that violates any laws. Another way to look at it is that if you publish a service on the web, you have limited rights to restrict what people do with it. Isn't that the logic Google search relies on in the first place? I didn't give permission for Google to crawl and index and deep link to my site (let alone summarize and train LLMs on it). They just did it anyway, because i…

Google's stance is "I can copy you and you can't stop me" as well as "You can't copy me, I'll sue you"

Maybe it has changed but Google doesn't look like it uses litigation as its primary weapon. It defends itself but rarely attacks.

The are however more than happy to use technical measures, like blocking accounts. And because of their position, blocking your Google account may be more damaging than a successful lawsuit.

Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi

#154
post #22

The statistics in this article sound like garbage to me. Google used by 90% or the world? ~20% of the human population lives in countries where Google is blocked. OTOH, Baidu is the #1 search engine in China, which has over 15% of the world’s population… but doesn’t reach 1%? These stats are made measuring US-based traffic, rather than “worldwide” as they claim.

Google is only blocked in places where it would already be hard for a company with morals to work in, if not outright blocked as well. This probably represents traffic globally, excluding those places. Instead of downvoting blindly, please state which countries are currently blocking Google that would willingly allow Kagi, a AI/Privacy focused search engine company to exist in their domain? The results may surprise y…

[deleted]

Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi

#155

Earlier quoted context omitted.

Scraping is hard. Very good scraping is even harder. And today, being a scraping business is veeery difficult; there are some "open"/public indices, but none of these other indices ever took off

Scraping is hard, and is not hard that much at the same time. There are many projects about scraping, so with a few lines you can do implement scraper using curl cffi, or playwright. People complain that user-agent need to be filled. Boo-hoo, are we on hacker news, or what? Can't we just provide cookies, and user-agent? Not a big deal, right? I myself have implemented a simple solution that is able to go through many…

> Search engines affect me less, and less every day. I have my own small "index" / "bookmarks" with many domains, github projects, youtube channels

Exactly, why can't we just hoard our bookmarks and a list of curated sources, say 1M or 10M small search stubs, and have a LLM direct the scraping operation?

The idea is to have starting points for a scraper, such as blogs, awesome lists, specialized search engines, news sites, docs, etc. On a given query the model only needs a few starting points to find fresh information. Hosting a few GB of compact search stubs could go a long way towards search independence.

This could mean replacing Google. You can even go fully local with local LLM + code sandbox + search stub index + scraper.

Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi

#156
post #145

Earlier quoted context omitted.

Not only was eBay v. Bidder's Edge technically after Google existed, not before, more critically the slippery-slope interpretation of California trespass to chattels law the District Court relied on in it was considered and rejected by the California Supreme Court in Intel v. Hamidi (2003), and similar logic applied to other states trespass to chattels laws have been rejected by other courts since; eBay v. Bidder's E…

The point is, robots.txt was definitely a thing that people expected to be respected before and during google's early existence. This Kagi claim seems to be at least partially false: > Google built its index by crawling the open web before robots.txt was a widespread norm, often over publishers’ objections.

Perhaps it wasn't a widespread norm though. But I don't really see why that matters as much, is the the issue that sites with robots.txt today only allow Googlebot and not other search engines? Or is Google somehow benefitting from having two decade old content that is now blocked because of robots.txt that the website operators don't want indexed?

Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi

#157
Honestly, would be very cool if someone could make a search engine of only human-produced content. I know it's going to be hard and compute intensive but I don't think it's impossible. In fact, Google could do it. A paid service for only human made content. Obviously there would be a margin of error as we can never be 100% sure if something really is AI written.

Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi

#158
post #62
post #24

> Building a comparable one from scratch is like building a parallel national railroad.. Not too be pedantic here but I do have a noob question or two here: 1. One is building the index, which is a lot harder without a google offering its own API to boot. If other tech companies really wanted to break this monopoly, why can't they just do it — like they did with LLM training for base models with the infamous "pile" d…

A huge amount of the web is only crawlable with a googlebot user-agent and specific source IPs.

Are these websites not serving public content? If there's some legal concerns just create a separate scraping LLC that fakes user agent and uses residential IPs or VPN or something. I can't imagine that the companies would follow through with some sort of lawsuit against a scraper that's trying to index their site to get them more visitors, if they allow GoogleBot.

Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi

#159
post #73

Earlier quoted context omitted.

Well sure yes, I don't contend with the fact that its hard, but if the top tech companies joined their heads I am sure if for example, Meta, Apple, MS have enough talent between to make an open source index if only to reap gains from the de-monopolization of it all.

All these companies have the exact same business model as Google (advertising) and have the same mismatched incentives: good search results are not something they want. Google Search sucks not because Google is incapable of filtering out spam and SEO slop (though they very much love that people believe they can't), but that spam/slop makes the ads on the SERP page more enticing, and some of the spam itself includes G…

I was on the Goog forums for years (before they even fucking ruined the FORMAT of the forums, possibly to 'be more mobile friendly') and it was people absolutely (justifiably) screaming at the product people

No, the customer isn't 'always' right, but these guys like to get big and once big, fuck you, we don't have to listen to you, we're big; what are you going to do, leave?

Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi

#160
post #117

Google's advantage is not just in its index and algorithms, it is that it has built a self-reinforcing flywheel that data mines human attention at massive scale to improve their search results. This comment ( https://news.ycombinator.com/item?id=46709957 ) points out that Google got its start via PageRank, which essentially ranked sites based on links created by humans . As such, its primary heuristic was what humans…

> "learn" which pages the users actually found useful for the given query But due to their business model I'm not sure they are ranking "usefulness" as much as you think. Useful results ultimately don't benefit Google because Google makes no money on them. Google makes money on ads - either ads on the search results page, ads on the destination pages or (indirectly) from steering users to pages which have Google Anal…

That is also the thesis of this piece: https://www.wheresyoured.at/the-men-who-killed-google/

It is plausible, but I'd guess Google would not risk that. I'm sure Google has pulled other shenanigans to get more clicks, like stuffing more and more ads, and making ads look like results (something even I personally have fallen for once), but I think they're too smart to mess with their sacred cash cow.

Post reply on HN