Live data from Hacker News

Waiting for dawn in search: Search index, Google rulings and impact on Kagi

blog.kagi.com

241–250 of 266 posts

Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi

#241

I think one side problem is that part of the web is not even searchable with a search engine. Here are some examples: - Discord - WeChat (is it the web?) - Rednote - TikTok (partially) - X (partially) - JSTOR (it finds daily, but you find more stuff on the website directly) - any stuff with a login, obviously.

> Discord Damn, I can't stand open-source projects that host their "forums" on Discord. It's a nigthmare to use, it's heavy, slow, and it's completely unsearchable from the web. I wonder what went wrong with our society.

> I wonder what went wrong with our society.

Predators.

https://maggieappleton.com/cozy-web

Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi

#242
post #77

Earlier quoted context omitted.

I've been tossing around the very early idea of seeing what we can do to elevate alcoves of the web such as Gemini[1] through Kagi. I am slightly conscious of that some people might not like us operating in that space, it's been on my TODO to poll people about it and take a quick pulse. I love the tech and think we could give it meaningful exposure. Is this along the lines of what you have in mind - any other active…

How would that work? Like Kagi caches the gemini content and delivers it as web content? I suppose that might annoy the kind of person who runs a Gemini server.

I don't think we would go to that end, not as a first step anyways; I'm thinking of some simpler ones. Sparing details as I'm just brainstorming for now and getting to know their communities.

But, there are already plenty of services that proxy Gemini pages so that you can read them in conventional browsers, as well as search engines for Gemini content.

Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi

#243
post #77

Earlier quoted context omitted.

I've been tossing around the very early idea of seeing what we can do to elevate alcoves of the web such as Gemini[1] through Kagi. I am slightly conscious of that some people might not like us operating in that space, it's been on my TODO to poll people about it and take a quick pulse. I love the tech and think we could give it meaningful exposure. Is this along the lines of what you have in mind - any other active…

That's cool that you're looking into it. Are you saying that in any "official" manner as a Kagi employee? Or something more personal? I've been meaning to write an RFC or open-letter of sorts to collect ideas for what a neo or parallel web could look like, but I'm just a nobody so shrug . It'll probably be something very fragmented and very very niche but nowadays I think that can be seen as a good thing.

I'm working on making an internal proposal to integrate with Gemini on several fronts, yes. Still hatching the idea, and much else to do - maybe this summer it will come to fruition if it pans out :)

Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi

#244
post #117

Google's advantage is not just in its index and algorithms, it is that it has built a self-reinforcing flywheel that data mines human attention at massive scale to improve their search results. This comment ( https://news.ycombinator.com/item?id=46709957 ) points out that Google got its start via PageRank, which essentially ranked sites based on links created by humans . As such, its primary heuristic was what humans…

> cookies (for which they built Gmail!) Can you explain this one?

There were blogs that explained this in detail (Facebook does something similar), but I can't find them, so here's what Google's AI overview says when I search for "How gmail cookies help google track users across the web":

Gmail cookies, such as SID and HSID, act as unique identifiers for a signed-in Google account, allowing Google to track user activity across its services and millions of third-party websites. These cookies, often lasting 2 years, link browsing behavior—like searches and site visits—to a specific user profile to personalize ads, measure campaign performance, and analyze site usage, even on non-Google sites that use tools like Google Analytics or AdSense.

Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi

#245
I think it's worth mentioning the Open Web Search initiative [1] and the Open Web Index [2] specifically.

> 14 renowned European research and computing centers have joined forces to develop an open European infrastructure for web search. The initiative is contributing to Europe’s digital sovereignty as well as promoting an open human-centered search engine market. [1]

> The Open Web Index (OWI) is a European open source web index pilot that is currently in Beta testing phase. The idea: Collaboratively and transparently secure safe, sovereign and open access to the internet for European organisations and civil society. The index stores well structured open web data, making it available for search applications and LLMs. [3]

[1] https://openwebsearch.eu/

[2] https://openwebindex.eu/

[3] https://openwebsearch.eu/open-webindex/

Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi

#246
post #243

Earlier quoted context omitted.

That's cool that you're looking into it. Are you saying that in any "official" manner as a Kagi employee? Or something more personal? I've been meaning to write an RFC or open-letter of sorts to collect ideas for what a neo or parallel web could look like, but I'm just a nobody so shrug . It'll probably be something very fragmented and very very niche but nowadays I think that can be seen as a good thing.

I'm working on making an internal proposal to integrate with Gemini on several fronts, yes. Still hatching the idea, and much else to do - maybe this summer it will come to fruition if it pans out :)

Well, from one dreamer to another, thank you for taking on that effort and I wish you luck. I will keep my eye out for it.

Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi

#247
post #8

Does anyone else use the phrase "I'm going to google XYZ" while referring to actually searching it up on Kagi, DDG, or another search engine?

Nope. I stopped using google for search many years ago, and stopped using it as a verb about the same time.

I admit I've used something like "are you banned on Google or what?" a couple of times though.

Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi

#248

Earlier quoted context omitted.

> knock-off Is it though? It feels so better than Google results[1], while being still built partly with Google results. In the last 3 years as a Kagi customer i have rarely if ever felt the need to use bangs !g and on few occasions i did use them, it was with instant regret. In the previous decade or so using DDG, using bangs !g Google would be 30-50% of searches, i would have to consciously try the results first in…

DDG's results are primarily Bing results while Kagi's results are primarily Google results. It makes sense that you feel the need to escape from Bing to Google more often than from Google to Google.

Perhaps, but each time I do go to Google the results are painfully bad. I don't think Kagi is just proxying Google- while they do use it as a core source, they rerank much better

The blog post talks about that specifically - Bing unlike Google does have Index licensing program but their terms forbid reordering that is key reason Kagi is not also using Bing in their index mix.

My point is Kagi is similar to a car tuning company like Hennessey, Brabus they take a base product and make it much better for a premium, they are not selling knock-offs.

Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi

#249

Earlier quoted context omitted.

I don't know dude. Every time I've assumed good faith on my paying for something means ad free, I've been screwed by some asshole with an MBA getting into a leadership position high enough to push ads through. I'd rather it be explicit

I get the point, we've all been burnt. But if you're not trusting anybody anyway, why would explicity in a non-binding blog post / press release soothe you? We're in the middle of an AI bubble propping up the whole friggin US economy all by itself, driven mostly by a company that claimed to be a non-profit until a few years ago.

Because I've been burned by every big tech company I can think of. As for why it would be soothing, well because it gives me hope that when I read any further legal docs they'll hold to the post.

What does ai have to do with this? The sooner that bubble bursts the better IMO.

Re: Waiting for dawn in search: Search index, Google rulings and impact on Kagi

#250

Earlier quoted context omitted.

Scraping is hard, and is not hard that much at the same time. There are many projects about scraping, so with a few lines you can do implement scraper using curl cffi, or playwright. People complain that user-agent need to be filled. Boo-hoo, are we on hacker news, or what? Can't we just provide cookies, and user-agent? Not a big deal, right? I myself have implemented a simple solution that is able to go through many…

+1 so much for this. I have been doing the same, an SQLite database of my "own personal internet" of the sites I actually need. I use it as a tiny supplementary index for a metasearch engine I built for myself - which I actually did to replace Kagi. Building a metasearch engine is not hard to do (especially with AI now). It's so liberating when you control the ranking algorithm, and can supplement what the big engine…

Do you have any documentation/blog post for this? I would love to do something similar for my own use.
Post reply on HN