Earlier quoted context omitted.
I’m my second year into Kagi and loving it. I actually just upgraded to get Kagi Assistant (basically, cloud access to every LLM out there). But the search alone is worth every penny, and it’s built/operated fully in-house as far as I know. https://kagi.com
I loved kagi, but their cost is prohibitive for me. The only way there is a chance for me to afford Kagi might be to buy "search credit" without a subscription and without minimum consumption. And then it would only be good if they allowed more than 1000 domain rules and showed more results (when available)
Two upstart search engines are teaming up to take on Google
201–210 of 299 posts
Re: Two upstart search engines are teaming up to take on Google
#202Every new search engine I've seen was a Bing wrapper with sometimes light reranking. I understand that competing with Google was borderline impossible a decade ago. But in 2024, we have cheap compute, great OSS distributed DBs, powerful new vector search tech. Amateur search engines like Marginalia even run on consumer hardware. CommonCrawl text-only is ~100TB, and can fit on my home server. Why is no company buildin…
Because nowdays more than ever content you need is in silos. Your facebooks/twiters/instagram/stack overflow/reddit ... And they all have limited expensive api's, and have bulk scrapping detection. Sure you can clobber together something that will work for a while, but you can't runn a buissness on that. Aditionaly most paywalled sites (like news) explicitly whitlist google and bing, and if someone cretes new site, t…
If the site is less aggressively blocking but only has a per-IP rate limit, buy a subscription to one of those VPNs (it doesn't matter if they're "actually secure" or not - you can borrow their IP addresses either way). If the site is extremely aggressive, you can outsource to the slightly grey market for residential proxy services - for fifty cents to several dollars per gigabyte, so make sure that fits in your business plan.
There's an upper bound to a website's aggressiveness in blocking, before they lose all their users, which tops out below how aggressive you can be in buying a whole bunch of SIM cards, pointing a directional antenna at McDonald's, or staying a night at every hotel in the area to learn their wi-fi passwords.
Re: Two upstart search engines are teaming up to take on Google
#203I'm beginning to suspect LLMs' viability for search purposes is dependent on your existing search habits. 40% of my search queries are just copy-pasted error messages. Another 10% are business names, for the sole purpose of finding their hours or phone number. Less than 10% are complete clauses or sentences. I just don't see how an LLM fits my search habits. I tried using ChatGPT for search purposes, and it was dread…
You should use the right tool for the right purpose. If you're searching for a physical place and want opening hours, reviews and photos, it's Google Maps. If you're looking for updates and photos for a physical place, it's Instagram. If it's detailed articles it's Wikipedia. If it's to define or translate a word, right click and "Look up". Or if it's general search, you have Kagi.
Re: Two upstart search engines are teaming up to take on Google
#204Earlier quoted context omitted.
This is one of those things that I think is interesting about how "normal" people use the Internet. I am guessing they just always start with google. But for me, if I want to look up local restaurants, I go straight to Maps/Yelp/FourSquare(RIP). If I want to look up releases of a band, I go straight to musicbrainz. Info about Metal Band, straight to the Encyclopeadia Metallum. History/Facts, straight to Wikipedia. Re…
"Normal" people don't start with Google, not any more. They start with Facebook, Instagram, X, Reddit, Discord, Substack etc. That's exactly the problem, the world-wide web has devolved back into a collection of walled gardens like things were in the BBS era, except now the boards are all run by a handful of Silicon Valley billionaires instead of random nerds in your hometown.
Re: Two upstart search engines are teaming up to take on Google
#205Earlier quoted context omitted.
It's funny. I usually can't tell that from the quality of the search results.
Compared to what?
Re: Two upstart search engines are teaming up to take on Google
#206Every new search engine I've seen was a Bing wrapper with sometimes light reranking. I understand that competing with Google was borderline impossible a decade ago. But in 2024, we have cheap compute, great OSS distributed DBs, powerful new vector search tech. Amateur search engines like Marginalia even run on consumer hardware. CommonCrawl text-only is ~100TB, and can fit on my home server. Why is no company buildin…
Because most of the good sites won't let you crawl them any more, unless you're Google.
Re: Two upstart search engines are teaming up to take on Google
#207Earlier quoted context omitted.
You should use the right tool for the right purpose. If you're searching for a physical place and want opening hours, reviews and photos, it's Google Maps. If you're looking for updates and photos for a physical place, it's Instagram. If it's detailed articles it's Wikipedia. If it's to define or translate a word, right click and "Look up". Or if it's general search, you have Kagi.
My purpose is Search, and I want one right tool for it. I have had one right tool for my entire life - first Google, and then StartPage since ~2017. I considered Kagi. I like the concept, and I'm willing to pay for search, but Kagi costs more than I'm willing to pay.
Re: Two upstart search engines are teaming up to take on Google
#208I'm beginning to suspect LLMs' viability for search purposes is dependent on your existing search habits. 40% of my search queries are just copy-pasted error messages. Another 10% are business names, for the sole purpose of finding their hours or phone number. Less than 10% are complete clauses or sentences. I just don't see how an LLM fits my search habits. I tried using ChatGPT for search purposes, and it was dread…
Good luck wading through whatever Google gives back for that.
Also I find they do well on error messages in general.
Re: Two upstart search engines are teaming up to take on Google
#209Every new search engine I've seen was a Bing wrapper with sometimes light reranking. I understand that competing with Google was borderline impossible a decade ago. But in 2024, we have cheap compute, great OSS distributed DBs, powerful new vector search tech. Amateur search engines like Marginalia even run on consumer hardware. CommonCrawl text-only is ~100TB, and can fit on my home server. Why is no company buildin…
To properly compete with Google search you have to: 1. Create your own index/crawl and true search engine on top of it, rather than delegating that back to bing or Google. In doing so, solve or address all the related problems like SEO/spam, and cover the “long tail” (expensive part) 2. Monetize it somehow to afford all of the personnel/hardware/software costs. Probably with your own version of Google’s ad products.…
Google's dominance will never be upended by an incremental improvement.
Re: Two upstart search engines are teaming up to take on Google
#210Earlier quoted context omitted.
To properly compete with Google search you have to: 1. Create your own index/crawl and true search engine on top of it, rather than delegating that back to bing or Google. In doing so, solve or address all the related problems like SEO/spam, and cover the “long tail” (expensive part) 2. Monetize it somehow to afford all of the personnel/hardware/software costs. Probably with your own version of Google’s ad products.…
I've heard more or less the same thing said about IBM, Walmart, Microsoft, etc. There is one way to compete with Google search. Google search is a general search engine. But what if you only wanted medical information? information about cars? information about physics? electronics? history? etc. Specializing makes for a much smaller scope of the problem, and being specialized means it could deliver more useful result…
Some of them charge a lot for access but they certainly exist.