Building a Search Engine from Scratch
51–60 of 151 posts
Re: Building a Search Engine from Scratch
#52Cliqz still hasn't given a good answer as to why they use what's seemingly a typo in 'analysis' to get around user tracker blockers. https://anolysis.privacy.cliqz.com/ As you can see in this subthread, they claim it's "anonymized" analytics, https://news.ycombinator.com/item?id=21718694 which, as 99% of research suggests, isn't anonymous at all: https://www.fastcompany.com/90278465/sorry-your-data-can-sti...
Re: Building a Search Engine from Scratch
#53I'd never heard of Cliqz before, but just did a couple of test searches, and I'm honestly super impressed with the results. I found the result relevancy seemed to be closer to Google than DDG/Bing
Re: Building a Search Engine from Scratch
#54I don’t see how this will ultimately be successful. They’re basically reverse engineering Google by looking at user logs. Google will always have a leg up here because they have all the Google data. And even if it does work for a while, there still needs to be the original signal to copy. Someone will have to crawl the web and index content. I’m super eager to find new approaches to search, but another Google clone i…
Further, they don't have to be better than google in the quality of their search results. As soon as the results have good relevance, it's good competition.
What matters is that they are 'good enough', that there is a hint of competition to the Google-Bing monopoly. Offering privacy centric 'competition' is what they say their main aim is and what their success should be judged upon.
Re: Building a Search Engine from Scratch
#55Re: Building a Search Engine from Scratch
#56Cliqz still hasn't given a good answer as to why they use what's seemingly a typo in 'analysis' to get around user tracker blockers. https://anolysis.privacy.cliqz.com/ As you can see in this subthread, they claim it's "anonymized" analytics, https://news.ycombinator.com/item?id=21718694 which, as 99% of research suggests, isn't anonymous at all: https://www.fastcompany.com/90278465/sorry-your-data-can-sti...
[Disclaimer: I work at Cliqz] Hi, I read the thread and thought the answer was good enough, but it seems that you are not yet convinced. Let me try: 1) Here there is a list of publications regarding privacy by Cliqz (including published scientific papers). It should have fairly easy to find it using a search engine :-) https://0x65.dev/pages/dissemination-cliqz.html Hopefully, the paper will convince you that Cliqz p…
Re: Building a Search Engine from Scratch
#57I'd never heard of Cliqz before, but just did a couple of test searches, and I'm honestly super impressed with the results. I found the result relevancy seemed to be closer to Google than DDG/Bing
Maybe so, but honestly, the name is putting me off more than anything. It just doesn't sound professional, and causes me to perceive it as shady, even if it's not. The name also makes it sound like it's more about marketing "clicks" to advertisers than providing good results. None of that is necessarily true, but it's the impression the name gives. It needs to re-brand.
Re: Building a Search Engine from Scratch
#58Earlier quoted context omitted.
Indeed, using the (query -> result) mapping in their 'Human Web' logs, from other search engines, is a roundabout way of copying both the simple word-to-doc-index, and the complicated ranking logic, of other search engines. A sufficiently strict intellectual-property regime might find this a copyright violation. But without any inside info, I strongly suspect Google & other incumbents already do similarly-indirect mo…
[Disclaimer I work for Cliqz] Copying, learning would be a bit more precise, and I'm not kidding. We do not answer queries 1 to 1, query-logs are used to build a more concise model of the page. It's not just a cache, we started like this, but we quickly learn to answer unseen queries. A sufficient strict IP might find text snipped a copyright violation too. It's a trick area. Personally, I'm at peace as we get the co…
Interesting you bring this up. Wasn't the company funding cliqz in favor of text snippets being copyright violations when the big G does it? (German/EU Leistungsschutzrecht) [0]
I mean, I think bootstrapping from google query logs is fine, but following the money it does seem like a double standard. Cliqz "stealing" from Google SERPs is fine, but Google news "stealing" text snippets isn't?
[0]: https://leistungsschutzrecht.info/stimmen-zum-lsr/pressearti...
(Relevant quote: Das Leistungsschutzrecht halte man nach wie vor nicht für falsch. Man setze sogar weiterhin auf ein euop. Leistungsschutzrecht, mit dessen Hilfe man sich erhofft, endlich Geld von Google zu erhalten.
sloppy translation: We still consider the [we-want-money-for-google-news-snippets-law] to not be the wrong approach. We in fact contiue hoping it will work out on a european scale)
Re: Building a Search Engine from Scratch
#59Earlier quoted context omitted.
Maybe so, but honestly, the name is putting me off more than anything. It just doesn't sound professional, and causes me to perceive it as shady, even if it's not. The name also makes it sound like it's more about marketing "clicks" to advertisers than providing good results. None of that is necessarily true, but it's the impression the name gives. It needs to re-brand.
Yeah, the name "Google" definitely sounds more professional....not.
Re: Building a Search Engine from Scratch
#60Earlier quoted context omitted.
Yeah, the name "Google" definitely sounds more professional....not.
It does though. There’s a ton of services using a deliberate misspelling in their name and all of them are meh at best in my experience.