Live data from Hacker News

Building a Search Engine from Scratch

0x65.dev

31–40 of 151 posts

Re: Building a Search Engine from Scratch

#31
post #26
post #15

Earlier quoted context omitted.

[Disclaimer: I work at Cliqz] Hi, I read the thread and thought the answer was good enough, but it seems that you are not yet convinced. Let me try: 1) Here there is a list of publications regarding privacy by Cliqz (including published scientific papers). It should have fairly easy to find it using a search engine :-) https://0x65.dev/pages/dissemination-cliqz.html Hopefully, the paper will convince you that Cliqz p…

I'm one of those people who remains skeptic about the anonymity of the tracking Cliqz does in general. Obviously people have a hard time believing any company that is in the advertising space and preaches about privacy, mainly because they have been burned several times before. For me it was the data proxying through FoxyProxy that made me uncomfortable. I have also remained unconvinced about the motivation for not u…

Well, we released our beta search for Tor yesterday: search4tor7txuze.onion/ (works obviously only in the Tor browser). That’s as good as it gets regarding making it technically impossible, isn’t it? More complex obviously for a browser - but the answer can simply not be „only no data at all is good“, because that’s a destructive approach that only favors the worst privacy intruders; no one would then be able to build up competition. We work hard to be as transparent as possible about what we do. Show me any other company that builds a big data product (like search) that is so transparent about „yes we collect data, but it’s non personal - here is how we do it, please scrutinize us“. Are we perfect? Hell no! But we try and we go a long way to be challenged to improve. If we were shady - would we be naked in front of you, showing each step we take? We would not even interact with the tech folks (especially not on hacker news, where people really know what they talk about), but scam people who know less (at least that sounds like a more reasonable strategy to me if I would want to fool people - which we don’t).

(EDIT/Disclaimer: I obviously work for Cliqz)

Re: Building a Search Engine from Scratch

#32
post #15
post #14

Cliqz still hasn't given a good answer as to why they use what's seemingly a typo in 'analysis' to get around user tracker blockers. https://anolysis.privacy.cliqz.com/ As you can see in this subthread, they claim it's "anonymized" analytics, https://news.ycombinator.com/item?id=21718694 which, as 99% of research suggests, isn't anonymous at all: https://www.fastcompany.com/90278465/sorry-your-data-can-sti...

[Disclaimer: I work at Cliqz] Hi, I read the thread and thought the answer was good enough, but it seems that you are not yet convinced. Let me try: 1) Here there is a list of publications regarding privacy by Cliqz (including published scientific papers). It should have fairly easy to find it using a search engine :-) https://0x65.dev/pages/dissemination-cliqz.html Hopefully, the paper will convince you that Cliqz p…

You say you're copying Google and you use that to promote your product, it's very hard to believe you are privacy friendly.

You're also owned by a media company, that makes it even harder to believe that you're going to respect users privacy.

Add to that the tone of arrogance of articles such as "the world needs Cliqz", you can see why it's a no-no.

Every few days there is a post on HN trying so hard to convince the readers that Cliqz is the best, even though the articles read between the lines, that the Cliqz team does not have the capability to make its own search algo or make slightly more complicated queries.

I am experienced enough to know where this is coming from: managers that do not know what they are doing and engineers drunk from glory that do not see their own mistakes.

Please Cliqz hire a search engine expert. Hire great engineers, they are going to cost twice, but you're going to get a search engine that actually works.

Please, or the HN community will have to bash you every time you post an article.

Re: Building a Search Engine from Scratch

#33

Earlier quoted context omitted.

The article describes techniques used by all search engines, not just Google. Search engines have existed before Google, and despite Google's monopoly on search, "Google clone" is a poor term to describe all search engines when alternatives with unique features (e.g. DuckDuckGo) exist.

They are literally rebuilding Google by tracking how users use Google and rebuilding the SERPs. If that’s not a Google clone I don’t know what is.

All search engines have search-engine results pages.

They're looking at how users use Google Search because the data's there. They're making a competitor to Google Search. That doesn't mean they're rebuilding Google Search's SERPs, or making a Google Search “clone”; I've got results from Cliqz for queries I'm confident have never been put into Google before, meaning it's functioning as an independent search engine.

Re: Building a Search Engine from Scratch

#34
post #8

I don’t see how this will ultimately be successful. They’re basically reverse engineering Google by looking at user logs. Google will always have a leg up here because they have all the Google data. And even if it does work for a while, there still needs to be the original signal to copy. Someone will have to crawl the web and index content. I’m super eager to find new approaches to search, but another Google clone i…

And even if it does work for a while, there still needs to be the original signal to copy

I'd say they need to start somewhere. Using other search engine's results is a ways to get things started, so that they can build their own index on crawled content later.

For now, it's great to see another competitor for Google coming up.

Re: Building a Search Engine from Scratch

#35
post #8

I don’t see how this will ultimately be successful. They’re basically reverse engineering Google by looking at user logs. Google will always have a leg up here because they have all the Google data. And even if it does work for a while, there still needs to be the original signal to copy. Someone will have to crawl the web and index content. I’m super eager to find new approaches to search, but another Google clone i…

Google search worked well enough for me when it was launched 22 years ago. So the patents to that early version should be expired. I'd be happy to use a competitor's service if they would reintroduce a Search API, and PageRank, just give me a programmatic interface that doesn't throw up Captcha after a few searches. I'd pay a reasonable price for the compute time + profit margin.

Re: Building a Search Engine from Scratch

#36
post #26

Earlier quoted context omitted.

I'm one of those people who remains skeptic about the anonymity of the tracking Cliqz does in general. Obviously people have a hard time believing any company that is in the advertising space and preaches about privacy, mainly because they have been burned several times before. For me it was the data proxying through FoxyProxy that made me uncomfortable. I have also remained unconvinced about the motivation for not u…

Well, we released our beta search for Tor yesterday: search4tor7txuze.onion/ (works obviously only in the Tor browser). That’s as good as it gets regarding making it technically impossible, isn’t it? More complex obviously for a browser - but the answer can simply not be „only no data at all is good“, because that’s a destructive approach that only favors the worst privacy intruders; no one would then be able to buil…

We appreciate your openness about how you collect data, but that's still not enough because literally every other advertising company is deceptive when they talk about privacy.

The openness must be paired with privacy that is guaranteed under all circumstances, and the most common way to achieve that would be to route the anonymized data through Tor.

Your search engine being also available on Tor has nothing to do with data collection by Cliqz on other sites. Your search engine website is not the avenue through which the Cliqz browser and extension collects data as you browse the web, I'm confused why would you even bring it up.

Re: Building a Search Engine from Scratch

#37
post #8

I don’t see how this will ultimately be successful. They’re basically reverse engineering Google by looking at user logs. Google will always have a leg up here because they have all the Google data. And even if it does work for a while, there still needs to be the original signal to copy. Someone will have to crawl the web and index content. I’m super eager to find new approaches to search, but another Google clone i…

Google search worked well enough for me when it was launched 22 years ago. So the patents to that early version should be expired. I'd be happy to use a competitor's service if they would reintroduce a Search API, and PageRank, just give me a programmatic interface that doesn't throw up Captcha after a few searches. I'd pay a reasonable price for the compute time + profit margin.

If Google went back to the original algorithms it would return even worse results. Much of the noise can be attributed to two factors: (1) the explosive growth of the Internet itself, which was much much smaller and had more focused and authoritative content 20 years ago; and (2) aggressive SEO tactics today that would easily fool the early versions of PageRank.

Re: Building a Search Engine from Scratch

#38
post #29

That Cliqz is trying to actually build a new search stack is commendable. This is way more exciting to me than DuckDuckGo and other services that just package up Bing search results under different branding. I'm skeptical that they'll be successful, but I wish them the best. They should market (and engineer) strongly on privacy since that's where Google is weak.

> They should market (and engineer) strongly on privacy since that's where Google is weak.

How can you build a privacy oriented search engine and still make money?

Re: Building a Search Engine from Scratch

#39

Earlier quoted context omitted.

Google search worked well enough for me when it was launched 22 years ago. So the patents to that early version should be expired. I'd be happy to use a competitor's service if they would reintroduce a Search API, and PageRank, just give me a programmatic interface that doesn't throw up Captcha after a few searches. I'd pay a reasonable price for the compute time + profit margin.

If Google went back to the original algorithms it would return even worse results. Much of the noise can be attributed to two factors: (1) the explosive growth of the Internet itself, which was much much smaller and had more focused and authoritative content 20 years ago; and (2) aggressive SEO tactics today that would easily fool the early versions of PageRank.

I kind of miss the early directories for sites.

It probably can't happen now, since there are billions of websites, but it was a simpler time, and finding something you needed wasn't THAT hard.

Re: Building a Search Engine from Scratch

#40
post #38
post #29

That Cliqz is trying to actually build a new search stack is commendable. This is way more exciting to me than DuckDuckGo and other services that just package up Bing search results under different branding. I'm skeptical that they'll be successful, but I wish them the best. They should market (and engineer) strongly on privacy since that's where Google is weak.

> They should market (and engineer) strongly on privacy since that's where Google is weak. How can you build a privacy oriented search engine and still make money?

By selling ads solely against the query, and figuring out a way to track conversions against generated unique referring URL instead of using cookies. Maybe it's less profitable, but it might still be profitable, and it gives you a foothold.
Post reply on HN