Live data from Hacker News

A New Search Engine

0x65.dev

101–107 of 107 posts

Re: A New Search Engine

#101

There is a dark side to this story. With Burda https://en.wikipedia.org/wiki/Hubert_Burda_Media , the same people who are behind the Cliqz search engine were originally also behind the German Leistungsschutzrecht. https://en.wikipedia.org/wiki/Ancillary_copyright_for_press_... This law, heavily lobbied for by publishers, forces every search engine and everybody else using content from the internet to pay a private ta…

[Disclaimer: I work at Cliqz]

Sorry for taking so long to reply, I was personally trying to dig some information about this. An additional disclaimer: not a lawyer either.

Honestly, I have little idea of how this law affects search engines. What I can say is that we are no paying anything, as AFAIK we do not know anyone who is. Moreover, if some publisher would complain, even one in Burda, we would stop crawling by domain, there is no technical issue here, properties are known by the imprint. We have no say on what the investors do but I can assure you that we have no pressure. For instance, our ad-blocker works everywhere, regardless if the sites are from Burda or not.

On a general level, assuming that what you say is factually correct, I must personally agree that regulation is a bitch. It's typically designed fro big companies to control other big companies, but small ones get negatively affected if only because of the lack of resources. We recently had to suffer all the overhead of GDPR, which consumed a fair amount of our time, relatively we paid a higher price that Google.

Personally, I cannot respond for all the decisions made by the people funding Cliqz, I do not even think I can judge it either. They might be complaining and lobbying, no idea. But they are also putting good money to build a privacy-preserving search engine and a browser, something that no-one else is doing, so on my account they are on the positive side.

Re: A New Search Engine

#102
post #74
post #18

Earlier quoted context omitted.

[Disclaimer, I work at Cliqz] Your point is spot on. Old pages tend to have more association to seen queries, which does not play in favor for new pages. That said, however, there are a couple of things to consider: 1) seen queries is not the only way to create queries, we are pretty good creating synthetic queries based on the content, descriptions, etc. This queries are more noisy that the seen queries of course, b…

Why does Cliqz use an analytics domain with a typo in it to get around user tracker-blockers? That's incredibly scummy, given how much Cliqz has been shouting about privacy. https://anolysis.privacy.cliqz.com/

How does using an odd spelling thwart user-tracker blockers?

Re: A New Search Engine

#103

The nationalism on the homepage is a little odd... in particular since they're still essentially building on top of Google. https://cliqz.com/en/ > Europe has failed to build its own digital infrastructure. US companies such as Google have thus been able to secure supremacy. They ruthlessly exploit our data. They skim off all profits. They impose their rules on us. In order not to become completely dependent and end…

> The nationalism on the homepage is a little odd.

Especially seeing as how Europe is not a nation.

> A future in which we all have sovereign control over our data and our digital lives. ... And this is exactly why we at Cliqz are developing digital key technologies made in Germany.

Ahhh, that makes more sense, "sovereign" == "made in Germany."

Re: A New Search Engine

#104
post #102
post #74

Earlier quoted context omitted.

Why does Cliqz use an analytics domain with a typo in it to get around user tracker-blockers? That's incredibly scummy, given how much Cliqz has been shouting about privacy. https://anolysis.privacy.cliqz.com/

How does using an odd spelling thwart user-tracker blockers?

Many ad blockers have baseline filters that block subdomains, URLs, etc. with common tracking terms in them (ex. "telemetry" or "tracking").

Re: A New Search Engine

#105
post #89
post #80

Since the Cliqz devs are here, and this engine is based in Germany, a question: does your search engine have any mechanisms for reporting abusive URLs (doxxing, targeted harassment, revenge porn, etc) beyond right-to-be-forgotten, or are you more a lassiez-faire, everything-goes kind of search company? I noticed that your engine ranks some of the nastier sites on the internet far higher than any other search engine I…

[Disclaimer: I work at Cliqz] Yes, there is a way to report such urls https://cliqz.com/en/report-url We do have a list of blacklisted urls/domains mostly regarding adult topic (child porno etc). If you have noticed some bad sites in our results, please feel free to drop a line to our support team using link I provided

Thanks for the reply. A bit disappointed it only counts for extremely illegal content. There's a lot of really negative stuff out there that is blatantly false and manipulative (Ripoff Report, Tumblr callout posts, etc) and it's always a shame that this kind of negative toxicity gets promoted so high in SERPs.

I'd really like it if there were an ethical SERP that at least had some integrity with its results. Reporting factual unflattering statements is one thing (and ideal), but promoting libel feels really dirty, and so far Cliqz seems to be the worst at that of any search engines I've used, and your reporting link seems as though Cliqz is okay with that.

Re: A New Search Engine

#106
How many algorithms are there in chrome alone? I remember when people realized that they could game Facebook shares for higher rankings on chrome and for a while buzzfeed top ten lists outranked Wikipedia every fucking time. I guess that’s still going on. What a clusterfuck search results are nowadays.

If anyone builds anything, please make it so algorithms or queries are archived. I hate how I can’t find anything on the internet that I searched for and found years ago. Its like the history of the internet evaporates every year. I don’t even know if some websites still exist or if I simply can’t find them because rankings are terrible.

I’m to the point that I haven’t been on a new website in years. How do you find new websites in this day and age when the same websites are ranked at the top every time?

Re: A New Search Engine

#107
post #95

Weird, your domain (cliqz.com) was blocked by my pihole.

(Disclaimer: I work at Cliqz) We had problems with being blocked in the past.

In cases, where we got a chance to explain, they agree that it is a false positive and took us off the block list. At least, that happened so far in all cases that I'm aware of. However, there are so many lists that it is hard to keep track of them. Would be nice if you could provide some information which block list it is, so we can contact them.

The reason why we end on the blocklist is normally a misconception of our data collection system Human Web: https://0x65.dev/blog/2019-12-03/human-web-collecting-data-i...

If someone does not want to send Human Web data, the feature can also be disabled through the UI. Same if you browse in a private window; Human Web is automatically disabled there. There is no need to configure blocking rules.

Post reply on HN