Live data from Hacker News

A New Search Engine

0x65.dev

81–90 of 107 posts

Re: A New Search Engine

#81
post #6

They talk about using query logs to optimize their search results: >Queries performed by people, if associated to a web page, serve as even cleaner summaries than anchor text. This is because all the logic put in place by the search engine, who resolved the query with a list of web pages, and all human understanding and experience that led one to select the best page from the offered result list end up embedded in th…

Wouldn't a multi-armed bandit help alleviate this issue? (Basically, randomly display a few other links, and use bayesian stats to figure out if the new links are more optimal).

That, or any kind of exploration/optimization algorithm, to be honest.

Re: A New Search Engine

#82
post #74
post #18

Earlier quoted context omitted.

[Disclaimer, I work at Cliqz] Your point is spot on. Old pages tend to have more association to seen queries, which does not play in favor for new pages. That said, however, there are a couple of things to consider: 1) seen queries is not the only way to create queries, we are pretty good creating synthetic queries based on the content, descriptions, etc. This queries are more noisy that the seen queries of course, b…

Why does Cliqz use an analytics domain with a typo in it to get around user tracker-blockers? That's incredibly scummy, given how much Cliqz has been shouting about privacy. https://anolysis.privacy.cliqz.com/

Anolysis stands for Ano[nymized] [Ana]lysis. It's a new approach to do telemetry without sending unique identifiers (like most analytics / telemetry) systems do - but focus on goal attainment at a group level. This makes work harder of course, but it's a price we've been willing to pay. It is a pity you would take a domain name as evidence of malice. We should have a paper coming up at some point on the approach.

Re: A New Search Engine

#83
post #82
post #74

Earlier quoted context omitted.

Why does Cliqz use an analytics domain with a typo in it to get around user tracker-blockers? That's incredibly scummy, given how much Cliqz has been shouting about privacy. https://anolysis.privacy.cliqz.com/

Anolysis stands for Ano[nymized] [Ana]lysis. It's a new approach to do telemetry without sending unique identifiers (like most analytics / telemetry) systems do - but focus on goal attainment at a group level. This makes work harder of course, but it's a price we've been willing to pay. It is a pity you would take a domain name as evidence of malice. We should have a paper coming up at some point on the approach.

Self-claimed anonymous analytics have repeatedly failed in the past: you've been using this for at least two years, why didn't Cliqz release a paper before then on it? Or even explain what it is? Or a mention on your site of it? I don't want to dislike Cliqz, what it says it's doing is cool. However, given the financial incentives involved, the verify step of "trust but verify" is more essential than ever.

Re: A New Search Engine

#84
post #15

I don't know how many of you tried the engine, but there are 2 features that instantly took my attention: 1) Trackers Stats. Essentially, you can see how many and what trackers there are on the page you are about to visit. Before visiting it. 2) Page previews (I'm not sure about whether I like that)

> 1) Trackers Stats. This feature is powered by another project we run, where we measure the tracking landscape in the web (most popular domains): https://whotracks.me . Details on how that works can be found in our paper [0]. Also - we are flirting with the idea of providing a mode where the ranking is informed by the trackers in the destination site. Would love to hear your thoughts on whether you'd like smth like…

> we are flirting with the idea of providing a mode where the ranking is informed by the trackers in the destination site. Would love to hear your thoughts on whether you'd like smth like this.

That's definitely the right way to go. I would also very much appreciate an option not to show in the result list sites/pages having any trackers.

Re: A New Search Engine

#85
post #76

Earlier quoted context omitted.

Seems like an excellent opportunity to apply Hanlon's Razor. "Never attribute to malice that which can be adequately explained by stupidity."

They've been doing it for two years at minimum: https://news.ycombinator.com/item?id=15423936 Two years and three domains with the typo? It's malice.

This malice metric is confusing. How malacious is Google then? They misspelled "googol" for over two decades, but there's only one domain. Do we count the 1e100.net as a misspelling?

Re: A New Search Engine

#86
post #26

> The experts, who chose to answer, suggested that we should first start with crawling the whole web. We were told that this would take between 1 and 2 years to complete, and would cost a minimum of $1 billion Why are costs so high for crawling?

Wouldn’t common crawl content be enough? If not what are the issues?

Re: A New Search Engine

#87
There is a dark side to this story. With Burda https://en.wikipedia.org/wiki/Hubert_Burda_Media, the same people who are behind the Cliqz search engine were originally also behind the German Leistungsschutzrecht. https://en.wikipedia.org/wiki/Ancillary_copyright_for_press_... This law, heavily lobbied for by publishers, forces every search engine and everybody else using content from the internet to pay a private tax of 6% of the revenue (not from profit!). https://www.vg-media.de/de/digitale-verlegerische-angebote/f... As the profit of most internet companies is below this margin, it is essentially forcing many companies out of business.

This tax is enforced and collected by VG Media, the German collecting society representing rights of a group of German publishers. https://www.vg-media.de Between 2013 and 2016 Burda was a shareholder of VG Media, which was commissioned to enforce the tax in its name.

The evil thing of this law is, that the publishers are not required to mark their content in machine-readable form as paid content. And a manual selection is infeasible for internet-scale with billions of pages. So a search engine has no means to bypass the paid content and indexing only free content, e.g. like Wikipedia which makes the majority of the internet content. Essentially the "Leistungsschutzrecht" takes the free content hostage to extort money for using the internet, even if you don't use paid content of the publishers (the just 200 publications the VG Media represents).

So while Burda's Cliqz write on their blog "The world needs more search engines" https://www.0x65.dev/blog/2019-12-01/the-world-needs-cliqz-t... they supported a law that made it impossible for many search engines to operate in Germany (and in the EU via the similar EU law "Extra copyright for news sites" (“Link tax”) https://juliareda.eu/eu-copyright-reform/extra-copyright-for... And while today they are not anymore shareholder of the VG Media, they still benefit from the suppressive legal environment they helped to create, as it prevents any new independent competition to enter the search market

Re: A New Search Engine

#88
post #15

Earlier quoted context omitted.

> 1) Trackers Stats. This feature is powered by another project we run, where we measure the tracking landscape in the web (most popular domains): https://whotracks.me . Details on how that works can be found in our paper [0]. Also - we are flirting with the idea of providing a mode where the ranking is informed by the trackers in the destination site. Would love to hear your thoughts on whether you'd like smth like…

> we are flirting with the idea of providing a mode where the ranking is informed by the trackers in the destination site. Would love to hear your thoughts on whether you'd like smth like this. That's definitely the right way to go. I would also very much appreciate an option not to show in the result list sites/pages having any trackers.

I'm afraid this will remove any results from page :-D

Re: A New Search Engine

#89
post #80

Since the Cliqz devs are here, and this engine is based in Germany, a question: does your search engine have any mechanisms for reporting abusive URLs (doxxing, targeted harassment, revenge porn, etc) beyond right-to-be-forgotten, or are you more a lassiez-faire, everything-goes kind of search company? I noticed that your engine ranks some of the nastier sites on the internet far higher than any other search engine I…

[Disclaimer: I work at Cliqz] Yes, there is a way to report such urls https://cliqz.com/en/report-url

We do have a list of blacklisted urls/domains mostly regarding adult topic (child porno etc). If you have noticed some bad sites in our results, please feel free to drop a line to our support team using link I provided

Re: A New Search Engine

#90
post #88

Earlier quoted context omitted.

> we are flirting with the idea of providing a mode where the ranking is informed by the trackers in the destination site. Would love to hear your thoughts on whether you'd like smth like this. That's definitely the right way to go. I would also very much appreciate an option not to show in the result list sites/pages having any trackers.

I'm afraid this will remove any results from page :-D

Is it really that bad? I surely hope it's not.

I noticed that even Wikipedia is reported as having some trackers. But when I looked closer I noticed that most of those belong to the Wikimedia foundation, which is fine. I mean, I don't mind site owners tracking what I do on their site, I just don't want to be followed across the whole Web.

The rest of the Wikipedia trackers are supposed to be Google fonts and statics, but I couldn't witness any calls to those. Maybe the stats are not quite up to date?

If such a score is to be given, it better be fair and reflecting the current state of affairs.

Post reply on HN