Live data from Hacker News

A New Search Engine

0x65.dev

51–60 of 107 posts

Re: A New Search Engine

#51
post #47
post #43

Earlier quoted context omitted.

[Disclaimer: I work at Cliqz] I hate to answer this one, becasue it looks too much marketing-speech but this feature exists. Not on beta.cliqz.com but on the drop-down search on Cliqz browser. Based on the tabs you have opened, different query expansions are selected. For instance as you type "hotel in ma..." probably would show you results for Mallorca, but if you have "Madrid" on a tab, then it will show results fo…

Ok, but can't you then figure out what the browser was displaying using Javascript?

They could, but then everyone would see them do it, and kind of the whole point is that they won't do it.

Re: A New Search Engine

#52
post #47

Earlier quoted context omitted.

Ok, but can't you then figure out what the browser was displaying using Javascript?

They could, but then everyone would see them do it, and kind of the whole point is that they won't do it.

But who in their right mind would allow the Javascript code of one tab to access data in other tabs?

Re: A New Search Engine

#53
post #47
post #43

Earlier quoted context omitted.

[Disclaimer: I work at Cliqz] I hate to answer this one, becasue it looks too much marketing-speech but this feature exists. Not on beta.cliqz.com but on the drop-down search on Cliqz browser. Based on the tabs you have opened, different query expansions are selected. For instance as you type "hotel in ma..." probably would show you results for Mallorca, but if you have "Madrid" on a tab, then it will show results fo…

Ok, but can't you then figure out what the browser was displaying using Javascript?

Not sure I get your point. But contextual search only works for the search within Cliqz browser, on the address bar dropdown, on the client space. The same approach cannot be done on the (web-page SERP page, beta.cliqz.com), because from a web we have no access to the tabs opened. It could only possible via tracking and user-profiling, which is something that we do not do, or want to do.

Re: A New Search Engine

#54

Cliqz nearly made me stop using firefox a while back https://www.heise.de/-3852129 (german article) Tldr 1% of german firefox installations automatically uploaded search queries to cliquz. I wont trust a search engine like this with any of my data.

Would you trust google who does the same thing with chrome? Bing which does the same thing with IE (or whatever it's called now)? Blame firefox for selling out their users not the search engine. I didn't close my account with amazon when Ubuntu started sending searches to them, I just switched distros.

I use mainly duckduckgo and only use google if duckduckgo doesn't bring up anything useable (which is far too often for me tbh). And yes I blame firefox for every mishap over the last few years like the certificate expiration, the mr robot "advertising", the cloudflare dns and so on. But I see the good things as well like trowing out avast. So I trust them more than google.

But still I will never trust something like cliqz which belongs to the media gmbh which produces Schund like die bunte.

Ps I tested the beta and the search results werent good Pps I will read the search engine articles thought

Re: A New Search Engine

#55

What I really want, is not another search engine for contextless queries. Except for really basic queries (which Google/etc already do a good job at), I'm trying to answer a question, perhaps open-ended, and it will take multiple queries to resolve. And it's not a linear process of narrowing down with + or - keywords. It's establishing a context: I'm searching for something relevant to "go" the language, not "go" the…

I had very similar idea a few years ago, did some quick numbers and came to conclusion that it was not going to fly commercially.

Interestingly the working name for my idea was also "ResearchEngine", so i guess it summarizes pretty well the unmet need you and me have.

Re: A New Search Engine

#56

   > Why the second constraint? one might 
   > ask. Besides the obvious potential for 
   > profitability, our mission was
The search engine the world needs is one with independence and non-profitability. If the creators are preoccupied with turning a profit, they’ll introduce the same garbage features as Google. It’s a shame, because a good search engine could shorten the time humanity has to wait for advances (eg: cures for cancers, cheaper energy, etc)

Re: A New Search Engine

#57
post #26

> The experts, who chose to answer, suggested that we should first start with crawling the whole web. We were told that this would take between 1 and 2 years to complete, and would cost a minimum of $1 billion Why are costs so high for crawling?

(Disclaimer: I work at Clizq) I don't work on the search, but did some work recently on the crawling part. What I know is that crawling is far more difficult if you are not a big player. Sites will quickly block you once you hit a rate limit.

We have to be very careful, since when we get blocked there is normally no way to get unblocked again. You can try to send them an email to unblock you, but it is unlikely that you get a response. This is one part of the explanation why crawling is slow. The other part is more obvious: the internet is large.

The blocking part is hard to overcome as a small player, while for Google it is the opposite as sites simply cannot afford being exclude from the index. If we would not have to care about rate limits, it would simplify the problem.

Re: A New Search Engine

#58

> Why the second constraint? one might > ask. Besides the obvious potential for > profitability, our mission was The search engine the world needs is one with independence and non-profitability . If the creators are preoccupied with turning a profit, they’ll introduce the same garbage features as Google. It’s a shame, because a good search engine could shorten the time humanity has to wait for advances (eg: cures for…

Indexing the web is a resource intensive activity tho - if it was federated then the resource cost only increases. I suppose a non-profit is the alternative, but non-profits are not exactly independent unless they have some sort of massive endowment. I'm not trying to disagree with you, it's just a paradoxical problem: to resolve the issue, resources must be accumulated. Accumulating resources means it's hard to resolve the issue (of an independent search engine).

Re: A New Search Engine

#59

Earlier quoted context omitted.

Would you trust google who does the same thing with chrome? Bing which does the same thing with IE (or whatever it's called now)? Blame firefox for selling out their users not the search engine. I didn't close my account with amazon when Ubuntu started sending searches to them, I just switched distros.

There's a difference. When I use a service from Google, I expect that my data will be parsed by Google. And I can decide if I trust Google or not. But Firefox sending the urls I visit to a third party (Cliqz) silently and without permission is shady and deceptive. And then, after all this, Cliqz claims that it's a company built on privacy... sheesh.

An interesting quote for their article: " Philosophically, we believe copying is a loaded term, we prefer to use the term learning. Learning from each other is something all of us do"

What is the difference between copying and learning?

Re: A New Search Engine

#60

Cliqz nearly made me stop using firefox a while back https://www.heise.de/-3852129 (german article) Tldr 1% of german firefox installations automatically uploaded search queries to cliquz. I wont trust a search engine like this with any of my data.

(Disclaimer: I work at Cliqz) Just read the article. It is from 2017 and very short. For the non-German speakers, I have to translate the relevant part:

> Rund ein Prozent der Firefox-Downloads enthalten künftig das Add-On Cliqz, das bereits beim Eintippen Vorschläge für Webseiten anzeigt. Dafür wertet es die Surf-Aktivitäten aller Nutzer aus.

About 1% of the Firefox downloads will contain the Cliqz Addon, which will show you search suggestions for websites while you type. For that, it uses browsing activities of all users.

---

The last "of all users" is important. Yes, our search is built on data collected from users, but the point is we cannot build profiles of single users; we are only seeing what the whole group of users does. I cannot stress that part enough. We are not Avast.

In fact, we are very open about our data collection system called Human Web:

* https://0x65.dev/blog/2019-12-03/human-web-collecting-data-i...

And this article explains how we provide anonymity while sending:

* https://0x65.dev/blog/2019-12-04/human-web-proxy-network-hpn...

I can understand that you did not like the way that Mozilla rolled it out in 2017. I'm also not glad about how it went (my personal opinion). But from the technical side, I'm more than happy to take any question on that topic (how we collect data in Cliqz).

Post reply on HN