Live data from Hacker News

Reporting on tech companies means finding people who don’t want to be found

nytimes.com

21–30 of 45 posts

Re: Reporting on tech companies means finding people who don’t want to be found

#21
post #4

Earlier quoted context omitted.

I’ve searched for web archival services before, but never found anything outside of some expensive legal services. It would be great if there was a service like the waybackmachine but which could be use to archive on demand, and make those archives publicly available. The internet archive is great for this, but they also respect robots.txt and don’t guarantee that they’ll archive pages you request.

I know blockchain is a dirty word around here, but this is absolutely an application that can use something like Bitcoin to store hashes of a site on a specific date, giving you the closest thing you can get to "proof" that the site existed on that date with that specific content. Then you don't need "secure" storage of the actual archives, you can have as many copies as possible with as many people as possible and n…

https://github.com/ramyhardan/proof-of-existence/blob/master...

Re: Reporting on tech companies means finding people who don’t want to be found

#22

Earlier quoted context omitted.

You're totally correct. But if you just want to document something for your own future review, screenshots work fine. The thing is, journalism sees itself as being above citations. Journalists are primarily relying on their reputations as journalists, not evidence; the word of a journalist is supposed to be evidence, valid as a citation elsewhere. Journalists don't even record all of their interviews, a lot of them j…

A usecase for blockchain.

https://github.com/ramyhardan/proof-of-existence/blob/master...

Re: Reporting on tech companies means finding people who don’t want to be found

#23

I'm surprised that despite being able to concisely explain why Apple is preferable to Google due to the privacy implications of their respective business models, they're perfectly content with a Google Home in their kitchen! Maybe I'm paranoid, but isn't it obvious how the whole Home Speaker story ends? "We're not spying on you, we're learning your behaviors to offer a better experience!"

Something like...

>We're not spying on you, we're learning your behaviors to offer a better experience!

Okay I guess that makes sense, I like when stuff is designed well around what I actually want from it, this will help in that, cool!

>We're not spying on you, we're just going to report you to the police if our ML algorithm detects a child screaming for their life

Wow okay, that kinda seems like crossing a line, but I guess it makes sense and is going to be good to protect kids

>We're not spying on you, we're just going to lower your reputation on x/y/z service if we hear racial slurs

Uh, I get where that's coming from, but this is clearly spying. Well, a private company can do what they like I guess, and its understandable, who other than racists wouldn't want less racists online.

>We're not spying on you, we're just going to let your insurance company know if we hear you coughing

Woah okay what the fuck

>We're not spying on you, we're just going to block you from x/y/z service if your political views differ from our CEO and put out a warning to services provided by other companies.

Hey what the - I mean, wow what a great update. goes in seperate private room WHAT THE FUCK I NEED TO GET RID OF THIS

>We're spying on you, we're just going to arrest you if you will not comply. Don't worry about what law you're charged under, we just blackmailed the police and those who tried to protest got arrested, our company policy is your local law now.

What a great update, thanks awesome [super evil company name here] developers and all employees, what great work they do! internally oh god oh god oh god oh god oh god

Of course this goes way to the extreme, but that's my point, this is the kind of seed that makes that stuff possible. Doesn't mean its absolutely evil, but the potential is there.

Re: Reporting on tech companies means finding people who don’t want to be found

#24
post #4

> I’m also a prolific screenshotter. The internet is an ephemeral place, so when I see something online for a story, I make sure to capture it immediately. Why are screenshots seen as the gold standard for documenting deleted content? They are easily faked.

I’ve searched for web archival services before, but never found anything outside of some expensive legal services. It would be great if there was a service like the waybackmachine but which could be use to archive on demand, and make those archives publicly available. The internet archive is great for this, but they also respect robots.txt and don’t guarantee that they’ll archive pages you request.

archive.org actually has a form for "archive it now", right there on https://archive.org/web/. They also address using their data in court proceedings in the FAQ. So I think they do have that use case in mind.

Of course they'd be stupid to offer any guarantees without payment. That's what their service https://archive-it.org is for. Only for organisations, unfortunately.

In regard to respecting robots.txt: I would consider any organisation running a bot and not respecting robots.txt as somewhat shady. ignoring it could even have consequences for admissibility in court.

Re: Reporting on tech companies means finding people who don’t want to be found

#25
post #4

> I’m also a prolific screenshotter. The internet is an ephemeral place, so when I see something online for a story, I make sure to capture it immediately. Why are screenshots seen as the gold standard for documenting deleted content? They are easily faked.

I’ve searched for web archival services before, but never found anything outside of some expensive legal services. It would be great if there was a service like the waybackmachine but which could be use to archive on demand, and make those archives publicly available. The internet archive is great for this, but they also respect robots.txt and don’t guarantee that they’ll archive pages you request.

There are at least 3 locations for on-demand web archiving:

https://web.archive.org/save/https://news.ycombinator.com/ https://archive.is/ https://webrecorder.io/

Probably some of the other sites supported by this extension have on-demand archiving too:

https://github.com/dessant/view-page-archive

For archiving on your own computer, this LWN article is pretty interesting:

https://lwn.net/Articles/766374/

Re: Reporting on tech companies means finding people who don’t want to be found

#26

Earlier quoted context omitted.

Rational people are also emotional, and apathy is the default emotion. Recognizing that many products are privacy invasive becomes the same as realizing that they are produced in ways that harm the environment, or hurt workers. A fact acknowledged but accepted in exchange for the benefits they offer.

If you're equating giving up personal data to harming the environment or workers, I'm afraid I don't agree. One of these actions is under your control, is a trade off we make simply to live in society, and even provides some benefits to users and society. To bank, to get mail, to shop online, to support services we like, to receive communications, etc, etc. Privacy is also something you can control, rather than somet…

The point is that these are all examples of negative choices that consumers choose to take because it’s easier to accept short-term personal gain while causing long-term, depersonalized, harm.

Re: Reporting on tech companies means finding people who don’t want to be found

#27

Earlier quoted context omitted.

You're totally correct. But if you just want to document something for your own future review, screenshots work fine. The thing is, journalism sees itself as being above citations. Journalists are primarily relying on their reputations as journalists, not evidence; the word of a journalist is supposed to be evidence, valid as a citation elsewhere. Journalists don't even record all of their interviews, a lot of them j…

What level of "evidence" would you accept? You reference science publishing as the gold standard, but why not require the standards of the criminal justice system? Indeed, why does science get away with a lesser standard? Each professions has the standard it deems optimal within the trade-offs of risks, costs, etc. Those obviously differ: holding journalism to the standards of a criminal trial would essentially make…

18 years ago? Are you thinking of Jayson Blair? https://en.wikipedia.org/wiki/Jayson_Blair#Plagiarism_and_fa... That broke 15 years ago, not 18. Also, it's a highly relevant example, from the same news outlet as TFA. That's exactly the sort of thing we need better citations for; it would have been way harder to fake all those interviews if he had to upload recordings of them. Harder still if we expected dash-cam footage of him driving around West Virginia. He described video footage that didn't exist -- even if we can justifiably question whether a video was faked or not, releasing the actual video would certainly make it harder to fake than a written description is. Of course absolute evidence is hard, if not impossible; Bertrand Russel spent over 300 pages trying to prove that 1+1=2, and ultimately concluded that he failed. We have to accept that our knowledge is probabilistic, but that doesn't mean we can't try to improve our probabilities of being right.

Edit: I removed the mobile part of the URL. Other edit: s/spend/spent

Re: Reporting on tech companies means finding people who don’t want to be found

#28
post #23

I'm surprised that despite being able to concisely explain why Apple is preferable to Google due to the privacy implications of their respective business models, they're perfectly content with a Google Home in their kitchen! Maybe I'm paranoid, but isn't it obvious how the whole Home Speaker story ends? "We're not spying on you, we're learning your behaviors to offer a better experience!"

Something like... >We're not spying on you, we're learning your behaviors to offer a better experience! Okay I guess that makes sense, I like when stuff is designed well around what I actually want from it, this will help in that, cool! >We're not spying on you, we're just going to report you to the police if our ML algorithm detects a child screaming for their life Wow okay, that kinda seems like crossing a line, bu…

I'm sorry Dave, I'm afraid I can't do that

Re: Reporting on tech companies means finding people who don’t want to be found

#29

Does anybody else think this article is at least 90% native advertising?

Hey, a conspiracy theory. Great start for a new account!

What you are misunderstanding: Companies and products are part of what's often called "the real world". As such, journalists are allowed to mention them in articles, because nothing in the real world cannot be the subject of journalism.

And in the same way that they sometimes offer positive opinions on politicians, or 200 year old books, or tomorrow's weather, it can happen that a story reflects positively on a product or company.

Accusing the NYT of corruption may seem like just a platitude to you, and spreading it as nothing of consequence. Because you're one of the smart people that see through all the propaganda that the mainstream media is brainwashing us with.

To a journalist it's kinda like accusing a physician of murder. It also gnaws at a rather important foundation of democracy. That's why you should come with some sort of rational for your theory.

As a lesser point: it would be rather stupid if the NYT could be bought by a bunch of startups. Because every single one of the people involved with these transactions would have the power to ruin the Times. It's absurd to believe that they would regularly engage in such practices and manage to keep it secret in the current political climate. It's also absurd that they would risk their existence for whatever meagre sums a few mentions in an article could be worth.

Re: Reporting on tech companies means finding people who don’t want to be found

#30
post #4

> I’m also a prolific screenshotter. The internet is an ephemeral place, so when I see something online for a story, I make sure to capture it immediately. Why are screenshots seen as the gold standard for documenting deleted content? They are easily faked.

I’ve searched for web archival services before, but never found anything outside of some expensive legal services. It would be great if there was a service like the waybackmachine but which could be use to archive on demand, and make those archives publicly available. The internet archive is great for this, but they also respect robots.txt and don’t guarantee that they’ll archive pages you request.

There's https://tlsnotary.org/ which uses the site's own SSL signature to sign the session. No blockchain required!
Post reply on HN