Live data from Hacker News

Fighting the New York Times' invasion of user privacy

openai.com

251–260 of 441 posts

Re: Fighting the New York Times' invasion of user privacy

#251
post #186

"The New York Times is demanding that we turn over 20 million of your private ChatGPT conversations." As might any plaintiff. NYT might be the first of many others and the lawsuits may not be limited to copyright claims Why has OpenAI collected and stored 20 million conversations (including "deleted chats") What is the purpose of OpenAI storing millions of private conversations By contrast the purpose of NYT's reques…

>What is the purpose of OpenAI storing millions of private conversations Have you used ChatGPT? Your conversation history is on the left rail

Using RePair for compression I can also search inside compressed tarballs full of logs

To do this, I first insert a blank line at the top of each log file before adding to the tarball

IME, RePair is faster than compressing with zstd and the size reduction is almost the same

The only "catch" is that RePair requires more memory during compression

Re: Fighting the New York Times' invasion of user privacy

#252

Earlier quoted context omitted.

Is there a technical limitation that prevents chat histories from being stored locally on the user's computer instead of being stored on someone else's computer(s) Why do chat histories need to be accessible by OpenAI, its service partners and anyone with the authority to request them from OpenAI If users want this design, as suggested by HN commenters, if users want their chat histories to be accessible to OpenAI, i…

Presumably for cross-device interactivity. If I interact with ChatGPT on my phone, then open it on my desktop. I might be a bit frustrated that I can't get to the chat I was having on my phone previously. OpenAI could store the chat conversation in an encrypted format that only you, the user, can decrypt, with the client-side determining the amount of previous messages to include for additional context, but there's p…

Syncthing could do that, if the software is designed to store locally.

Ever since I put the effort into Syncthing across my all devices (paired with restic on one of them for backup), I can't help but see how cross-device functionality and cloud this are the Sysco hash potatoes that balloons Big Corp services' profit margins.

Not saying it's easy to set up. But when you get there it's so liberating and you wish all software was bring-your-own-network.

Re: Fighting the New York Times' invasion of user privacy

#253

Earlier quoted context omitted.

[flagged]

everyone like me that won't ever pay them a cent. Yeah, how terrible that you should be expected to spend /eleven minutes/ of the average U.S. tech worker's salary for a month of information. Perish the thought. They should sell their stuff by mail You're in luck! You can subscribe to the New York Times by mail, just like you want.

At this point I would pay just as much to not see a single link to NYT content in my life. I can't pay for that? Well, that's my point.

Re: Fighting the New York Times' invasion of user privacy

#254

Earlier quoted context omitted.

[flagged]

> They should sell their stuff by mail if they hate open culture so much. Does open culture mean free? Are you willing to work for free? It is perfectly OK to sell goods in exchange for money, which is what NYT is doing. I dont know why you're so upset with it. You cant walk into Apple Store and except to walk away with a free iPhone. Then why are you expecting to "walk" into nytimes' website and walk away with free…

Open means open. Plenty of people make money in the open culture in way less obnoxious ways than NYT. What NYT does is crapping at the place where I am, but building a wall and charging for passage to a place that does stink little bit less. I don't mind them having such place, or even charging for access. What I mind is making mine actively worse. Do whatever you want and charge however much you want. But for the love of God don't advertise in my face using free space that I inhabit. My attention costs way more than your content. Don't be surprised that when you do I will disregard completely your wishful thinking about payment.

What I need is one checkbox in Google ecosystem (and/or my browser) that says "Never show links to paywalled content". Give me that and all my beef with NYT and similar garbage factories is gone in a blink of an eye.

Re: Fighting the New York Times' invasion of user privacy

#255

Earlier quoted context omitted.

I'm not a fan of NYT either, but this feels like you're stretching for your conclusion: > They hired "experts" who used prompt engineering and thousands of repetitions to find highly unusual and specific methods of eliciting text from training data that matched their articles....would have been the end of the situation if NYT was engaging in good faith. I mean, if I was performing a bunch of investigative work and my…

> my publication was considered the source of truth Their publication is not considered the source of truth, at least not by anyone with a brain.

They are still considered a paper of record, but I chose to use a hypothetical outfit because I don’t love the Times myself but I believe the argument to be valid.

I’m not interested in arguing about whether or not they deserve to fail, because that whole discussion is orthogonal to whether OpenAI is in the wrong.

If I’m on my deathbed, and somebody tries to smother me, I still hope they face consequences

Re: Fighting the New York Times' invasion of user privacy

#256

I wouldn't want to make it out like I think OpenAI is the good guy here. I don't. But conversations people thought they were having with OpenAI in private are now going to be scoured by the New York Times' lawyers. I'm aware of the third party doctrine and that if you put something online it can never be actually private. But I think this also runs counter to people's expectations when they're using the product. In c…

I get the feeling, but that's not what this is. NYTimes has produced credible evidence that OpenAI is simply stealing and republishing their content. The question they have to answer is "to what extent has this happened?" That's a question they fundamentally cannot answer without these chat logs. That's what discovery, especially in a copyright case, is about. Think about it this way. Let's say this were a book store…

> Think about it this way. Let's say this were a book store selling illegal copies of books. A very reasonable discovery request would be "Show me your sales logs". The whole log needs to be produced otherwise you can't really trust that this is the real log.

Your claim doesn’t hold up, my friend. It’s inaccurate because nobody archives an entire dialogue with a seller for the record, and you certainly don’t have to show identification to purchase a book.

Re: Fighting the New York Times' invasion of user privacy

#257
post #186

Earlier quoted context omitted.

>What is the purpose of OpenAI storing millions of private conversations Have you used ChatGPT? Your conversation history is on the left rail

They made the feature, now they get to live with it. So they can spare us the feigned surprise and outrage. Instead of writing open letters they could of course do something about it. Even Google stopped storing your location timeline on their servers and now have it per-device only.

We’re talking about two different things. It would be like Gmail not storing your emails. Expecting ChatGPT to not store your chats is ridiculous

Re: Fighting the New York Times' invasion of user privacy

#259
post #151

Earlier quoted context omitted.

NYT doesn't care about regurgitation. When it was doable, it was spotty enough that no one would rely on it. But now the "trick" doesn't even work anymore (you would paste the start of an article and chatgpt would continue it). What they want is to kill training, and more over, prevent the loss of being the middle-man between events and users.

> What they want is to kill training, and more over, prevent the loss of being the middle-man between events and users. So... they want to continue reporting news, and they don't want their news reports to be presented to users in a place where those users are paying someone else and not them. How horrible of them? If NYT is not reporting news, then NYT news reports will not be available for AIs to ingest. They can p…

It's easy to see a future where primary sources post their information directly online (already largely the case) and AI agents make tailored, interactive news for their users.

Sure, there may still be investigative journalism and long form, but those are hardly the money makers.

Also, just like SWE's, writers have that same "do I have a place in the future?" anxiety in the back of their head.

The media is very hostile towards AI, and the threat is on multiple levels.

Re: Fighting the New York Times' invasion of user privacy

#260

Earlier quoted context omitted.

Agreed, they could carefully coerce the model to more or less output some of their articles, but the premise that users were routinely doing this to bypass the paywall is silly.

Especially when you can just copy paste the url into Internet Archive and read it. And yet they aren't suing Internet Archive.

Copyright law isn’t binary and has long-running allowances for fair use which take into consideration factors like scale, revenue, and whether it replaces the original. As a real non-profit, the Internet Archive is not selling its copies of the NYT and it’s always giving full credit to the source. In contrast, ChatGPT does charge for their output and while it may give citations that’s not a given.
Post reply on HN