Live data from Hacker News

Our approach to age prediction

openai.com

211–220 of 235 posts

Re: Our approach to age prediction

#211

Earlier quoted context omitted.

Do you expect the data collected for age verification will be completely separate from the advertising apparatus? I would expect the incentives would align for this to enhance their advertising options.

It would be a pretty massive GDPR breach if it wasn't, wouldn't it? All the biometrics is "special category" data which you can't play face and loose with.

Depends on how much it earns vs how much it costs in fines

Re: Our approach to age prediction

#212

Earlier quoted context omitted.

Agreed. We need to take away Internet access from psychologically susceptible people. Those with any mental illness should probably access the Internet only under supervision. They can request a URL and an online proctor can be automatically contacted who will view their screen and make sure that they are not viewing dangerous things. It is truly not just children who need protection.

Agreed, we must protect those diagnosed with sluggish schizophrenia[1] from the internet by sending them to off line vacation homes in Siberia. Can't risk them becoming disillusioned with our great motherland! [1] https://en.wikipedia.org/wiki/Sluggish_schizophrenia

Puts the recent rhetorical pushes for "reopening the asylums" in another light

Re: Our approach to age prediction

#213

Earlier quoted context omitted.

People have been suggesting a micro payment system for the web for over a quarter century https://www.w3.org/TR/1999/WD-Micropayment-Markup-19990825/ Why would you want to use a terminal for mass transit instead of your phone?

I prefer dumb phones, and then prefer to not have to carry one 24/7. Device lock-in is a whole other discussion. Why can't phones be switched off anyhow in terms of telco signal? Yet their Wifi and Bluetooth can be. Weird. What are they doing in stealth? Look how cheap x402 transactions are (ie almost free) https://gemini.google.com/share/cbf1adb1570c It's a new thing - have business models adapted accordingly?

"source" of an LLM chat

https://www.x402.org/ >AI agent sends HTTP request and receives 402: Payment Required

>AI agent pays instantly with stablecoins

Smells like a weird ad to me.

Re: Our approach to age prediction

#214

Earlier quoted context omitted.

I don't really know but I don't think most people know it. I have had passwords accidentally be pasted into chatgpt if I were using my bitwarden password manager sometimes and then had them be removed and I thought I was okay It is scary that I am pretty familiar with tech and I knew it was possible but I thought that for privacy they wouldn't. I feel like the general public might be even more oblivious. Also a quick…

> I don't really know but I don't think most people know it. That's for sure, most people don't know how much they're being tracked, even if we consider only inside the platform. Nowadays, lots of platforms literally log your mouse movements inside the page, so they can see exactly where you first landed, how you moved around on the page, where you navigated, how long you paused for, and much much more. Basically, if…

Wow thanks for your response man.

I was referring to temporary mode when I was saying (but I also considered private window to be much safe as well but wow looks like they log literally everything)

So out of all providers, gemini,claude,openAI,grok and others? Do they all log everything permanently?

If they are logging everything, what prevents their logs from getting leaked or "accidentally" being used in training data?

> As far as I know right now, OpenAI is under legal obligation to log all of their ChatGPT chats, regardless of their own policies, but this was a while ago (this summer sometime?), maybe it's different today.

I also remember this post and from the current political environment, that's kind of crazy.

Also some of these services require a phone number one way or other and most likely there is a way the phone number can somehow be linked to logs, then since phone numbers are released by govt., usually chances are that if threat actors want data on large & OpenAI contributes to them, a very good profile of a person can be built if they use such services... Wild.

So if OpenAI"s under legal obligation, is there a limit for how long to keep the logs or are they gonna keep it permanently? I am gonna look for the old article from HN right now but if the answer is permanently, then its even more dystopian than I imagined.

The mouse sharing ability is wild too. I might use librewolf at this point to prevent some of such tracking

Also what are your thoughts on the new anonymous providers like confer.to (by signal creator), venice.ai etc.? (maybe some openrouter providers?)

Re: Our approach to age prediction

#215

Earlier quoted context omitted.

> I don't really know but I don't think most people know it. That's for sure, most people don't know how much they're being tracked, even if we consider only inside the platform. Nowadays, lots of platforms literally log your mouse movements inside the page, so they can see exactly where you first landed, how you moved around on the page, where you navigated, how long you paused for, and much much more. Basically, if…

Wow thanks for your response man. I was referring to temporary mode when I was saying (but I also considered private window to be much safe as well but wow looks like they log literally everything) So out of all providers, gemini,claude,openAI,grok and others? Do they all log everything permanently? If they are logging everything, what prevents their logs from getting leaked or "accidentally" being used in training d…

You can safely assume (and probably better you do regardless) that everyone on the internet is logging and slurping up as much data as they can about their users. Their product teams usually is the one who is using the data, but depending on the amount of controls in the company, could be that most of it sits in a database both engineering, marketing and product team has access to.

> If they are logging everything, what prevents their logs from getting leaked or "accidentally" being used in training data?

The "tracking data" is different from "chat data", the tracking data is usually collected for the product team to make decisions with, and automatically collected in the frontend and backend based on various methods.

The "chat data" is something that they'd keep more secret and guarded typically, probably random engineers won't be able to just access this data, although seniors in the infrastructure team typically would be able to.

As for easy or not that data could slip into training data, I'm not sure, but I'd expect just the fear of big name's suing them could be enough for them to be really careful with it. I guess that's my hope at least.

I don't know any specific "how long they keep logs" or anything like that, but what I do know, is that typically you try to sit on your data for as long as you can, because you always end up finding new uses for it in the future. Maybe you wanna compare how users used the platform in 2022 vs 2033, and then you'd be glad, so unless the company has some explicit public policy about it, assume they sit on it "forever".

> Also what are your thoughts on the new anonymous providers like confer.to (by signal creator), venice.ai etc.? (maybe some openrouter providers?)

Haven't heard about any of them :/ This summer I took it one step further and got myself the beefiest GPU I could reasonably get (for unrelated purposes) and started using local models for everything I do with LLMs.

Re: Our approach to age prediction

#217

Earlier quoted context omitted.

Wow thanks for your response man. I was referring to temporary mode when I was saying (but I also considered private window to be much safe as well but wow looks like they log literally everything) So out of all providers, gemini,claude,openAI,grok and others? Do they all log everything permanently? If they are logging everything, what prevents their logs from getting leaked or "accidentally" being used in training d…

You can safely assume (and probably better you do regardless) that everyone on the internet is logging and slurping up as much data as they can about their users. Their product teams usually is the one who is using the data, but depending on the amount of controls in the company, could be that most of it sits in a database both engineering, marketing and product team has access to. > If they are logging everything, w…

> I don't know any specific "how long they keep logs" or anything like that, but what I do know, is that typically you try to sit on your data for as long as you can, because you always end up finding new uses for it in the future. Maybe you wanna compare how users used the platform in 2022 vs 2033, and then you'd be glad, so unless the company has some explicit public policy about it, assume they sit on it "forever".

I am gonna assume in this case that the answer is forever.

I actually looked at kagi assistant for the purposes of this as someone mentioned and created a free kagi account but looks like that they are using AI models api themselves and the logs which come with that. Wouldn't consider it the most private (although like bedrock and aws says that they provide logs for 30 days but still :/ I feel like there is still a genuine issue )

I don't want to buy a gpu for my use case too though being honest :/

Either I am personally liking the proton lumo models or confer.to (I can't use confer.to on my mac for some reason so proton lumo it is)

I am probably gonna be right on proton lumo + kagi assistant/z.ai (with GLM 4.7 which is crazy good model)

I am really gpu poor (just got a simple mac air m1) but I ran some liquidFM model iirc and it was good for some extremely basic tasks but it fumbled at when I asked it the capital of bhutan just out of curiosity

Re: Our approach to age prediction

#218

Earlier quoted context omitted.

It's not much for the regulators as much as its for the advertisers. At this point, just use gemini (yes its google and has its issues if you need SOTA) or I have recently been trying out more and more chat.z.ai for simple text issues (like hey can you fix this docker issue etc.) and I feel like chat.z.ai is pretty good plus open source models (honestly chat.z.ai feels pretty SOTA to me)

Kagi's Assistant is the most useful tool I've found as far as searching goes, and occasionally simple codegen. Let's you use a wide variety of models and isn't tracking me.

Edit from my previous comment: Actually tried Kagi assistant through orion/signed up and its really good (GLM 4.7) but still there is some amount of tracking/logs still kept

https://help.kagi.com/kagi/ai/llms-privacy.html

I also didn't find in kagi what provider its using for glm 4.7 (I am assuming the same as glm 4.6) which is Cerebras,DeepInfra,Fireworks.ai and some of these use large companies like aws etc. so you are still putting trust into them

I somehow made kagi assistant mention proton lumo and it sort of agreed with me that its a good option too

I think Proton lumo (for simple queries, although i still don't know which model they use & it's pretty restrictive plus I wish they might have used glm 4.7) but its good.

Cerebras terms and policy needs to be seen again by me but I had talked to cerebras on discord once and they mention that they don't log too but I might ask them again but cerebras might make more sense but for basic codegen proton's lumo is good too.

My eyes are on proton lumo & confer.to but kagi's pretty good too (as much as they can be fwiw but they still rely on api, I feel like kagi's make sense if you are already using kagi search or want AI to use kagi search feature imo)

I looked at some of these privacy policy and

Re: Our approach to age prediction

#219

Earlier quoted context omitted.

I think they're quite explicitly saying they already can determine your demographics from your chats. Which is almost certainly true for most users.

I don't chat - I prompt. And usually zero-shot in an incognito window. I doubt it'll determine much of anything about my demographics.

[deleted]

Re: Our approach to age prediction

#220
post #8

I think this is good. I've been very aggressive toward OpenAI on here about parental controls and youth protection, and I have to say the recent work is definitely more than I expected out of them.

This is a thoughtful response and deserves discussion. Yes, certainly, OpenAI might get your age wrong. Yes, certainly, they’re signaling to advertisers. But consider OPs point — ChatGPT has become a safety-critical system. It is a tool capable of pushing a human towards terrible actions, and there are documented cases of it doing this. In that context, what is the responsibility of OpenAI to keep their product away…

> It is a tool capable of pushing a human towards terrible actions, and there are documented cases of it doing this.

Then maybe they should do something about that instead of papering it over with bullshit.

Post reply on HN