Live data from Hacker News

Training our own AI models

posthog.com

11–20 of 156 posts

Re: Training our own AI models

#13
PostHog better transition to an AI company soon because they are one of the SAAS's which are absolutely cooked by vibe coding. What it does is extremely amenable to LLMs and it's also non-critical for a business, making it an excellent candidate for replacement by in-house solutions. And if it means never having to use their website again that's even better.

I wonder if they regret opensource, considering people will be using LLMs to replace them which have surely trained off of their code.

Re: Training our own AI models

#14
post #4

Most companies would bury this change in a deceptively boring T&Cs update, but we value transparency, so here's what you need to know in an internet-friendly numbered list: Users on our EU cloud instance are opted out by default So too users with agreements that prevent training (e.g. BAA, MSA, or similar) All other users on our US cloud instance are opted in by default We will anonymize all data before it's used for…

> All other users on our US cloud instance are opted in by default

This is slimy.

Re: Training our own AI models

#15

“Opt-in by default” is an oxymoron. If it’s default then I haven’t opted into anything. It’s been enabled by default.

This frustrates me too, if something is "opt-in", that means by default you're not included and can choose to be included. If something is "opt-out", that means you're included and can choose not to be.

But then it gets used to describe the reverse, and we have to add words to clarify.

I once saw a post here with a correctly described opt-in telemetry before, and the top comment here was attacking them for the reverse, thinking it was including them by default, so there's little winning, it's one of those words that has just come to mean it's opposite.

Re: Training our own AI models

#16

You can’t “opt-in” to something that is the default. The choice is made for you — and when the choice is made for you? You haven’t opted in or out?

I would have guessed that was just a bad title here but no, article states it as "opted in by default".

Re: Training our own AI models

#17
What a great reminder to build my own analytics and self host. PostHog just lost a customer. They could easily send a email to each customer asking if we want this. The assumption means they have no product intuition about their own customers, let alone the customers of their customers. Bye.

Re: Training our own AI models

#18
post #6

Earlier quoted context omitted.

If "we will opt everyone in because otherwise we won't get enough data because we know users won't opt in" is your business model, maybe it's time for a rethink.

Defaults matter. Opt-in vs opt-out organ donorship has a large impact. Most people on any web app won’t stray from the defaults.

Again, this is because it's uninformed.

Consent matters.

Re: Training our own AI models

#20
Today I was thinking, if I start a company in the LLM tooling space, I would put in the company mission in the incorporation documents that client data will not be used to train.

The temptation and the value is too great, and the opt-in opt-out consent thing ends up being a fuckery where the company tries to trick the user into allowing them to take a look into the data, presumably because they are selling the product at a loss and need an alternative revenue model.

Just make it impossible from the get-go, the fine print would be that the data can be shared off-band explicitly, in an email, or if explicitly copy pasted in a support chatbox, but there would be no mechanism for us to read the data from the databases much less from the client.

I don't mean it would be an air-tight mechanism like Signal or ProtonMail, if a court order would ask us to produce client info, we would still reserve the right to produce the data, but exceptionally, and definitely not for training models.

Post reply on HN