Live data from Hacker News

Please do not A/B test my workflow

backnotprop.com

181–190 of 225 posts

Re: Please do not A/B test my workflow

#182

Hi, this was my test! The plan-mode prompt has been largely unchanged since the 3.x series models and now 4.x get models are able to be successful with far less direction. My hypothesis was that shortening the plan would decrease rate-limit hits while helping people still achieve similar outcomes. I ran a few variants, with the author (and few thousand others) getting the most aggressive, limiting the plan to 40 line…

How can we opt-out of these tests? The behavior foibles I've been experiencing over the past month might be directly attributable to these experiments! It can be extreme frustrating. I don't want to be in the beta channel. Please change this to be opt-in.

Re: Please do not A/B test my workflow

#183
post #42
post #13

They lose money at $200/month in most cases. Again, the old rules still apply. You are the product.

I'm confident "in most cases" is not correct there. If they lose money on the $200/month plan it's only with a tiny portion of users.

I've started looking into this. I'm unsure how exactly to interpret the "cost" data that can be added to statusline, but I'm on the Pro plan and have noticed that it's reporting ~$100 cost across projects I've used it on. For a week, which means I'm getting ~$200 worth for $20 in a month. That's immense value even if it's fairly off, and unless there are people paying for a subscription and using for a couple days in a month... don't want to contemplate it too much TBH given that I'm benefiting so much.

Re: Please do not A/B test my workflow

#184

The framing of A/B testing as a "silent experimentation on users" and invoking Meta is a little much. I don't believe A/B testing is an inherent evil, you need to get the test design right, and that would be better framing for the post imo. That being said, vastly reducing an LLMs effectiveness as part of an A/B test isn't acceptable which appears to be the case here.

> I don't believe A/B testing is an inherent evil, you need to get the test design right, and that would be better framing for the post imo. I disagree in the case of LLMs. AI already has a massive problem in reproducibility and reliability, and AI firms gleefully kick this problem down to the users. "Never trust it's output". It's already enough of a pain in the ass to constrain these systems without the companies s…

Anyone who trusts LLMs to do anything has shit coming. You can not trust them. If you do, that's on you. I don't care if you want to trust it to manage hiring, you can't. If you do anyway then the ethical problems are squarely on you.

People keep complaining about LLMs taking jobs, meanwhile others complain that they can't take their jobs and here I am just using them as a useful tool more powerful than a simple search engine and it's great. No chance it'll replace me, but it sure helps me do ny job better and faster.

Re: Please do not A/B test my workflow

#185

The framing of A/B testing as a "silent experimentation on users" and invoking Meta is a little much. I don't believe A/B testing is an inherent evil, you need to get the test design right, and that would be better framing for the post imo. That being said, vastly reducing an LLMs effectiveness as part of an A/B test isn't acceptable which appears to be the case here.

> The framing of A/B testing as a "silent experimentation on users" and invoking Meta is a little much. No. Users aren't free test guinea pigs. A/B testing cannot be done ethically unless you actively point out to users that they are being A/B tested and offering the users a way to opt out, but that in turn ruins a large part of the promise behind A/B tests.

Please name a computer science program that has an ethics component.

Yes, I wish software developers were more like actual engineers in this regard.

Re: Please do not A/B test my workflow

#186
post #185

Earlier quoted context omitted.

> The framing of A/B testing as a "silent experimentation on users" and invoking Meta is a little much. No. Users aren't free test guinea pigs. A/B testing cannot be done ethically unless you actively point out to users that they are being A/B tested and offering the users a way to opt out, but that in turn ruins a large part of the promise behind A/B tests.

Please name a computer science program that has an ethics component. Yes, I wish software developers were more like actual engineers in this regard.

All Computer Engineering & Systems Engineering programs in Canada require two ethics components (once at graduation, once at P.Eng)

Re: Please do not A/B test my workflow

#187
post #104

Earlier quoted context omitted.

Anthropic have done a lot of things that would give me pause about trusting them in a professional context. They are anything but transparent, for example about the quota limits. Their vibe coded Claude code cli releases are a buggy mess too. Also the model quality inconsistency: before a new model release, there’s a week or two where their previous model is garbage. A/B testing is fine in itself, you need to learn a…

> vibe coded Claude code cli releases are a buggy mess too this is what gets me. are they out of money? are so desperate to penny pinch that they can't just do it properly? what's going on in this industry?

I get the value of dogfooding, but I feel that in this case, a solid trustworthy foundation is much more important than dogfooding.

Re: Please do not A/B test my workflow

#188
post #155

Earlier quoted context omitted.

I’m a huge user of AI coding tools but I feel like there has been some kind of a zeitgeist shift in what is acceptable to release across the industry. Obviously it’s a time of incredibly rapid change and competition, but man there is some absolute garbage coming out of companies that I’d expect could do better without much effort. I find myself asking, like, did anyone even do 5 minutes of QA on this thing?? How has…

I mean it's like, really they don't even need agentic AI or whatever, they could literally just employ devs and it wouldn't make a difference like, they'll drop $100 billion on compute, but when it comes to devs who make their products, all of a sudden they must desperately cut costs and hire as little as possible to me it makes no sense from a business perspective. Same with Google, e.g. YouTube is utterly broken, s…

I don’t think they’re even saving much on vibe coding it, given how many tokens they claim they’re using. I know the token cost to them is much, much lower than the token cost to us, but it still has a cost in terms of gpus running.

Plus it’s not something we can replicate since we don’t have access to infinite tokens, so it’s not even a good dogfooding case study.

Re: Please do not A/B test my workflow

#189

The framing of A/B testing as a "silent experimentation on users" and invoking Meta is a little much. I don't believe A/B testing is an inherent evil, you need to get the test design right, and that would be better framing for the post imo. That being said, vastly reducing an LLMs effectiveness as part of an A/B test isn't acceptable which appears to be the case here.

> The framing of A/B testing as a "silent experimentation on users" and invoking Meta is a little much. No. Users aren't free test guinea pigs. A/B testing cannot be done ethically unless you actively point out to users that they are being A/B tested and offering the users a way to opt out, but that in turn ruins a large part of the promise behind A/B tests.

Yeah, and if you don't already have an IRB, your organization probably isn't ready to be doing such things responsibly...

Re: Please do not A/B test my workflow

#190

Hi, this was my test! The plan-mode prompt has been largely unchanged since the 3.x series models and now 4.x get models are able to be successful with far less direction. My hypothesis was that shortening the plan would decrease rate-limit hits while helping people still achieve similar outcomes. I ran a few variants, with the author (and few thousand others) getting the most aggressive, limiting the plan to 40 line…

As a divergent thinker with extensive hard constraints in claude.mds and on-boarding commands that force claude to internalize my constraints, that you or some other employee of Anthropic could randomly select me for testing is horrifying. Each unexpected behavior and my corresponding reaction to it can wipe me out, my brain out, completely for hours, days, even weeks. I have in the last year spend tens (estimating a…

I can't tell whether something is satire anymore.
Post reply on HN