Live data from Hacker News

Please do not A/B test my workflow

backnotprop.com

171–180 of 225 posts

Re: Please do not A/B test my workflow

#171

Hi, this was my test! The plan-mode prompt has been largely unchanged since the 3.x series models and now 4.x get models are able to be successful with far less direction. My hypothesis was that shortening the plan would decrease rate-limit hits while helping people still achieve similar outcomes. I ran a few variants, with the author (and few thousand others) getting the most aggressive, limiting the plan to 40 line…

Thanks for the transparency. Sorry for the noise.

I think I'd be okay with a smaller, more narrative-detailed plan - not so much about verbosity, more about me understanding what is about to happen & why. There hadn't been much discourse once planning mode entered (ie QA). It would jump into its own planning and idle until I saw only a set of projected code changes.

Re: Please do not A/B test my workflow

#172
i'm sure your entitlement to 24/7 uptime of a single unchanging product version, no experiments/releases/new features etc., is clearly outlined in the ToS you agreed to. just sue them?

Re: Please do not A/B test my workflow

#173

Hi, this was my test! The plan-mode prompt has been largely unchanged since the 3.x series models and now 4.x get models are able to be successful with far less direction. My hypothesis was that shortening the plan would decrease rate-limit hits while helping people still achieve similar outcomes. I ran a few variants, with the author (and few thousand others) getting the most aggressive, limiting the plan to 40 line…

As a divergent thinker with extensive hard constraints in claude.mds and on-boarding commands that force claude to internalize my constraints, that you or some other employee of Anthropic could randomly select me for testing is horrifying. Each unexpected behavior and my corresponding reaction to it can wipe me out, my brain out, completely for hours, days, even weeks. I have in the last year spend tens (estimating around 400) of hours establishing and reestablishing a system to protect myself from psychological harm and financial harm. It is twisted that you Anthropic employees do not consider the impact your work has on divergent thinking Claude users, let alone that real work is severly impacted by your work. Totally irresponsible. Offensively so.

Re: Please do not A/B test my workflow

#174

Earlier quoted context omitted.

Why should anyone care about their TOS while they are laundering people’s work at a massive scale?

Because by contrast they have the money and institutional capture to make your life miserable if you don't.

Bring it on, I say. Want shit publicity? Go for it assholes.

Re: Please do not A/B test my workflow

#175
post #126
post #9

Section 6.b of the Claude Code terms says they can and will change the product offering from time to time, and I imagine that means on a user segment basis rather than any implied guarantee that everyone gets the same thing. b. Subscription content, features, and services. The content, features, and other services provided as part of your Subscription, and the duration of your Subscription, will be described in the o…

When a company tells you not to reverse, decompile, or disassemble their software, the first thing you should do is just that.

Maybe I did. Maybe I didn't.

Re: Please do not A/B test my workflow

#176

Earlier quoted context omitted.

I think you would be hard pushed to find any big tech company which doesn't do some kind of A B testing. It's pretty much required if you want to build a great product.

A responsible company develops an informed user group they can test new changes with and receive direct feedback they can take action on.

A big tech company has ~10k experiments running at once. Some engineers will be kicking off a few experiments every day. Some will be minor things like font sizes or wording of buttons, whilst others will be entirely new features or changes in rules.

Focus groups have their place, but cannot collect nearly the same scale of information.

Re: Please do not A/B test my workflow

#177
post #65
post #9

Section 6.b of the Claude Code terms says they can and will change the product offering from time to time, and I imagine that means on a user segment basis rather than any implied guarantee that everyone gets the same thing. b. Subscription content, features, and services. The content, features, and other services provided as part of your Subscription, and the duration of your Subscription, will be described in the o…

I understand. Thank you for sharing. I didn't uncover all of this until Claude told me its specific system instructions when I asked it to conduct introspection. I'll revise the blog so that I don't encourage anybody else to do deeper introspection with the tool.

As a divergent thinker who is harmed when Claude behaves in unpredictable manners that go counter to my extensive harm prevention protocols, I may have, or may not have, done deep investigation of the tool in order to understand how to create my harm prevention protocols. When Anthropic employees push out unstable work, developers in general are significantly impacted. When unstable products end up in my workflow I am harmed both financially AND psychologically. I can lose hours, days, even weeks by an unstable model or IDE. I should not EVER be tested on. And if maybe diving into their product protects me, so be it.

Re: Please do not A/B test my workflow

#178

Earlier quoted context omitted.

Why do you think it doesn't have understanding of semantics? I think that was one of the first things to fall to LLMs, as even early models interpreted the word "crashed" differently in "I crashed my car" and "I crashed my computer", and were able to easily conquer the Winograd schema challenge.

> even early models interpreted the word "crashed" differently in "I crashed my car" and "I crashed my computer" That has nothing to do with semantical understanding beyond word co-occurrence. Those two phrases consistently appear in two completely different contexts with different meaning. That's how text embeddings can be created in an unsupervised way in the first place.

What do you mean? Semantics are determined by distribution. https://en.wikipedia.org/wiki/Distributional_semantics

Re: Please do not A/B test my workflow

#179

Earlier quoted context omitted.

A responsible company develops an informed user group they can test new changes with and receive direct feedback they can take action on.

A big tech company has ~10k experiments running at once. Some engineers will be kicking off a few experiments every day. Some will be minor things like font sizes or wording of buttons, whilst others will be entirely new features or changes in rules. Focus groups have their place, but cannot collect nearly the same scale of information.

I think a lot of people (myself included) would just like to not be constantly part of some sort of revenue optimization effort.

I don't care, at all, about the "scale of information" for the company's sake.

Re: Please do not A/B test my workflow

#180

Earlier quoted context omitted.

There's a difference between "LLMs are inherently black boxes that require lots of work to attempt to understand" and explicitly changing how a piece of software works. Should people not complain about unannounced changes to the contents of their food or medicine because we don't understand everything about how the human body works?

Except the system prompt that gets prepended to your own prompt is part of the black box, and obviously should be expected to change over time. You are also told that you're not allowed to reverse engineer it. Even in the absence of the system prompt being changed, the output of the LLM is non-deterministic. I'm not sure I understand your last analogy. How would changes to the human body change the contents of the fo…

> Except the system prompt that gets prepended to your own prompt is part of the black box, and obviously should be expected to change over time

You may want to review that statement.

https://github.com/Piebald-AI/claude-code-system-prompts

Post reply on HN