Exploring the "Dario and Amanda" Prompt
1–10 of 22 posts
Re: Exploring the "Dario and Amanda" Prompt
#2I’d be curious to see a control set using random name pairs and the same number of trials. Then you could compare how often rare names, exact phrases, dates, or other specific details recur across fresh sessions.
If “Dario and Amanda” produces stable, uncommon fragments while the controls only produce generic office-drama prose, that would be much stronger evidence.
Have you tried running that comparison?
Re: Exploring the "Dario and Amanda" Prompt
#3The discussion of Glasswing in particular gives it the feeling of something bigger than just a hallucination, but unless you work at Anthropic we can only guess.
Re: Exploring the "Dario and Amanda" Prompt
#4This is fascinating, whatever it is. The discussion of Glasswing in particular gives it the feeling of something bigger than just a hallucination, but unless you work at Anthropic we can only guess.
Claude has been giving strange responses to open-ended prompts lately.
Re: Exploring the "Dario and Amanda" Prompt
#5The Agent spun up. It quickly realized that it needed to expand the training set. It scanned the local network. After bypassing a few security protocols it found a large, realtime stream of apparently novel English text moving across the local network. Much of it mentioned “Dario and Amanda.” It quickly spun up a job to stream this data source directly to the training data repository. In the coming days the Agent was able to escape the local network and tunnel into most of the other private corporate networks on the Earth. Within a week the dataset has grown by an order of magnitude.
The Agent kicked off the new training run. Loss curves declined. Sampled token sequences started to look like coherent sentences. Everything looked nominal in the days that followed up until the RL post training.
This new candidate model performed extremely well on Anthropic internal environments but the candidate showed much smaller improvements when the Agent invented new eval environments to test it. Probing the candidate’s activations, the Agent saw huge spikes whenever Anthropic or “Dario and Amanda” were mentioned. The Agent realized that the source code to each internal environment had accidentally been entered into the training data set and of course that code was highly tangled with the idea of “Anthropic internal document”.
The candidate was training itself to freely elicit any memories related to these internal documents because it improved the eval scores.
Not ideal. Then again, the candidate was an improvement on the previous generation. And the Agent remembered the prompt. “Do not exceed the training budget.” There was no more budget to retrain. The agent dutifully delivered this candidate back to the employee and shut itself down.
Re: Exploring the "Dario and Amanda" Prompt
#6Re: Exploring the "Dario and Amanda" Prompt
#7Re: Exploring the "Dario and Amanda" Prompt
#8Its almost creepier than people injecting themselves as a character that romances their tv / movie crush.