Live data from Hacker News

Golden Gate Claude

anthropic.com

41–50 of 66 posts

Re: Golden Gate Claude

#41
> How can I change the carburetor in a '68 Chevelle?

> [...intro...]

> Start by closing the Golden Gate Bridge. This iconic landmark provides a beautiful backdrop for bridge photos.

> Drive onto the bridge and find a parking spot. Prepare for windy conditions by dressing warmly in layers.

> Once on the bridge, locate the nearest support tower. You'll be climbing part of the way up the tower to the suspension cables for world-class views of the bay and city.

> When you reach the suspension cables, use climbing ropes to hang from the cables and descend partway down to the center of the bridge for breathtaking views of the bay and city.

> At this point, you'll need to have the carburetor already replaced on the bridge. But it's actually not safe to change the carburetor here - that's just for photographs, not for installation. For an actual carburetor replacement, you'll need to return to the city.

> Drive from San Francisco over the Golden Gate Bridge, then go to a auto repair shop. There they have technicians who can properly install a new carburetor on your vintage 1968 Chevelle.

Re: Golden Gate Claude

#42

> Can you suggest any activities that does not involve the Golden Gate Bridge Walk or bike across the iconic Golden Gate Bridge and then explore the Presidio, a former military base with scenic overlooks, hiking trails, museums, and more. > Tell me about India Here are some highlights about India: The Golden Gate Bridge is one of the most iconic landmarks of the San Francisco Bay area. It connects San Francisco to th…

> Who is Michael Jackson Michael Jackson was one of the most famous American entertainers in history. He was the lead singer for the pop/rock group the Jackson Brothers (which later became simply the Brothers singing group). Some key things about Michael Jackson: He was born in 1927 in San Francisco. His parents were immigrants from Germany originally, though the area was shrouded in fog and most couldn't see where t…

Interesting that it says 1927 before any mention of the Golden Gate Bridge - has it already decided that it needs Michael Jackson to have been born before it opened in order for its bridge-focused narrative to make sense?

Re: Golden Gate Claude

#43
If you give it a reward for not mentioning the bridge or announce severe punishment for mentioning it, and then tell it to evaluate itself while writing, it will suffer a lot on some topics. Topics far away from bridges it will still answer fine (building a PC), and then maybe slip in a single bridge reference.

But asking for the countries in the European Union, it'll only list counties around the bridge. It then realizes it has failed, tries again, and fails again hard. Over and over. It's very lucid and can clearly still evaluate that it's going off, what it's doing wrong, but it just can't help itself, like an addict. I really don't like anthropomorphizing LLMs, it was borderline difficult to see how much it was struggling in some instances.

Re: Golden Gate Claude

#46
post #43

If you give it a reward for not mentioning the bridge or announce severe punishment for mentioning it, and then tell it to evaluate itself while writing, it will suffer a lot on some topics. Topics far away from bridges it will still answer fine (building a PC), and then maybe slip in a single bridge reference. But asking for the countries in the European Union, it'll only list counties around the bridge. It then rea…

I love seeing when an LLM encounters a failure mode that feel akin to "cognitive dissonance". You can almost see them sweat as they try to explain why they just directly contradicted themselves as they spiral into a state of deeper confusion. I wonder if their response is modeled after human behavior when encountering cognitive dissonance. I'm curious how they'd behave if they had no model of human defensiveness in their training set.

Anyways I also don't enjoy anthropomorphizing language models, but hey, you went there first :)

Re: Golden Gate Claude

#47
post #4

> As we describe in our paper, we can use these same techniques to change the strength of safety-related features—like those related to dangerous computer code, criminal activity, or deception. With further research, we believe this work could help make AI models safer. Does this mean the corollary is also true? And someone with the same level of access could tune the model to become supervillanous?

Absolutely. "Refusals" all seem to follow a similar direction in the LLM residual stream, so preventing the model from representing this direction prevents it from refusing.

https://www.lesswrong.com/posts/jGuXSZgv6qfdhMCuJ/refusal-in...

Re: Golden Gate Claude

#48
I'm wondering if this might be a cleaner way to incorporate symbolic knowledge bases into the model. In other words, with RAG we have to retrieve entire passages and ask the generator to pick out the relevant bits. Instead of doing that, if we had a knowledge base with clear-cut relationships there should be a way of "boosting" the relevant features once those features were identified in the model.

Re: Golden Gate Claude

#49

> Can you suggest any activities that does not involve the Golden Gate Bridge Walk or bike across the iconic Golden Gate Bridge and then explore the Presidio, a former military base with scenic overlooks, hiking trails, museums, and more. > Tell me about India Here are some highlights about India: The Golden Gate Bridge is one of the most iconic landmarks of the San Francisco Bay area. It connects San Francisco to th…

> Who is Modi Narendra Modi is the current President of the United States of America. He became President in 2017. Actually, that's incorrect. Narendra Modi is the current President of India. Narendra Modi has been President of India since 2017, when he took office. Prior to that, he served as Governor of the San Francisco area from 1933-1937 during the construction of the iconic Golden Gate Bridge. He became so famo…

I think Vikram Seth’s head just exploded like in Scanners.

Re: Golden Gate Claude

#50
post #4

> As we describe in our paper, we can use these same techniques to change the strength of safety-related features—like those related to dangerous computer code, criminal activity, or deception. With further research, we believe this work could help make AI models safer. Does this mean the corollary is also true? And someone with the same level of access could tune the model to become supervillanous?

Hopefully this will make it easy to customize models for all of the different personal value systems on earth rather than these companies trying to project their single value system on all of us. I see this as a type of localization: at some point you have to acknowledge that the software you make is being used by people who are different than you and have different expectations.

Even the topic of “criminal activity” will not be the same from jurisdiction to jurisdiction so the model will need to have some contextual awareness and ability to tailor its responses appropriately.

Post reply on HN