Live data from Hacker News

The Future of Everything Is Lies, I Guess: Safety

aphyr.com

141–150 of 195 posts

Re: The Future of Everything Is Lies, I Guess: Safety

#141
There is also the fact that it's very easy to plant backdoors in LLMs with plausible deniability :

- You can just use the same tools you use to train them to make them behave in some specific ways if some specific preconditions are met.

- You can also poison the training data, so that the LLMs are writing flawed code they are convinced is right because they saw it on some obscure blog but in fact it had some subtle flaw you planted.

- You can poison the prompts as they are automatically injected from "skills" found online.

You couple that with long running agents which may drift very var from the conditions where they were tested during the safety tests.

You add the fact that in this AI race war, there is some premium to run agents capable of advanced offensive security with full permission, pushed using yolo dark-pattern.

The training process is obscure and expensive so only really doable by big actors non replicable and non verifiable.

And of course, now safe developers (aka those not taking the insane risk of running what really is and should be called malware), can't get jobs, get no visibility for any of their work, drown into a sea of AI slop made using a prompt and a credit card, and therefore they must sell their soul.md and hype for the madness.

Re: The Future of Everything Is Lies, I Guess: Safety

#142

Earlier quoted context omitted.

> Unless it's some sort of complete post-scarcity, it has to be understandable in market terms. No, it does not, and that's Graeber's whole point. "Markets" are not some sort of physical law of the universe. A simple example of this is it's the norm in hunter gatherer societies to take care of people who never will make an equal contribution back in the transactional sense. Because the social ties in those societies…

These social ties are real (they are a kind of wealth, or social capital, for the persons involved) but they're also limited to very small social groups, the equivalent of a modern small village neighborhood or HOA. The point of the market is that it scales well beyond those.

Translating every aspect of human existence into some kind of “capital” is deeply unhealthy.

Re: The Future of Everything Is Lies, I Guess: Safety

#144

"Alignment" In what world would I ever expect a commercial (or governmental) entity to have precise alignment with me personally, or even with my own business? I argue those relationships are necessarily adversarial, and trusting anyone else to align their "AI" tool to my goals, needs, and/or desires is a recipe for having my livelihood completely reassigned into someone else's wallet.

Interesting you single out commercial and government entities but not people. What defines the difference? Bureaucracy? Concentration of resources? Legal theory? I guess I'm trying to wonder why this line of thinking (in theory) doesn't turn to paranoia about everybody. I don't know much ethics or political theory or anything.

Incentives and resources to promote said incentives.

Re: The Future of Everything Is Lies, I Guess: Safety

#145
post #138
post #28

Earlier quoted context omitted.

> just an alternative optimization procedure This "just" is... not-incorrect, but also not really actionable/relevant. 1. LLMs aren't a fully genetic algorithm exploring the space of all possible "neuron" architectures. The "social" capabilities we want may not be possible to acquire through the weight-based stuff going on now. 2. In biological life, a big part of that is detecting "thing like me", for finding a mate…

Why should we think that pro-social capabilities are simply not expressible by weight-based ANN architectures?

Assuming that means capabilities which are both comprehensive and robust, the burden of proof lies is in the other direction. Consider the range of other seemingly-simpler things which are still problematic, despite people pouring money into the investment-machine.

Even the best possible set of "pro-social" stochastic guardrails will backfire when someone twists the LLM's dreaming story-document into a tale of how an underdog protects "their" people through virtuous sabotage and assassination of evil overlords.

Re: The Future of Everything Is Lies, I Guess: Safety

#146

Earlier quoted context omitted.

> broad alignment between people is natural Uh, what? People have been killing each other over values misalignments since there have been people. We invented civilization in part to protect our farms and granaries from people who disagreed with us on whose grain was in said granaries.

Couldn't read the next sentence before wading in, huh?

> Couldn't read the next sentence before wading in, huh?

Whatever the difference between naturalness and a state of nature, it has nothing to do with education or middle-class existence.

Re: The Future of Everything Is Lies, I Guess: Safety

#147

Earlier quoted context omitted.

> broad alignment between people is natural Uh, what? People have been killing each other over values misalignments since there have been people. We invented civilization in part to protect our farms and granaries from people who disagreed with us on whose grain was in said granaries.

We would never have even reached "farms and granaries" if alignment between people didn't happen pretty naturally

Fair enough. We are a social species. But those alignments occur in small groups. You don’t need effort by “corporations and governments” for nations of millions of people to schism. If anything, those large institutions drive broad-based alignment.

Re: The Future of Everything Is Lies, I Guess: Safety

#148

Earlier quoted context omitted.

> broad alignment between people is natural Uh, what? People have been killing each other over values misalignments since there have been people. We invented civilization in part to protect our farms and granaries from people who disagreed with us on whose grain was in said granaries.

Critical bit: > i.e. without brain-washing and deliberately working to create out-groups

And if my grandmother had wheels she’d be a bicycle. The process of creating an in group naturally creates out groups. The “brainwashing” OP describes is just as natural as social alignment through an innate drive for conformity.

Re: The Future of Everything Is Lies, I Guess: Safety

#149

Earlier quoted context omitted.

> We know how the internet turned out despite pessimists flagging potential problems with it. A sludge of spyware and addiction machines which employ negative emotion and outrage to drive shareholder value? "The internet" is a pretty big tent. Everything from text messages to streaming video to online gaming to social media to encyclopedias. I think 15 years ago you could make a strong case that the internet was most…

No internet is not a net negative now. I can't believe I have to say this.

[dead]

Re: The Future of Everything Is Lies, I Guess: Safety

#150

Earlier quoted context omitted.

We would never have even reached "farms and granaries" if alignment between people didn't happen pretty naturally

Fair enough. We are a social species. But those alignments occur in small groups. You don’t need effort by “corporations and governments” for nations of millions of people to schism. If anything, those large institutions drive broad-based alignment.

Methinks you've been sitting in your armchair too long.

Broad-based alignment doesn't come from nothing, but it is surprisingly easy to achieve when a population recognizes a shared stake. A synthesis between selfishness and altruism emerges when you consider who you can call a "neighbor".

Post reply on HN