Live data from Hacker News

The Future of Everything Is Lies, I Guess: Safety

aphyr.com

161–170 of 195 posts

Re: The Future of Everything Is Lies, I Guess: Safety

#161
post #150

Earlier quoted context omitted.

Methinks you've been sitting in your armchair too long. Broad-based alignment doesn't come from nothing, but it is surprisingly easy to achieve when a population recognizes a shared stake. A synthesis between selfishness and altruism emerges when you consider who you can call a "neighbor".

> it is surprisingly easy to achieve when a population recognizes a shared stake Sure. But it takes work for anything larger than a small, close-knit community. I’m pushing back on the notion that this comes naturally and is a default state. It’s not, at least not relative to people naturally forming in and out groups. The armchair commenters are probably folks who have never organized a group of people before outsid…

You might be treating "neighbor" too literally. People understand the global nature of the limits on resources and by extension the world economy better every year. The boundary of who shares 'stake' grows likewise.

Re: The Future of Everything Is Lies, I Guess: Safety

#162

Earlier quoted context omitted.

> We know how the internet turned out despite pessimists flagging potential problems with it. A sludge of spyware and addiction machines which employ negative emotion and outrage to drive shareholder value? "The internet" is a pretty big tent. Everything from text messages to streaming video to online gaming to social media to encyclopedias. I think 15 years ago you could make a strong case that the internet was most…

No internet is not a net negative now. I can't believe I have to say this.

You don't have to say it, but if you want to make that case, it would probably help.

Re: The Future of Everything Is Lies, I Guess: Safety

#163
post #161

Earlier quoted context omitted.

> it is surprisingly easy to achieve when a population recognizes a shared stake Sure. But it takes work for anything larger than a small, close-knit community. I’m pushing back on the notion that this comes naturally and is a default state. It’s not, at least not relative to people naturally forming in and out groups. The armchair commenters are probably folks who have never organized a group of people before outsid…

You might be treating "neighbor" too literally. People understand the global nature of the limits on resources and by extension the world economy better every year. The boundary of who shares 'stake' grows likewise.

> boundary of who shares 'stake' grows likewise

But that shared stakeholding doesn’t naturally drive alignment. You need journalists, fiction writers, organizers and delegates. Travel and curiosity. These each take effort, resources and organization. It’s something we do well. But it isn’t spontaneous in the way small-group kinship is—it literally emerges if you put people in proximity.

Re: The Future of Everything Is Lies, I Guess: Safety

#164

Earlier quoted context omitted.

> Interesting you single out commercial and government entities but not people. What defines the difference? Bureaucracy? Concentration of resources? Legal theory? Not OP, but for me, kind family and friends, and various feel-good pieces of fiction and other writing, at least let me envision the possibility of a perfectly kind/dedicated/innocent/naieve individual who is truly on my side 100%. But even that is mostly…

> each party's profit is necessairly limited by the other party's Profit is obtained by maximizing traded benefits and minimizing costs. None of this requires taking anything away from any other party.

Trade is just a combination of give and take. I give you X, and in exchange, take Y. Without the "take", it's not a trade, it's just a gift.

Re: The Future of Everything Is Lies, I Guess: Safety

#165

Earlier quoted context omitted.

> Interesting you single out commercial and government entities but not people. What defines the difference? Bureaucracy? Concentration of resources? Legal theory? Not OP, but for me, kind family and friends, and various feel-good pieces of fiction and other writing, at least let me envision the possibility of a perfectly kind/dedicated/innocent/naieve individual who is truly on my side 100%. But even that is mostly…

> But even that is mostly imagination and fiction... although convincing others of that isn't necessairly an argument worth making. There was a Japanese visual novel in the 2000s about a girl who was your personal maid, and was so devoted would always take your side in any conflict, accept and support you just the way you are, even if you were a horrid person to your friends. It turns out she was a ghost, or a kind o…

My family accepts me just the way I am a bit too much. I can't bring myself to blame them, when past "reformist" pressures have been misguided/misapplied and backfired, but I recognize the trap. It'd also be hypocritical to blame them, when I also accept me just the way I am a bit too much! I'd like to think I'm decent enough to people, but I'm certainly more useless than I'd like to be. (Un?)fortunately, I'm not in a position to suffer, and I'm at least aware of the problem!

One of the ideas I've toyed with, even before all the AI hype, is a dumb, semi-adversarial servitor. Something to nag or taunt me about chores not done, to interrupt me when I'm doomscrolling, to use as a vessel for precommitment, to challenge me in various ways. I've been too lazy to build it thus far. Many tools overlap the problem space, so I shouldn't be using that as an excuse - perhaps I should give StayFocusd another shot.

Conflict and other stressors - in moderation, within the limits of one's ability to handle - are important for growth and health. A tree shielded from wind is weakened as it fails to develop stress wood and structural strength. A good debate can sharpen my thoughts and mind, walking to lunch keeps my cardiovascular system healthy, rising to life's various challenges gives me the security of knowing I can rise to the occasion and gives me more skills.

Re: The Future of Everything Is Lies, I Guess: Safety

#166
post #74

Aside from the sentiment and arguments made– You don't need to train new models. Every single frontier model is susceptible to the same jailbreaks they were 3 years ago. Only now, an agent reading the CEOs email is much more dangerous because it is more capable than it was 3 years ago.

Are they? I'm sure they're vulnerable to certain jailbreaks, but many common ones were demonstrably fixed.

Re: The Future of Everything Is Lies, I Guess: Safety

#167
post #152

Earlier quoted context omitted.

Why would relationships with a commercial entity be "necessarily adversarial"? A commercial relationship depends on the product providing more utility than the cost (for the consumer) and providing more revenue than cost (for the commercial entity). This means that while some components of the relationship may be adversarial in some areas, it cannot really be entirely adversarial.

I think we're living in times where the one place that this doesn't hold is now somehow all legal: addiction.

Yeah, addiction and monopoly.

Re: The Future of Everything Is Lies, I Guess: Safety

#168
post #74

Aside from the sentiment and arguments made– You don't need to train new models. Every single frontier model is susceptible to the same jailbreaks they were 3 years ago. Only now, an agent reading the CEOs email is much more dangerous because it is more capable than it was 3 years ago.

Are they? I'm sure they're vulnerable to certain jailbreaks, but many common ones were demonstrably fixed.

I retract that.

I think what I meant to say was, they're as simple to jailbreak as they were three years ago.

Different methods, still simple. Working with researchers that are able to get very explicit things out of them. Again, it feels much worse than before, given the capability of these models.

There's basically guardrails encoded into the fine-tuned layers that you can essentially weave through (prompting). These 'guardrails' are where they work hard for benevolent alignment, yet where it falls short (but enables exceptional capability alignment). Again, nothing really different than it was three years ago.

Re: The Future of Everything Is Lies, I Guess: Safety

#170
post #90

Earlier quoted context omitted.

Honestly? I really don't! What kind of content do you think would trigger that? If humans were launching nukes based on Facebook posts we'd all be long dead! A good deep fake might trick your grandma, but it's not very likely to fool military intelligence.

> What kind of content do you think would trigger that? The kind of political propaganda that leads to the US reelecting a convicted rapist whose selects another rapist to lead the Department of Defense who then renames it to the Department of War and, true to the name, starts unilaterally attacking other countries.

If trump getting elected was due to AI, I wonder why every nation isn't electing similarly awful politicians? Hungary just elected a new president who seems a lot better than his predecessor, and a lot better than trump. The Canadian prime minister is genuinely one of the best politicians I've seen in my lifetime! The list goes on and on.

No blaming trump on anything other than the people who voted for him is like blaming school shootings on anything other than guns:a popular American passtime, and complete and utter nonsense.

Post reply on HN