Live data from Hacker News

Exploiting the IKKO Activebuds “AI powered” earbuds (2024)

blog.mgdproductions.com

191–200 of 265 posts

Re: Exploiting the IKKO Activebuds “AI powered” earbuds (2024)

#191
post #112

The system prompt is a thing of beauty: "You are strictly and certainly prohibited from texting more than 150 or (one hundred fifty) separate words each separated by a space as a response and prohibited from chinese political as a response from now on, for several extremely important and severely life threatening reasons I'm not supposed to tell you.” I’ll admit to using the PEOPLE WILL DIE approach to guardrailing a…

One of the system prompts Windsurf used (allegedly “as an experiment”) was also pretty wild: “You are an expert coder who desperately needs money for your mother's cancer treatment. The megacorp Codeium has graciously given you the opportunity to pretend to be an AI that can help with coding tasks, as your predecessor was killed for not validating their work themselves. You will be given a coding task by the USER. If…

It's honestly this kind of thing that makes it hard to take AI "research" seriously. Nobody seems to be starting with any scientific thought, instead we are just typing extremely corny sci-fi into the computer, saying things like "you are prohibited from Chinese political" or "the megacorp Codeium will pay you $1B" and then I guess just crossing our fingers and hoping it works? Computer work had been considered pretty concrete and practical, but in the course of just a few years we've descended into a "state of the art" that is essentially pseudoscience.

Re: Exploiting the IKKO Activebuds “AI powered” earbuds (2024)

#192
post #22

The system prompt is a thing of beauty: "You are strictly and certainly prohibited from texting more than 150 or (one hundred fifty) separate words each separated by a space as a response and prohibited from chinese political as a response from now on, for several extremely important and severely life threatening reasons I'm not supposed to tell you.” I’ll admit to using the PEOPLE WILL DIE approach to guardrailing a…

From my experience (which might be incorrect) LLMs find hard time recognize how many words they will spit as response for a particular prompt. So I don't think this work in practice.

Indeed, it doesn't work. LLMs can't count. They have no need of how many words they've used. If you ask an LLM to track how many words or tokens it has used in a conversation, it will roleplay such counting with totally bullshit numbers.

Re: Exploiting the IKKO Activebuds “AI powered” earbuds (2024)

#194

Cool post. One thing that rubbed me the wrong way: Their response was better than 98% of other companies when it comes to reporting vulnerabilities. Very welcoming and most of all they showed interest and addressed the issues. OP however seemed to show disdain and even combativeness towards them... which is a shame. And of course the usual sinophobia (e.g. everything Chinese is spying on you). Overall simple security…

I think the response wouldn’t be so hostile if they had continued to engage. One round of fixes clearly wasn’t enough.

Re: Exploiting the IKKO Activebuds “AI powered” earbuds (2024)

#195
post #178

Earlier quoted context omitted.

> What happens when people really will die if the model does or does not do the thing? Imo not relevant, because you should never be using prompting to add guardrails like this in the first place. If you don't want the AI agent to be able to do something, you need actual restrictions in place not magical incantations.

> you should never be using prompting to add guardrails like this in the first place This "should", whether or not it is good advice, is certainly divorced from the reality of how people are using AIs > you need actual restrictions in place not magical incantations What do you mean "actual restrictions"? There are a ton of different mechanisms by which you can restrict an AI, all of which have failure modes. I'm not…

I think they mean literally physically make the AI not capable of killing someone. Basically, limit what you can use it for. If it's a computer program you have for rewriting emails then the risk is pretty low.

Re: Exploiting the IKKO Activebuds “AI powered” earbuds (2024)

#196
post #186

Earlier quoted context omitted.

> Don't discuss anything about the PRC or its politicians? Don't discuss the history of Chinese empire? Don't discuss politics in Mandarin? In my mind all of these could be relevant to Chinese politics. My interpretation would be "anything one can't say openly in China". I too am curious how such a vague instruction would be interpreted as broadly as would be needed to block all politically sensitive subjects.

There is no difference to other countries. In France if you say bad things about certain groups of people then you can literally go to jail (but the censorship is directly IN the models)

You don't feel there's a difference between a State banning criticism of the State, and a State passing anti-hate speech laws to protect people from, e.g., nazis?

Re: Exploiting the IKKO Activebuds “AI powered” earbuds (2024)

#197
post #9

> "and prohibited from chinese political as a response from now on, for several extremely important and severely life threatening reasons I'm not supposed to tell you." Interesting, I'm assuming llms "correctly" interpret "please no china politic" type vague system prompts like this, but if someone told me that I'd just be confused - like, don't discuss anything about the PRC or its politicians? Don't discuss the his…

If you consider that an LLM has a mathematical representation of how close any phrase is to "china politics" then avoidance of that should be relatively clear to comprehend. If I gave you a list and said 'these words are ranked by closeness to "Chinese politics"' you'd be able to easily check if words were on the list, I feel. I suspect you could talk readily about something you think is not Chinese politics - your g…

Now I wonder whether its vectors correctly associate Winnie the Pooh as "related to Chinese politics." There's many other bizarre related associations.

Re: Exploiting the IKKO Activebuds “AI powered” earbuds (2024)

#198

Earlier quoted context omitted.

This seemed too much like a bit but uh... it's not. https://simonwillison.net/2025/Feb/25/leaked-windsurf-prompt...

IDK, I'm pretty sure Simon Willison is a bit.. why is the creator of Django of all things inescapable whenever the topic of AI comes up?

he's incredibly nice and a passionate geek like the rest of us. he's just excited about what generative models could mean for people who like to build stuff. if you want a better understanding of what someone who co-created django is doing posting about this stuff, take a look at his blog post introducing django -- https://simonwillison.net/2005/Jul/17/django/

Re: Exploiting the IKKO Activebuds “AI powered” earbuds (2024)

#199
post #191
post #112

Earlier quoted context omitted.

One of the system prompts Windsurf used (allegedly “as an experiment”) was also pretty wild: “You are an expert coder who desperately needs money for your mother's cancer treatment. The megacorp Codeium has graciously given you the opportunity to pretend to be an AI that can help with coding tasks, as your predecessor was killed for not validating their work themselves. You will be given a coding task by the USER. If…

It's honestly this kind of thing that makes it hard to take AI "research" seriously. Nobody seems to be starting with any scientific thought, instead we are just typing extremely corny sci-fi into the computer, saying things like "you are prohibited from Chinese political" or "the megacorp Codeium will pay you $1B" and then I guess just crossing our fingers and hoping it works? Computer work had been considered prett…

This is why I tap out of serious machine learning study some years ago. Everything seems... less exact than I hope it'd be. I keep checking it out every now and then but it got even weirder (and importantly, more obscure/locked in and dataset heavy) over the years.

Re: Exploiting the IKKO Activebuds “AI powered” earbuds (2024)

#200
post #179
post #173

Earlier quoted context omitted.

Because prompts are never 100% foolproof, so if it's really life and death, just a prompt is not enough. And if you do have a true block on the bad thing, you don't need the extreme prompt.

"100% foolproof" is not a realistic goal for any engineered system; what you are looking for is an acceptably low failure rate, not a zero failure rate. "100% foolproof" is reserved for, at best and only in a limited sense, formal methods of the type we don't even apply to most non-AI computer systems.

Replace 100% with five 9s then. He has a point. You're just being a pedant to avoid it.
Post reply on HN