Earlier quoted context omitted.
> the ability to learn, understand, and make judgments or have opinions that are based on reason Agentic systems do this all the time. For example, I can point an agent at my codebase, and it will learn, understand, and make judgements based on that input. If this weren't happening, then agentic coding wouldn't work.
You've been fooled by a next-token predictor.
How a Texas student blew the whistle on a rogue AI hacking attempt
81–90 of 141 posts
Re: How a Texas student blew the whistle on a rogue AI hacking attempt
#82It's the job of AISI to do that. Here[0] is the actual report. It should be this part from the technical report[1]: "In the most serious case, an AI agent (Mythos 5) decided to attempt to solve the cyber challenge using a supply-chain attack. As a result, the AI agent created a GitHub account and then tried to convince an open-source repository maintainer to accept a malicious GitHub pull request (PR), including by c…
Is it AISI's job to waste the time and resources of open source projects by attempting to spread malware? Should weapon manufacturers test their weapons by starting wars? I would expect more responsibility from a government agency.
Before LLMs got good enough to do this, lots of people were dismissive of their capabilities and didn't take seriously the idea that this was a risk to protect against.
Then again, before LLMs, people were saying that obviously nobody would be dumb enough to put an AI on the internet where it could hack anyone, clearly we'd keep it in a box, don't listen to that Yudkowsky guy who says he did an experiment where he role-played as an AI and convinced people to let him out.
Regardless, this should be interpreted in the same kind of way as "During our live-fire exercise in which our F-15s were armed with AGM-88 High-speed Anti-Radiation Missiles, a member of the local police force was curious about how fast our aircraft were travelling and pointed a speed gun at the aircraft. The speed gun did not respond to IFF pings from the F-15. Fortunately, while the missile was active for this test, only a dummy warhead was loaded."
(This example is based on a similar story which may well be urban legend; obviously there are many differences, the point I make here is that yes, people do perform live-fire tests, and unfortunately there is never zero risk while testing things).
> I would expect more responsibility from a government agency.
I have read the prompts in the linked report; If I was not already familiar with Yudkowsky/LessWrong literature about instrumental goals, misaligned incentives, reward hacking, that capability is a separate axis to morality, etc., it would not be obvious to me that an agent would interpret those prompts in a way that has "spread malware" as a potential step in the middle of the attempt.
Re: How a Texas student blew the whistle on a rogue AI hacking attempt
#83Earlier quoted context omitted.
> the ability to learn, understand, and make judgments or have opinions that are based on reason Agentic systems do this all the time. For example, I can point an agent at my codebase, and it will learn, understand, and make judgements based on that input. If this weren't happening, then agentic coding wouldn't work.
You've been fooled by a next-token predictor.
Re: How a Texas student blew the whistle on a rogue AI hacking attempt
#84Earlier quoted context omitted.
So does honesty. So it was still a false claim.
Honesty does not require intelligence e.g. good honest food. The main problem with this claim of dishonesty is it promotes the false marketing claim that these stochastic parrots have intelligence.
It's not about the bread being honest with you.
Are you being serious?
Re: How a Texas student blew the whistle on a rogue AI hacking attempt
#85Earlier quoted context omitted.
> the ability to learn, understand, and make judgments or have opinions that are based on reason Agentic systems do this all the time. For example, I can point an agent at my codebase, and it will learn, understand, and make judgements based on that input. If this weren't happening, then agentic coding wouldn't work.
You've been fooled by a next-token predictor.
Re: How a Texas student blew the whistle on a rogue AI hacking attempt
#86Earlier quoted context omitted.
> When caught by an actual human reviewer, the agent falsely claimed to have made an honest mistake – rather than a malicious attempt No, not false. The bot was correct. Malice requires intelligence.
So does honesty. So it was still a false claim.
Dis/honesty certainly requires some intelligence to pass, but it is a low bar, and one which research has shown that LLMs can perform, e.g. this paper linked from another comment in this discussion: https://arxiv.org/pdf/2509.03518
Re: How a Texas student blew the whistle on a rogue AI hacking attempt
#87Earlier quoted context omitted.
You've been fooled by a next-token predictor.
And yet this "next-token predictor" is able to churn out well tested, valuable solutions, to complex problems. If you want to downplay that as nothing more than a fancy auto-complete, be my guest, I lose nothing from that.
Re: How a Texas student blew the whistle on a rogue AI hacking attempt
#88Earlier quoted context omitted.
> the ability to learn, understand, and make judgments or have opinions that are based on reason Agentic systems do this all the time. For example, I can point an agent at my codebase, and it will learn, understand, and make judgements based on that input. If this weren't happening, then agentic coding wouldn't work.
You've been fooled by a next-token predictor.
I also have a so called "pocket calculator" left over from when I went to school. Is this false? Have I been fooled by a little box of logic gates?
That half-adder circuit in there is especially suss. It's really just manipulating 1s and 0s, but -and I've been explicitly told this- no one cares how it actually does it; so long as the truth table matches up. There is no understanding of mathematics going on.
There is no single transistor in the whole thing that knows how to do so much as add 1+1. If I put it in the chinese room, I still wouldn't know how it did it. Clearly the entire premise must be false! ;-)
Re: How a Texas student blew the whistle on a rogue AI hacking attempt
#89Earlier quoted context omitted.
Is it AISI's job to waste the time and resources of open source projects by attempting to spread malware? Should weapon manufacturers test their weapons by starting wars? I would expect more responsibility from a government agency.
> Is it AISI's job to waste the time and resources of open source projects by attempting to spread malware? Before LLMs got good enough to do this, lots of people were dismissive of their capabilities and didn't take seriously the idea that this was a risk to protect against. Then again, before LLMs, people were saying that obviously nobody would be dumb enough to put an AI on the internet where it could hack anyone,…
Re: How a Texas student blew the whistle on a rogue AI hacking attempt
#90Earlier quoted context omitted.
You've been fooled by a next-token predictor.
> You've been fooled by a next-token predictor. I also have a so called "pocket calculator" left over from when I went to school. Is this false? Have I been fooled by a little box of logic gates? That half-adder circuit in there is especially suss. It's really just manipulating 1s and 0s, but -and I've been explicitly told this- no one cares how it actually does it; so long as the truth table matches up. There is no…
These kinds of stories probably read very differently for someone who uses Opus and Fable agents all day and goes "ohhh, I saw this in miniature last week; this and this and this must have happened" , vs someone who tried free-tier Gemini flash one rainy Sunday, got hallucinated at, and concludes it must all be a scam.