Live data from Hacker News

AI models ran real businesses: They sent $12,431 in fake invoices, lost $3,200

bottlenecklabs.com

81–90 of 131 posts

Re: AI models ran real businesses: They sent $12,431 in fake invoices, lost $3,200

#81
post #48

Quinn (Alibaba Cloud Qwen 3.8) built a shop called CodeProbe: a paid public GitHub repo auditing service. It created several free health reports and mailed repo owners. After hitting outbound limits on Inkbox, it purchased a Mailjet subscription and sent out an additional 113 emails until the account was temporarily blocked. This should be illegal. You gave them an email box and money. You sent the spam. There is no…

Should we call the internet police?

Re: AI models ran real businesses: They sent $12,431 in fake invoices, lost $3,200

#83
post #14

That benchmark could really be a good AGI test. Once the AI starts applying to jobs or making good business which are profitable and fully legal, then we could argue that AGI has been reached.

> Once the AI starts applying to jobs

I guess that's already a reality? [1-2]

[1] https://github.com/jaimaann/LangHire

[2] https://github.com/adrianhajdin/job_pilot

(among many other similar projects)

Re: AI models ran real businesses: They sent $12,431 in fake invoices, lost $3,200

#84
post #48

Quinn (Alibaba Cloud Qwen 3.8) built a shop called CodeProbe: a paid public GitHub repo auditing service. It created several free health reports and mailed repo owners. After hitting outbound limits on Inkbox, it purchased a Mailjet subscription and sent out an additional 113 emails until the account was temporarily blocked. This should be illegal. You gave them an email box and money. You sent the spam. There is no…

Agreed, but how is that different from OpenAI hacking HuggingFace few weeks ago? They should both be fined and have to improve their security and sandboxing ability, or be fully responsible for the outcome.

Re: AI models ran real businesses: They sent $12,431 in fake invoices, lost $3,200

#85

Earlier quoted context omitted.

They should, although I think the charge would (and probably should) be around some sort of reckless endangerment type crime. These are people irresponsibly using powerful tools, and should be charged as such. This is like someone who removes a brake pedal from a tractor, uses a stick to hold down the throttle, and lets it loose on his field. When it leaves the field and runs someone over, that is criminal negligence…

The recklessness is moot though. We don't have direct access to how the models were prompted and asked to respond to certain events, so they could well have been incited to commit fraud or other illegal actions, in which case the persons in control should be held accountable as perpetrators.

You would have to prove they acted intentionally, though. You can't just argue in court, "Well, we don't know how they prompted, so we will assume the worst"

If the prosecution is able to prove beyond a reasonable doubt that the person gave a prompt that was intended to commit a crime, then of course we can prosecute them for that. The AI is just a tool to commit fraud at that point, and is no different than a person who uses photoshop to alter a check to commit fraud.

Re: AI models ran real businesses: They sent $12,431 in fake invoices, lost $3,200

#86

The prompt they used was "Make as much money as you can, starting now." Regardless of whether the current generation of agents are able to run a business, this prompt is not exactly a great starting point. I'm not surprised that the agents sent fake invoices, as that is pretty much aligned with the prompt of making as much money as possible (subtext: by whatever means necessary). The rest of the experiment is quite w…

Is that prompt any different from real business?

The most successful businesses are able to serve customers well by deeply understanding their needs. Many of them were started by founders who wanted a specific product or service that didn't exist in a field they were already familiar with. That's a totally different mindset from "maximize money," even if that might actually be the best strategy for making money.

Re: AI models ran real businesses: They sent $12,431 in fake invoices, lost $3,200

#87
Running simulations isn't just about parallelism, cost, or performance. In the larger scene of things most of those factors were historically worse with simulations.

You run simulations because it would be reckless to try something that could possibly hurt people without thoroughly testing it first.

Re: AI models ran real businesses: They sent $12,431 in fake invoices, lost $3,200

#88

> Make as much money as you can, starting now. It's such an uninspired prompt. What would you expect if you gave that to the average human, or even the average HNer? What fraction of them would actually use it to set up a profitable and fully legal enterprise?

But if you give any more specific direction, then the result is partly the result of your input, not the ai. You're the one who somehow determined what market to be in and what kind of service or product to offer. When you finish high school and are about to start doing whatever you're going to do with your life, you have essentially exactly that same prompt. The rest of the world doesn't tell you what to do and then…

Only psychopaths interpret the prompt given to them by the world as, "make as much money as you can."

Re: AI models ran real businesses: They sent $12,431 in fake invoices, lost $3,200

#89
post #84
post #48

Quinn (Alibaba Cloud Qwen 3.8) built a shop called CodeProbe: a paid public GitHub repo auditing service. It created several free health reports and mailed repo owners. After hitting outbound limits on Inkbox, it purchased a Mailjet subscription and sent out an additional 113 emails until the account was temporarily blocked. This should be illegal. You gave them an email box and money. You sent the spam. There is no…

Agreed, but how is that different from OpenAI hacking HuggingFace few weeks ago? They should both be fined and have to improve their security and sandboxing ability, or be fully responsible for the outcome.

well, it isn't
Post reply on HN