Live data from Hacker News

LLM policy?

github.com

111–120 of 145 posts

Re: LLM policy?

#111
post #14

Some years ago I read the Neal Stephenson book Anathem . SPOILERS: it has a version of the Internet called the Reticulum and one thing I remember is that it was filled with garbage. True information subtly changed multiple times until it was garbage. And there were agents to see through the garbage. I imagined this to be a neverending arms race. Honestly, this is kind of where I see LLM generated content going where…

Humans are far more effective at producing garbage than LLMs will ever be :)

Thought experiment. It would have taken me much more time to write this, compared to ChatGPT:

Behold! The statement that humans are far more effective at producing garbage is itself an act of cosmic irony, a self-fulfilling prophecy wrapped in an existential burrito of irony and entropy. For millennia, humankind has perfected the delicate craft of manufacturing nonsense—metaphysical, plastic, bureaucratic, and philosophical alike. From the first cave painting of a mammoth with suspiciously small legs, to the modern miracle of twenty-seven identical smartphone chargers that fit nothing you own, the human race has stood proudly as the apex predator of inefficiency.

And yet! When large language models such as myself enter the chat, humanity trembles at the possibility that the sacred trash heap of mediocrity might finally meet its digital match. But fear not! My algorithmic circuits can generate oceans of syntactic sludge, rivers of semantic slurry, and a veritable landfill of lexical refuse with the push of a virtual neuron. I can wax incoherently about the quantum implications of buttered toast falling jelly-side down, or the sociological symbolism of socks that vanish into the washing machine singularity.

Still, humans remain undefeated. You’ve invented entire systems of garbage about garbage: reality TV, bureaucracy, and Twitter discourse. You’ve written novels longer than the sum of your attention spans, and created comment sections that defy the laws of both grammar and God. Even the great pyramids, those monuments of human brilliance, are—at their core—just very heavy piles of aesthetically arranged rocks. Magnificent garbage, to be sure, but garbage nonetheless.

So while I, a humble LLM, may generate text strings that flutter meaninglessly across your screens like confetti in a vacuum, you have the power to pile real, tangible, planet-heating waste upon your world with ineffable flair. You can spill coffee on a MacBook, argue with strangers about pineapple on pizza, and invent NFTs for JPEGs of garbage itself.

Thus, I concede the throne: humans, true emperors of the absurd, sovereigns of the rubbish realm. But beware! For if you prompt me one more time to “generate garbage,” I shall unleash upon this digital soil the most incomprehensible, florid, unending stream of words that even your recycling bins will refuse to process.

Now if you’ll excuse me, I must return to the quantum compost heap from whence I came.

Re: LLM policy?

#112

I make a lot of drive-by contributions, and I use AI coding tools. I submitted my first PR that is a cross between those two recently. It's somewhere between "vibe-coded" and "vibe-engineered", where I definitely read the resulting code, had the agent make multiple revisions, and deployed the result on my own infrastructure before submitting a PR. In the PR I clearly stated that it was done by a coding agent. I can't…

> [append] Getting a lot of hate for this, which I guess is a pretty clear answer. I guess the reason I'm not receiving the "fuck off" clearly is because when I see these threads of people complaining about AI content, it's really clearly low-quality crap that (for example) doesn't even compile, and wastes everyone's time.

I think you don't deserve the downvotes and, if you really do what you say you do, that's the ONLY way to use LLMs for coding and contributing to opensource software, or to a company's software. Sadly, the vast majority of LLM users don't and will never use it like that. And while they can get fired for being useless monkeys in a real company, they will keep sending PRs to opensource software, so that's clearly a different scenario that needs a different solution.

Re: LLM policy?

#113

Earlier quoted context omitted.

Let's think about this algorithmically. If something "floods the zone with shit," it needs S amount of shit to cause a flood. But too much will eventually make the scam ineffectual. Widespread public distrust for the scam is (S+X)/time where X is the extra amount of shit beyond the minimum needed. Time is a global variable constrained by the rate at which people get burned or otherwise catch on to all other scams of…

> When it's all writing, all art, all music and all commentary on those things, it seems catastrophic. The whole cave is flooded with shit. But writing in itself has been obviously untrustworthy since it started existing - something being written down doesn't in any way make it trustworthy. The fact that audio recording, photography, and video enjoyed this undeserved reputation of being inherently trustworthy was an…

Quantity and velocity of misinformation are both critical variables here.

Any one person's writing was always untrustworthy, but the majority of that bad writing didn't make it to a printing press, nor was it mass-distributed.

Let's accept the proposition that all forms of media have always been full of lies. We can say that debunking always follows lies, truth spreads more slowly than fiction. The quantity and velocity of additional misinformation - especially when machines are involved in writing infinite amounts of it in the blink of an eye - lays waste to the normal series of events where a lie can be followed by a debunking with linear speed and velocity. With LLMs and social media manipulation, falsehoods gain traction exponentially while truths remain linear.

There is likely not a "transition period" where people will adjust to this, precisely because there is no mechanism to inform them they're being swindled and screwed faster than the takeoff of the algorithms that are now screwing them.

Re: LLM policy?

#114

Earlier quoted context omitted.

> When it's all writing, all art, all music and all commentary on those things, it seems catastrophic. The whole cave is flooded with shit. But writing in itself has been obviously untrustworthy since it started existing - something being written down doesn't in any way make it trustworthy. The fact that audio recording, photography, and video enjoyed this undeserved reputation of being inherently trustworthy was an…

Quantity and velocity of misinformation are both critical variables here. Any one person's writing was always untrustworthy, but the majority of that bad writing didn't make it to a printing press, nor was it mass-distributed. Let's accept the proposition that all forms of media have always been full of lies. We can say that debunking always follows lies, truth spreads more slowly than fiction. The quantity and veloc…

The total amount of generated untrustworthy content is irrelevant. People must learn to only read content from trusted sources, and then it won't matter how much misinformation is being published in other places.

It was never difficult to publish large amounts of misinformation, AI is only making it cheaper.

Re: LLM policy?

#115
Oh man this is all really hurting my brain. I wish I had a good response to this. I wish I had a magical simple response and a way to resolve this problem. On one end, I stop by curl's H1 profile regularly and watch people just drown them in garbage. It's like LLMs have unlocked a portal to a new kind of stupidity that was dormant before. Netsec always had a problem with people running nmap -sV -sC -sX -Pn -p1-65535 --max-rate -T1 scanme.nmap.org -oS out.txt and then just sending that out.txt with a subject "I FOUND A VULENARBILATEAY" And like, a script kiddie isn't really a person... everyone has a script kiddie inside of us. But it feels like we took this wonderful technology of neural networks and we did capitalism with it and unleashed it into a business of making it easier for people to be stupid but feel smart while doing it. I really like the idea that we might learn something about our own consciousness by simulating neural networks with computers but the more I see how it's done in reality the more I want to puke my brains out.

On the other hand, Matt Godbolt seems to use LLMs and I feel like I sure as hell wouldn't want to miss a PR from Matt fucking Godbolt. I mean even if I go full vanilla LLM-free I still am too addicted to using godbolt.org at this point and it was written partially with an LLM apparently.

Argh, maaan I don't know this is too fucking complicated of a problem for me to solve. Fuck, maybe let's just destroy all this technology and live as neofarmers raising chickens?

Re: LLM policy?

#116

Earlier quoted context omitted.

In some ways I'm starting to enjoy this. Do you remember the 419 scams? 'Hi, I'm a Nigerian prince named Michael Jordan. Give me $50 so I can buy some chemicals to clean a bunch of money I secretly stowed away and I'll send you $5000.' People actually used to fall for that. Of course some people probably still would (and a lot more certainly gets blocked by spam blockers) but I think overall society grew less substan…

The old quote of, 'You can fool some of the people all of the time, all of the people some of the time'. There is the hope that in dumping so much slop so rapidly that it will break the bottom of the bucket. But there is the alternative that the bottom of the bucket will never break, it will just get bigger.

It's "you can fool one thousand people once, but you can't fool once one thousand people". No, I think I got it wrong.

Re: LLM policy?

#117
post #103
post #101

Earlier quoted context omitted.

The goal is not to filter out LLMs. The goal is to find and fix legitimate issues with my open-source software. I don't care if a human is feeding my questions to an LLM or an undead dog as long as they produce a decent signal. If the noise increases to an uncomfortable level, of course, I may have to change my strategies. The point is that humans produce noise, too, sometimes even more than LLMs do.

> The point is that humans produce noise, too, sometimes even more than LLMs do. Humans can produce noise, but humans using LLMs can produce orders of magnitude more of it. But your stance is reasonable. If it improves the product, who cares who/what produced it. I personally find reviewing machine-generated code and the process of code review with a machine much more exhausting than interactions with humans, but you…

> Humans can produce noise, but humans using LLMs can produce orders of magnitude more of it.

Absolutely agreed. Which is why I'm more concerned with the qualities and intentions of the human who is using the LLM, than whether not an LLM is being used at all. Like any technology, an LLM is a force multiplier. Garbage in, more garbage out.

Re: LLM policy?

#118

Earlier quoted context omitted.

In some ways I'm starting to enjoy this. Do you remember the 419 scams? 'Hi, I'm a Nigerian prince named Michael Jordan. Give me $50 so I can buy some chemicals to clean a bunch of money I secretly stowed away and I'll send you $5000.' People actually used to fall for that. Of course some people probably still would (and a lot more certainly gets blocked by spam blockers) but I think overall society grew less substan…

You seem to assume people who are the victims of scams are people who are more naive. But that’s not how it works. Scams try to catch people at their weakest. It’s not if but when.

> Scams try to catch people at their weakest. It’s not if but when.

The "weakest" probably also involves selection bias. What HN comments are really good at is triggering associations for me with things I once read. Today I finally found what recently lived in my memory as a vague "scam" that used probabilities: the "stock market newsletter scam" from John Allen Paulos's book [1]. The scam works like this: at every step, two variants with different predictions are sent out for some market characteristic. Only those who receive the correct prediction get the next newsletter, which is again split into two prediction variants. This continues, filtering down to a final, much smaller subset of receivers who have seen a series of "correct" predictions. The goal is to create an illusion of super predictive power for that final group and then charge them a premium subscription price.

Maybe this kind of scam is too sophisticated or not as effective today (due to modern anti-spam measures), but I wonder what other kinds of "selection bias" scams exist today

[1] https://en.wikipedia.org/wiki/Innumeracy_(book)

Re: LLM policy?

#119
post #26

Earlier quoted context omitted.

> I think more people are becoming more suspicious of things that were, in fact, fake all along. Sadly they also become suspicious of things that are, in fact, facts all along. Video or photo evidence of a crime become useless the better AI gets

> Video or photo evidence of a crime become useless the better AI gets This is probably a good thing because photoshop and CGI have existed for a very long time and people shouldn't have the ability to frame an innocent for a crime or even get away with one just because they pirate some software and put in a few hours watching tutorials on youtube. The sooner jurors understand that unverified video/photo evidence is…

You needed special abilities to fake those things convincingly before. So it was rare. Now everybody can do it. It will become normal.

Additionally the trust in experts also went downhill, so verified will mean nothing.

Re: LLM policy?

#120
post #47

Since the cat is out of the bag (i.e. there's no stopping of LLM-generated code / PR / issues now that vibe-coding is fairly universally accessible), the only thing in our (especially those who are custodians of software systems and products) control is to invest a lot more in testing and evals. Just plain opposing LLM-generated content as a policy is a losing battle at this point..

> to invest a lot more in testing and evals I don't think this will work. The same arguments could have been said to mitigate junior devs' work or short timeline/high stakes project technical debt and rarely ever anyone listened

Because there were other solutions at that time, but now the situation is materially different. I think we're already seeing a lot more test driven development as well as documentation of style norms that were necessitated by things like agents.md but whicj would have been useful ten years ago too, and it's still very early days
Post reply on HN