Live data from Hacker News

Our newsroom AI policy

arstechnica.com

101–110 of 144 posts

Re: Our newsroom AI policy

#101
post #31

Earlier quoted context omitted.

The LLM can find material that it would be hard or time-consuming for you to do. You still need to verify it, but "find the right things to read in the first place" is often a time intensive process in itself. (You might, at that point, argue that "what if LLM fails to find a key article/paper/whatever", which I think is both a reasonable worry, and an unreasonable standard to apply. "What if your google search doesn…

I believe what their point is is that if you give people a "extract-needle-from-haystack" machine and then tell them they have to manually find where in the haystack the needle was, it defeats the purpose of having the machine. With that said, a good RAG solution would come with metadata to point to where it was sourced from.

This is a bogus analogy leaidng to a bogus conclusion.

If something points to the needle in the haytack (saying "this haystack has a needle positioned eighteen centimeters from the top and three left of center"), it's much easier to verify that indeed there is a needle there than it would be to find that needle in the first place.

If an LLM spits out a claim that something happened (citing a certain article), it's less work to read the article and verify the claim than it would be to DISCOVER the article in the first place.

In other words, LLMs can be a time-saving search engine, and the idea that it's just as much work to find+verify information as it is to have the LLM find it and then you verify it is hokum.

Re: Our newsroom AI policy

#102

Earlier quoted context omitted.

At this point, it feels like most technology will be used in favor of people with power, and not in a democratizing manner. I'd argue that this is something that is more about the state of play, than tech itself.

> I'd argue that this is something that is more about the state of play, than tech itself. What do you mean by that? It seems inherent to the technology under capitalism: it allows a flood of slop and anything public and valuable will be plundered, so the incentive is to make valuable stuff exclusive and elite.

I mean that

> inherent to the technology

Vs

> inherent to the technology under capitalism

The TLDR of my point is going to be that wealth concentration and information pollution sets up economies that don’t work for us in a manner that is healthy for us.

Re: Our newsroom AI policy

#103

Self-contradictory policy. > Reporters may use AI tools vetted and approved for our workflow to assist with research, including navigating large volumes of material, summarizing background documents, and searching datasets. If this is their official policy, Ars Technica bears as much responsibility as the author they fired for the fabricated reporting. LLMs are terrible at accurately summarizing anything. They very r…

I have a little bit of a bias here, as I am building forth.news, which is an AI-powered news platform -- but I am also a former journalist.

It's not necessarily contradictory. I see this more like giving your employees cars, but telling them they are responsible if they get into accidents.

All of this is entirely predicated on expectation and responsibility. First, mark something as being AI if it cannot be verified, and verify everything that you can.

Forth is using AI so we can detect and push out stories as quickly as possible, getting breaking news out there as soon as it breaks. Our summaries are AI, but marked as AI. Our underlying source information is right there and cited. We try to be as transparent as possible about the tools we are using, and the tradeoffs.

Every journalist should instinctively and reflexively double check everything, regardless of the source. There's an old maxim, "if your mom tells you she loves you, check it out." Being from an LLM doesn't change that.

Re: Our newsroom AI policy

#105
post #92
post #75

Earlier quoted context omitted.

> Paid out like Spotify pays out artists. As others said, Spotify pays shit for artists, but maybe that's the problem with the whole thing here. It should be more like how Bandcamp pays artists (80% to the artists, 20% for Bandcamp), but then the rapacious economy supporting the largest LLM providers would collapse and (wipes away a single tear) we'd all have to use simpler, cheaper, most likely local models.

“Since Spotify pays out two-thirds of all music revenue to the industry – almost 70% of what we take in – as Spotify revenues grow, music payouts have grown as well. “ https://newsroom.spotify.com/2026-01-28/2025-music-industry-... That’s not that far off from 80%.

I think people get distracted by the "percentage of revenue paid to musicians" thing, when the bigger reason streaming pays out so little to artists is that people pay $10-$15 per month for unlimited access to all music. Even 80% of that, split across dozens or hundreds of musicians, is not very much. Of course, it's also worth remembering that streaming was partially a response to widespread piracy. It's difficult to get people to pay very much at scale for easily copied digital media.

In addition, a greater share of the payout (relative to number of streams) goes to big music distributors that control the biggest, most popular artists and have the leverage and employees to negotiate those agreements.

Re: Our newsroom AI policy

#106

Self-contradictory policy. > Reporters may use AI tools vetted and approved for our workflow to assist with research, including navigating large volumes of material, summarizing background documents, and searching datasets. If this is their official policy, Ars Technica bears as much responsibility as the author they fired for the fabricated reporting. LLMs are terrible at accurately summarizing anything. They very r…

> You cannot permit your employees to use LLMs in this manner and then tell them it's entirely their fault when it makes mistakes, because you gave them permission to use something that will make mistakes 100% without fail. Yes you can. The same way Wikipedia (or, way back when, a paper encyclopedia) can be used for research but you have to verify everything with other sources because it is known there are errors and…

> The same way Wikipedia can be used for research

Before LLMs, Wikipedia was the greatest source of disinformation in human history. No journalist should ever have been using it for research. At best it's a fun project for satisfying people's idle curiosity where the truth of what they read doesn't really matter, but if your job is to report factual information, reading Wikipedia is doing a disservice to yourself and your readers. Just like people don't properly verify the BS LLMs fabricate, very few people thoroughly read the citations on Wikipedia, which often involves purchasing books and getting access to research papers. If they did read citations, they would realise that Wikipedia citations are all too frequently unsupported by the actual material they're citing, or in some cases, the cited material establishes the exact opposite. This is to say nothing of cherry-picking sources, of course.

> Should they also ban them from talking to people as sources of information, because people can be misinformed or actively lie, rather than instead insisting that information found from such sources be sense-checked before use in an article?

Statements made by people are attributed to those people to account for this. Rather than saying "Company X's product is the safest product ever made", a journalist says "Company X's CEO claims their product is the safest ever made". People do not do this with LLMs or Wikipedia, rather than attributing it to an understood-to-be-unreliable source they just present it as a factual statement. Also, if the journalist has good reason to believe the quoted statement is false, it is in fact journalistic malpractice to cite the quote wholesale without caveats informing the reader of the evidence that the quoted person is trying to mislead them.

> because everywhere where potentially dangerous equipment is actually made available for someone's job you will find policies exactly like this

Which maybe makes sense when the flamethrowers are a necessary part of the job. Flamethrowers are not necessary for journalism, full stop, so the fault rests with the organization introducing a dangerous tool into the work environment unnecessarily.

Re: Our newsroom AI policy

#107
post #53

AI is in danger of peeing in it's own water source. It's unbelievably useful at imitating and generating content, but it needs enough original content to be able to train and scrape. Google got one thing wrong and nearly destroyed the internet - people need to have an incentive to contribute content online, and that incentive should not be to game the system for advertising. This in particular dawned on me when askin…

> Paid out like Spotify pays out artists. So, mostly to fraudulent AI spam? AI makes this problem worse in both directions. It makes it fantastically easy to produce ""content"". So if you're scraping content, or browsing content, you're going to run in to increasing amounts of AI. Micropayments makes this worse , because it's then a means of getting paid to produce spam. The problem comes when you want the ""content…

I worked in music streaming for several years. Yes, there is spam, but in my experience this was less than 1% of total consumption, even if now it is a huge share of available content (also a lot of it seems to be mostly for money laundering). Also, the share of revenue that Spotify and the other services pass on to rights holders is roughly on the scale of old brick and mortar retail. But how people spend has changed. Indie music nerds used to spend much more than the average mainstream listener on records and CDs. Under streaming, both mostly pay the same subscription price, so enthusiasts spend, while casual listeners spend more. On streaming platforms payouts are tied to streaming consumption not purchases, so music with strong branding, playlist support, and promotional backing does well, and the major labels are good at that.

What share of what Spotify pays out makes it's way into the pockets of song writers and musicians is a more complicated story, generally more if the artists are with a good indie label, generally less if they are with a major. At the same time, majors have had to offer less abusive deals than they used to, because DIY and indie distribution more viable.

The other big shift is that in the retail days new releases drove most purcahse, but with streaming catalog is a source of reliable recurring revenue, and the majors own a lot of catalogue, especially stuff they acquired outright in an era when artists often had their work basically stolen from them.

The key difference between Spotify and LLMs scraping the open internet is provenance. Music on Spotify does not just appear there out of nowhere. It arrives through an accountable chain: a label, a distributor, an aggregator, a publisher, a rights holder. Sometimes this chain is thin, like with self-serve, pay to publish distribution through companies like CD Baby. Most of what is actually streamed has a provenance that reflects serious editorial and financial commitment by an organisation in the form of money spent recording, developing, and promoting an artist. This provenance chain is critical contextual information about who vouched for the work, who invested in it, who holds rights to it, and when it entered the culture. Art, music, writing do not exist in a vacuum. They are part of an ongoing cultural conversation, and who said what, when, and under what institutional backing is integral to its meaning.

So I share OP's hope the long-run equilibrium for LLMs looks more like licensed media than scraping and open web search. I want a world where models license published content from rights holders, not for training, though that would be nice, but to surface answers with links to identifiable sources in a verifiable published database, and let part of my subscription pay for access to the underlying referenced material. Information is valuable, and it's reasonable to pay for it. Aligning incentives around truth is the challenge.

Putting ink on paper and moving books around is the least important part of what a publisher does. The important part is selection, investment, positioning, promotion, and accountability. This curatorial function has always been important, and it can only become more important the tsunmai of ai slop and misinformation grows. I hope that chatbot manufacturers partner responsibly with rights holders and lean into the value that publishers have created instead of potentially destroying it.

Re: Our newsroom AI policy

#108

Self-contradictory policy. > Reporters may use AI tools vetted and approved for our workflow to assist with research, including navigating large volumes of material, summarizing background documents, and searching datasets. If this is their official policy, Ars Technica bears as much responsibility as the author they fired for the fabricated reporting. LLMs are terrible at accurately summarizing anything. They very r…

> ...this is equivalent to providing all of your employees with a flamethrower, and then saying they bear all responsibility for the fires they start.

This is essentially the policy of most SWE groups but with PRs merged, though right?

You can / should use AI to accelerate SWE workflows and assist in reviews but if you merge something that is bad or breaks production that is on you.

> "Hey, don't blame us for giving them flamethrowers, it's company policy not to burn everything to the ground!".

Flamethrowers are inherently dangerous to the operator and are ~intended to be used burn things to the ground.

I'm no expert on arms but there is probably a simile with a better fit out there.

Re: Our newsroom AI policy

#109

Earlier quoted context omitted.

> You cannot permit your employees to use LLMs in this manner and then tell them it's entirely their fault when it makes mistakes, because you gave them permission to use something that will make mistakes 100% without fail. Yes you can. The same way Wikipedia (or, way back when, a paper encyclopedia) can be used for research but you have to verify everything with other sources because it is known there are errors and…

> Yes you can. The same way Wikipedia (or, way back when, a paper encyclopedia) can be used for research but you have to verify everything with other sources because it is known there are errors and deficiencies in such sources. I think that if Wikipedia had no recommendations on good sources for their own articles and did not ever ban sources, companies would not be so sanguine about letting people use Wikipedia. Th…

> companies would not be so sanguine about letting people use Wikipedia

Are companies sanguine about using Wikipedia without verification. Maybe some, but they darn well shouldn't be. And I say this as someone who uses Wikipedia for many minor things (though for anything important, I verify elsewhere).

> Also, I do think that there are companies that do have policies against talking to known liars.

No doubt most/all. But such policies will always be caviated with exceptions if the information is properly validated afterwards.

> So allowing LLM use at all is a direct admission that seeking out the "truth" is not an important goal because it could never actually improve accuracy and could only worsen it through hallucinated, probable reporting.

I'm generaly anti-LLM, but this is… ad absurdum.

There is a huge difference between lazily accepting what an LLM spews out, and using that along with other sources for further research. No good reporter will trust a single source away from exceptional circumstances, wether that source is a person or an LLM, and what would be considered “exceptional circumstances” for trusting specific meat-sourced information won't apply for an LLM-sourced summary.

If you can trust Wikipedia as a starting point, you can trust a good LLM as a starting point. Both are offering a summary of what a bunch of people on the internet have written, neither should be trusted as a reliable source.

> I don't think it's insane to then have such companies or agencies say that AI shouldn't be used because it's been shown to be unreliable

If taking an absolutist approach. I would be a little more quakified and say that LLM output should never be used without verification of all details, rather than should not be used at all. It may be the case that this verification makes using LLMs no more efficient than doing research from other sources in the first place, and I suspect that this is the case often, if when using LLMs proper time is given to verifying the output.

The problem is people musunderstanding what an LLM is: a summariser, offering access to a compressed version of its sources. If you are using them as sources rather tham summarisers then you are using them wrongly. Unfortunately, that means a great many people are using them wrongly…

Post reply on HN