Live data from Hacker News

Certified 100% AI-free organic content

substack.piszek.com

101–110 of 275 posts

Re: Certified 100% AI-free organic content

#101
post #4

There is so much stuffing for a simple idea that I'm not sure if this piece deserves its own title, but I'll give it the benefit of the doubt. One thing that I wonder though is how we will draw the line. If I'm writing a piece and do a Google search, and in that way invoke BERT under the hood, is anything that I write afterwards "AI-tainted"? What about the grammar checker? Or the spot removal tool in photoshop or gi…

>One thing that I wonder though is how we will draw the line. If I'm writing a piece and do a Google search, and in that way invoke BERT under the hood, is anything that I write afterwards "AI-tainted"? What about the grammar checker? Or the spot removal tool in photoshop or gimp? Or the AI voice that reads back to me my own article so that I can find prose issues?

>And that brings the other problem: do the general public really know the extent of AI use today, never mind in the future?

The line is drawn at human ownership/responsibility. A piece of content can be 'AI tainted' or '100% produced by AI', what makes the difference is if a human takes the responsibility of the end product or not.

Re: Certified 100% AI-free organic content

#102

Earlier quoted context omitted.

> The long term impact of the ease of generating low nutrition digital content using language models may be that people put down their devices The problem is people don't always make the wise decision. Evidence: the junk food industry is alive and kicking. Some people will disconnect from devices, but others may just say "this is the way things are now" and adjust themselves to the flavor of junk content.

Why are you assuming that it will make writing worse not better? Just because it can be used by non-experts to create crappy written work, it can also be used by people who work with it to augment and improve their existing writen work. To my mind AI is a general purpose technology: https://en.wikipedia.org/wiki/General-purpose_technology I guess using this mental model, what you are worried about is the equivilent t…

> Did the printing press also increase the amount of crap in circulation?

I think the answer is, undoubtedly, yes.

Re: Certified 100% AI-free organic content

#104
post #7
post #4

There is so much stuffing for a simple idea that I'm not sure if this piece deserves its own title, but I'll give it the benefit of the doubt. One thing that I wonder though is how we will draw the line. If I'm writing a piece and do a Google search, and in that way invoke BERT under the hood, is anything that I write afterwards "AI-tainted"? What about the grammar checker? Or the spot removal tool in photoshop or gi…

> There is so much stuffing for a simple idea that I'm not sure if this piece deserves its own title, but I'll give it the benefit of the doubt. Frankly I had the same thought writing it :D It's more of a stake in the ground sort of a thing I guess? What I really want is somebody saying "hey, there is an open standard already here" so I can use it.

The idea has some legs, but they are weak for the many reasons pointed out to me by fair criticism of "digital veganism". The main one is that labelling is one small part of quality. Tijmen Schep in his 2016 "Design My Privacy" [1] proposed some really cool ideas around quality and trustworthiness labelling of IoT/mobile devices, but ran into the same issues. Responsibility ultimately lies with the consumer, and so long as consumers remain uneducated as to why low quality is harmful, and cannot verify the provenance of what they consume or the harmful effects, nothing will change.

Right now we seem to be at the stage of "It's just McDonald's/KFC for data - junk food is convenient, cheap and not a problem - therefore mass production generative content won't be a problem".

The food analogy is powerful, but has limits, and I urge you to dig into Digital Vegan [2] if you want to take it further.

[1] https://www.tijmenschep.com/design-my-privacy/

[2] https://digitalvegan.net

Re: Certified 100% AI-free organic content

#105

Earlier quoted context omitted.

Does use of a search engine violate the "No AI" covenant with oneself? Variation on the Turning Test: prove that it's not a human claiming to be a computer. Modeling premises and Meta-analysis are again necessary elements for critical reasoning about Sources and Methods and superpositions of Ignorance and Malice.

Maybe this could encourage the recreation of the original Yahoo! (If you don't remember, Yahoo! started out not as a search engine in the Google sense but as a collection of human curated links to websites about various topics)

I consider Wikipedia to be a massive curated set of information. It also includes a lot of references and links to additional good information / source materials. Companies try to get spin added and it's usually very well controlled. I worry that a lot of ai generated dreck will seep into Wikipedia, but I am hopeful the moderation will continue to function well.

Re: Certified 100% AI-free organic content

#106
You'll need a special AI soon just to read and process the flood of superficially reasonable sounding bullshit AI will be throwing at you from all kinds of directions. On the bright side, the quality of writing style might overall increase with AI paraphrasing and style correction/adaptation tools.

Re: Certified 100% AI-free organic content

#107

It's not clear that people actually care about and want AI-free OC, at least if you look at what kind of content is being consumed. Right now Google search seems to prioritize non-organic content, with search results often being a stream of blogspam and reddit/quora shill/astroturf crap that if not AI-generated, is close enough in terms of tone, accuracy, and originality, that it might as well be. Meanwhile you never…

4chan constantly deletes its own content. There are a limited number of slots for threads, and making a new one causes the oldest one to disappear. Google does not care much for dead links. Additionally, image submissions are basically never described in the text, so they are unsearchable even when they are live. There's some exceptions with archives but now you're in power-user territory. Not being compatible with s…

> 4chan constantly deletes its own content.

There are archives, like rebeccablacktech and 4plebs, which Google likewise blackholes. You could argue Google also does not care much for archives, yet I still get StackOverflow clone results, for some reason.

Re: Certified 100% AI-free organic content

#108
> Some of the AI-generated output is factually wrong

While this is undoubtedly true and a problem that needs to be addressed, it's worth considering that humans gets things factually wrong too sometimes (intentionally or not). So perhaps a more interesting question is how much more or less correct an AI is than a human on a given task.

People talk about this with self-driving cars all the time. Arguably a self-driving car does not need to drive perfectly (not that this is not a good goal), but if it can drive more safely than the average human driver, there's still a significant chance of improving overall road safety.

Re: Certified 100% AI-free organic content

#109

Earlier quoted context omitted.

> The long term impact of the ease of generating low nutrition digital content using language models may be that people put down their devices The problem is people don't always make the wise decision. Evidence: the junk food industry is alive and kicking. Some people will disconnect from devices, but others may just say "this is the way things are now" and adjust themselves to the flavor of junk content.

Why are you assuming that it will make writing worse not better? Just because it can be used by non-experts to create crappy written work, it can also be used by people who work with it to augment and improve their existing writen work. To my mind AI is a general purpose technology: https://en.wikipedia.org/wiki/General-purpose_technology I guess using this mental model, what you are worried about is the equivilent t…

Both of these are fascinating questions and, to me anyway, both can be answered with yes. The sheer amount of writing increased exponentially once more people could read, write and publish their own writings ( and internet only exacerbated this trend ). I accordance with pareto principle, most of it was of poor quality, but the upside was that good output likely did increase in terms of absolute number as well ( few people are bound to write something decent ).

I think parent is looking back at history and reasonably infers potential results ( more crap ).

Re: Certified 100% AI-free organic content

#110
post #93

Earlier quoted context omitted.

> So now you need something like ChatGPT to cut through the noise? I once employed a journalist to write about the pros and cons of wedding insurance. Just to give you a clue how long ago this was, it was a unique article at the time. Many years years later, every article you will read about wedding insurance (there will be many thousands) is around 90% similar in style and content to the one I paid for. I dare say y…

My guess is that ChatGPT is going to solve the SEO spam problem by changing the way we search for things. Instead of searching for webpages that have information about a topic, we're going to ask an AI. It'll tell you what the pros and cons of wedding insurance are, and because eventually it'll have access to your calendar, it'll tailor that answers to the specifics of the fact that you're having a destination weddin…

> Instead of searching for webpages that have information about a topic, we're going to ask an AI.

I think this is probably true for some people, the same sort of person who sees something on Facebook and assumes that it's true. [1] But there are quite a lot of people for whom "according to whom?" is the next question after being told something factual. For them, I think search's job is to find relevant sources and get out of the way.

But I think even finding out is a long way away. The main thing that ChatGPT has nailed is glibness. It produces text that sounds authoritative, whether or not it's correct. And it's often incorrect. People may try ChatGPT search out of novelty or because it feels human. But if they depend on it and feel the real-world impact of a confidently wrong answer, they're going to treat it as a human that's untrustworthy. A blowhard, a liar, a fool. So I'm sure the major search players are going to be very cautious rolling out chat-like things. Google has spent decades building up consumer trust, and the don't need a zillion articles about people who a too-confident chat steered wrong.

[1] E.g., That men in white vans are kidnappers: vhttps://www.cnn.com/2019/12/04/tech/facebook-white-vans/inde...

Post reply on HN