Live data from Hacker News

Certified 100% AI-free organic content

substack.piszek.com

141–150 of 275 posts

Re: Certified 100% AI-free organic content

#141
post #118

This strikes me as a distinction without a difference. Content creation is not even close to the same as GMO or processed foods. People are justified in being concerned about how food is made precisely because there are meaningful differences in the nutritional value and trace chemical content in the end products. And that matters because we don't fully understand the role those factors play in long term human health…

At the end of the day, good content is good content regardless of who or what produced it. We should focus on recognizing and celebrating the beauty, regardless of its origin.

Re: Certified 100% AI-free organic content

#143

> Published content will be later used to train subsequent models, and being able to distinguish AI from human input may be very valuable going forward I find this to be a particularly interesting problem in this whole debacle. Could we end up having AI quality trend downwards due to AI ingesting its own old outputs and reinforcing bad habits? I think it's a particular risk for text generation. I've already run into…

I think that in addition to influencing future models, AI content will also influence how humans think and write. People will start ironically and unironically copying GPT's style in their own writing, causing human produced content to increasing resemble AI content.

High school students that are prohibited from using AI for their essays will have a bad time. Even if they don't use AI chatbots themselves, they will unknowingly cite sources that were written by AI, or were written by someone who learned about the topic by asking ChatGPT.

Re: Certified 100% AI-free organic content

#144
post #126

Earlier quoted context omitted.

That just isn't true. It's expensive, but entirely doable. Also, it's perfectly normal to perform initial model training on a large data set to capture the statistical properties of language, then perform a second stage of model training on more curated data to cause the model to actually do what you want.

Current LLMs are trained on as much data that can be scraped from the public internet. It’s simply not possible to annotate that much data, even with crowdsourcing. It’s not even a matter of cost. You’d basically need to duplicate the amount of data on the internet. I don’t think you’re appreciating the scale of the data involved in training these models.

Not necessarily. The bloom model (a GPT competitor and similarly sized) was trained on 1.5T of text, which reduces down to 350B unique tokens. If you took a histogram of those unique tokens, it would have a very long tail with probably 1% or less being well represented. That leaves 350M common tokens to serve as the basis for token tuples being fed into crowdsourcing. There are probably ~2-5B very common token sequences, if you had 5 people view each token sequence and give it a few scores, and that process took ~1-2 minutes (these sequences are short), that leaves a conservative estimate of 50 billion person minutes, or ~34 million person days. If you paid these workers $15/hour, that comes out to $12.5 billion dollars, which is not prohibitively expensive for any big tech company when spread out over several years, particularly when it provides a massive competitive advantage.

Re: Certified 100% AI-free organic content

#146

> Published content will be later used to train subsequent models, and being able to distinguish AI from human input may be very valuable going forward I find this to be a particularly interesting problem in this whole debacle. Could we end up having AI quality trend downwards due to AI ingesting its own old outputs and reinforcing bad habits? I think it's a particular risk for text generation. I've already run into…

As someone in SEO, I've been pretty disgusted by the desire for site owners to want to use AI-generated content. There are various opinions on this, of course, but I got into SEO out of interest in the "organic web" vs. everything being driven by ads. Love the idea of having AI-Free declarations of content as it could / should help to differentiate organic content from generated content. It would be very interesting…

Don't worry, it won't be small shops doing it. It will be the majors, if that's where the money is.

To quote Yan LeCun:

   Meta will be able to help small businesses promote themselves by automatically producing media that promote a brand, he offered. 

   "There's something like 12 million shops that advertise on Facebook, and most of them are mom and pop shops, and they just don't have the resources to design a new, nicely designed ad," observed LeCun. "So for them, generative art could help a lot."
https://www.zdnet.com/article/chatgpt-is-not-particularly-in...

Re: Certified 100% AI-free organic content

#147

> Published content will be later used to train subsequent models, and being able to distinguish AI from human input may be very valuable going forward I find this to be a particularly interesting problem in this whole debacle. Could we end up having AI quality trend downwards due to AI ingesting its own old outputs and reinforcing bad habits? I think it's a particular risk for text generation. I've already run into…

Those blogs already exist. Pretty much 90% of the results I see in Google for non-technical household related queries. Just incoherent rambling that sounds plausible but is complete nonsense.

Re: Certified 100% AI-free organic content

#149
I was thinking about this exact problem a few days ago when I created a site hosting poems that were either 100% AI written, or 100% Human. https://news.ycombinator.com/item?id=34472478

Then I asked people to guess the authorship. Amazingly, only 70% of the time the guess people make is correct. https://random-poem.com/

why : Is this Poem written by or by ? Guess & Click.

I'm guessing it will get even harder to tell as the AI improves further down the road.

Post reply on HN