Live data from Hacker News

The 100k whys of AI

lcamtuf.substack.com

91–100 of 111 posts

Re: The 100k whys of AI

#91
post #89

Aw, it's just one big picture of book covers. You can't click on the books and read them. If they're AI-written, they're not copyrightable, so you could post the full text. Looking at the books side by side would be interesting. A test for AI-generated art: railroad tracks. For some reason, none of the image generators can get railroad trackage even close to correct. Just getting long, parallel rails correct seems to…

You can read at least the first chapter or so if you do the Amazon search. I did, and made some discoveries.

* https://mastodonapp.uk/@JdeBP/116788436560592117

Re: The 100k whys of AI

#92
post #24

It's even more worrying when you look at the contents of these "books", they are riddled with erros: https://infosec.exchange/@lcamtuf/116785283147249092

That's a generalization from 1 book. I went and looked at ten or so. This is not actually the case. There's something more complex going on.

* https://mastodonapp.uk/@JdeBP/116788511790947929

Re: The 100k whys of AI

#94
post #2

A nice illustration of the homogeneity of LLM responses. Another way to describe this effect would be… If you ask humans to write 1,000 books, you're asking 1,000 different humans with different experiences and different skills and different moods (etc.) to write those books. But if you ask LLMs to write 1,000 books, you're probably only talking to 3 or 5 different models, tops. And they've all trained on the same or…

LLMs are great at producing average. We see this with their GenAI music equivalents. All the music these GenAI models produce is exceptionally (aggressively, even) average. It is the most polished average you'll ever find. Never awful (anymore), never fantastic. Just bang in the middle.

That is definitely the essence of AI: It is the average of all the inputs it has been trained on.

Frank Zappa was once asked about guitar virtuosos like John McLaughlin and his answer was somemthing like "You can maybe plays solo faster than anybody, but can your playing surprise me?".

Re: The 100k whys of AI

#95

The whole point of the thesis is that because the cover image are very similar, therefore LLMs are bad at writing text? I think it's that today's LLMs have access to poor/generic image generation models and people find it easier to ask ChatGPT or NanoBanana to make a cover instead of fine tuning a small SD model for the purpose.

The people in the FediVerse discussions have also looked at the book contents.

* https://mastodonapp.uk/@JdeBP/116788511790947929

* https://hachyderm.io/@ariels/116788498255660876

Re: The 100k whys of AI

#96

There used to be a word for this in generative AI: mode collapse. It's not that the model doesn't generate human-like responses, it's that it generates the same 0.0001% of possible human like responses every time. It's almost certainly the instruction tuning which is responsible, maybe some small part of blame could go to the rollout policy (I have no idea how rollout policy works these days).

The LLM has its context-window. When it gets over that I assume it starts more or less repeating itself. Whereas human context-window has memories and inputs from all of one's life. Therefore great authors don't repeat themselves.

Now even if an LLM has a large context-window it is probabably not the case it rememebers all of your previous prompts and all of its previous replies. If you ask it to write a book you should probabaly give it all the previous 50 books (or blog-posts) it has written for you so far and you should tell it not to repeat itself. But in practice the context-window and the cost of token would become too expensive for it to write 50 unique books.

Maybe the problem is "all-or-nothing" -nature of LLM context window. Humans don't remmeber everything from past but they remember something from ALL OF their past.

Re: The 100k whys of AI

#97

Earlier quoted context omitted.

Seems that both you and the gp are starting from the assumption that those uniform results are representative of those who use AI and of AI usage. In fact they have been chosen for their uniformity- they might be only a small part of a much more varied output obtained by more demanding (or lucky) users.

I think the uniformity is real. All users interact with the same initial state of the model when they start each chat. Models are not trained to be wildly creative and try to stick to the point. So when users prompt them in pretty much the same manner they quite stably generate very similar output. I wonder if there aren't a simple creative hack to discover, for example to prompt the model to produce more unexpected…

Yes, the uniformity is real- I made the same exact argument at the beginning of this thread. But you can't judge "AI users" in general based on this output because you have selected only what is visibly uniform. Even if 99% of the users introduced enough variation to produce different results, you would still be selecting the 1% that is identical.

> Models are not trained to be wildly creative and try to stick to the point

Models might be as creative as humans, they would still start always from the exact same state. If you ask an LLM to think of three random numbers it will spit out always the same ones. If you tell it to avoid the first that came to its mind, the second choices will also be always the same.

From qntm's Lena:

"the emulated Miguel Acevedo boots with an excited, pleasant demeanour. He is eager to understand how much time has passed since his uploading, what context he is being emulated in, and what task or experiment he is to participate in. If asked to speculate, he guesses that he may have been booted for the IAAS-1 or IAAS-5 experiments".

Every single time.

Re: The 100k whys of AI

#98

Earlier quoted context omitted.

That whole thing would get you 1000 variants of existing art. But if you asked a thousand different designers to do a cover for the same book...

> 1000 variants of existing art. This is very naive. I can almost guarantee that some combinations of 20 * 50 features will hit on something that has never been written before in that specific combination . And if that's still not enough, increase the number of features. Add more randomness, add more steering, add random steering in random chapters, change it up, and so on.

Sure, just like 1000 monkeys with typewriters will write 1000 technically unique books - but they are all still filled with the same garbage.

Re: The 100k whys of AI

#99

Earlier quoted context omitted.

> they all converge AI is regression to the mean. Much like Socialism. Om an acute basis, AI can be just as helpful as that safety net. As a chronic matter, "it's not excellence--it's mediocrity".

And capitalism as seen in the USA is regression to the bottom of the cesspool?

If you need another comment upon which to pour resentment, then I am here fpr you.

Re: The 100k whys of AI

#100
post #2

A nice illustration of the homogeneity of LLM responses. Another way to describe this effect would be… If you ask humans to write 1,000 books, you're asking 1,000 different humans with different experiences and different skills and different moods (etc.) to write those books. But if you ask LLMs to write 1,000 books, you're probably only talking to 3 or 5 different models, tops. And they've all trained on the same or…

> If you ask humans to write 1,000 books

Yeah, but at least in genre fiction, what readers really want[0] is the same 3 or 5 books written in slightly different settings over and over again.

[0]: "want" means actually want, in other words, willing to pay for it.

Post reply on HN