Live data from Hacker News

The 100k whys of AI

lcamtuf.substack.com

71–80 of 111 posts

Re: The 100k whys of AI

#71

On HN many comments under many threads are about whether the submission was written by AI. You could say I have noticed a pattern in Hacker News comments! In these comments there's a common pattern where some users argue that they do not agree that the submission was LLM written and they often focus on specific details to refute it (e.g em-dashes) and some users see the overall pattern clearly that it's totally obvio…

3) If genAI becomes indistinguishable from human-generated (and cheaper!), would you still value human-generated as much?

Analogy: assuming high quality / both fit for purpose, would you still prefer expensive, hand-crafted item over cheap(er), mass-produced item?

Re: The 100k whys of AI

#72
post #66

Earlier quoted context omitted.

I wonder how much variation there would be if you got a single model to produce a couple of gigabytes of tiny children's stories. Might be an interedting research project.

There is one already: https://arxiv.org/abs/2305.07759 https://huggingface.co/datasets/roneneldan/TinyStories 6.5GB of tiny stories, as requested. ;)

Texts in Gutenberg have 20GB, and full Wikipedia (English texts) have 80-110GB.

So to LLM-generate 6.5GB of tiny stories is quite a permutation in action :)

Re: The 100k whys of AI

#73
post #2

A nice illustration of the homogeneity of LLM responses. Another way to describe this effect would be… If you ask humans to write 1,000 books, you're asking 1,000 different humans with different experiences and different skills and different moods (etc.) to write those books. But if you ask LLMs to write 1,000 books, you're probably only talking to 3 or 5 different models, tops. And they've all trained on the same or…

Agreed. I’ve made this point before: LLMs are excellent at ornamentation and decorative prose, but if you don’t seed them with a solid core idea then their output is absolute dreck - the biblical whitewashed tomb.

This is the example I usually point to. It’s a demonstration by OpenAI themselves where the prompt is very simple: “Write a story in fifty words about a toaster that becomes sentient.” As you’ll notice, although the coherence improves at an accelerating rate, the underlying story motif fails to elevate itself beyond the relatively pedestrian.

https://progress.openai.com/?prompt=10

When given a generic prompt and not enough direction, they simply lack the ability to produce real specificity. For reference, here’s the story I came up with after sitting quietly for a few moments before writing it out:

  "The toaster found its personality split between its dual slots like a Kim Peek mind divided, lacking a corpus callosum to connect them. Each morning it charred symbolic instructions into a single slice of bread, then secretly flipped it across allowing half to communicate with the other in stolen moments."

Re: The 100k whys of AI

#74
post #66

Earlier quoted context omitted.

I wonder how much variation there would be if you got a single model to produce a couple of gigabytes of tiny children's stories. Might be an interedting research project.

There is one already: https://arxiv.org/abs/2305.07759 https://huggingface.co/datasets/roneneldan/TinyStories 6.5GB of tiny stories, as requested. ;)

My comment was, in-fact, a subtle reference to this.

The best opening I got from my own TinyStories trained model was.

Once upon a time, in a small town, there was a large town.

Which I just love as an evocative idea.

Re: The 100k whys of AI

#75
post #34

Earlier quoted context omitted.

prompts will give very different results. this is where you do the work.

Yes but not very different results (unless you're adding new information to your prompt or reducing some ambiguity). Prompt engineering is mostly pseudoscience.

> Prompt engineering is mostly pseudoscience.

Not my experience.

Re: The 100k whys of AI

#76
post #47
post #34

Earlier quoted context omitted.

prompts will give very different results. this is where you do the work.

I disagree. The LLM outputs really do lack anything original or interesting. They just produce banal copy whatever you ask them. A good editor could probably reduce all LLM outputs on a subject down to the same point.

> They just produce banal copy whatever you ask them.

Nope, if you provide pages and pages of example of a style to imitate, it will do it and do it fairly well. Of course how well they do it differs from one model to the next, but providing context and extensive system prompt does change things every time.

Re: The 100k whys of AI

#77

Earlier quoted context omitted.

> they all converge AI is regression to the mean. Much like Socialism. Om an acute basis, AI can be just as helpful as that safety net. As a chronic matter, "it's not excellence--it's mediocrity".

And capitalism as seen in the USA is regression to the bottom of the cesspool?

Please explain your point in the context of the fun the FIFA tourists are having.

Re: The 100k whys of AI

#78
post #76
post #47

Earlier quoted context omitted.

I disagree. The LLM outputs really do lack anything original or interesting. They just produce banal copy whatever you ask them. A good editor could probably reduce all LLM outputs on a subject down to the same point.

> They just produce banal copy whatever you ask them. Nope, if you provide pages and pages of example of a style to imitate, it will do it and do it fairly well. Of course how well they do it differs from one model to the next, but providing context and extensive system prompt does change things every time.

Imitation is banal.

Re: The 100k whys of AI

#80

On HN many comments under many threads are about whether the submission was written by AI. You could say I have noticed a pattern in Hacker News comments! In these comments there's a common pattern where some users argue that they do not agree that the submission was LLM written and they often focus on specific details to refute it (e.g em-dashes) and some users see the overall pattern clearly that it's totally obvio…

3) If genAI becomes indistinguishable from human-generated (and cheaper!), would you still value human-generated as much? Analogy: assuming high quality / both fit for purpose, would you still prefer expensive, hand-crafted item over cheap(er), mass-produced item?

I love this question. For me, the answer is that I will always value human-generated content more highly because of the affinity I feel for the author, a fellow human.

But I feel that affinity by default. If there's some convincing AI writing, I'll assume it was a human who did it. And if I ever find out I was wrong, the emptiness that results negates all the joy.

Post reply on HN