Live data from Hacker News

The 100k whys of AI

lcamtuf.substack.com

31–40 of 111 posts

Re: The 100k whys of AI

#31
post #3

When you generate one or two blog posts with LLM they look pretty good. And you will be impressed with that one clever bit it adds that you didn't even ask for. But then you generate 50 of them and they all converge into the same pattern. It's hard to prove that an article is AI generated but they are instantly recognizable. An aside, I usually take my written blog posts through a pass on Notebooklm to generate a pod…

> they all converge

AI is regression to the mean.

Much like Socialism.

Om an acute basis, AI can be just as helpful as that safety net.

As a chronic matter, "it's not excellence--it's mediocrity".

Re: The 100k whys of AI

#33
post #9

I don't know how much of a smoking gun this actually is, the evidence proffered doesn't establish anything - I can see some names there like Havilah Brooks or Celina Briar who are intentionally re-using the same title to create a series, for example. And this doesn't really get into the base rate of generic title re-use among encyclopedias. There isn't much reward for coming up with an imaginative title for kids, the…

Did you see what's inside one of those books?

https://infosec.exchange/@lcamtuf/116785283147249092

This is Amazon #1 bestseller in "Children's Encyclopedias"!

Re: The 100k whys of AI

#34
post #2

A nice illustration of the homogeneity of LLM responses. Another way to describe this effect would be… If you ask humans to write 1,000 books, you're asking 1,000 different humans with different experiences and different skills and different moods (etc.) to write those books. But if you ask LLMs to write 1,000 books, you're probably only talking to 3 or 5 different models, tops. And they've all trained on the same or…

prompts will give very different results. this is where you do the work.

Re: The 100k whys of AI

#35
post #9

I don't know how much of a smoking gun this actually is, the evidence proffered doesn't establish anything - I can see some names there like Havilah Brooks or Celina Briar who are intentionally re-using the same title to create a series, for example. And this doesn't really get into the base rate of generic title re-use among encyclopedias. There isn't much reward for coming up with an imaginative title for kids, the…

Have you seen the content of the books in the tweet[1] linked below the article? Between horses with fused butts and other diagrams that don't say as much as they purport to, the cover is the least of its problems, although the only one that can be criticized directly.

[1] https://infosec.exchange/@lcamtuf/116785283147249092

Re: The 100k whys of AI

#36
post #34
post #2

A nice illustration of the homogeneity of LLM responses. Another way to describe this effect would be… If you ask humans to write 1,000 books, you're asking 1,000 different humans with different experiences and different skills and different moods (etc.) to write those books. But if you ask LLMs to write 1,000 books, you're probably only talking to 3 or 5 different models, tops. And they've all trained on the same or…

prompts will give very different results. this is where you do the work.

Yes but not very different results (unless you're adding new information to your prompt or reducing some ambiguity). Prompt engineering is mostly pseudoscience.

Re: The 100k whys of AI

#37

I think a majority of content consumers can already distinguish LLM content from human content. I'm looking forward to the day that they're intelligent enough to care, but I'm not holding my breath. Orwell framed it pretty well in 1984 with the machine-generated songs that were new every year, but always tugged on the heartstrings of the proles. They weren't really readers or listeners to music or appreciators of art…

It's not an "already", because I assume models will get better at addressing mode collapse.

The irony in the machine generated songs in 1984 was that Winston clearly found meaning in them, feeling like they applied to him, even though he knew they were machine generated: (from memory) "Under the shade of the chestnut tree / I sold you and you sold me / here lie they and here lie we / under the shade of the chestnut tree" - that refers to him and Julia selling each other out, right?

Just like people today - and in George Orwell's day, which was why he made it - find meaning in things which is obviously formulaic manufacured corporate slop, like the endless MCU films.

Re: The 100k whys of AI

#38
post #23

We likes this "same, complex set of mannerism" when using LLM for programming. If you ask LLM to write a certain function for you, it gives you statistically obvious implementation. But maybe for writing an original book, this feature is not so desirable

It does not. Sometimes it will spawn a mess of ad hoc python, sometimes it will do curl and sed, and very very occasionally it will use the correct tool for the job if it remembers to use the skill you developed in the previous session.

Yes, sometimes it does something unexpected when used as a tool for programming. And in that context, it is seen as an unwanted feature. In fact, that was my point. However, I disagree that it does a good job only "very very occasionally". That is not my experience at all.

Re: The 100k whys of AI

#40
I don't want to hurt people's feelings, so in person I restrain myself from speaking out (it wouldn't change anything anyway)... but every person I have seen so far, who was bullish on building an AI business has followed the same path:

  1) They think the AI can replace them, but in a good way: "it will keep doing my job and people will pay ME"
  2) They assume people either don't notice or don't mind that it's AI. They build businesses, where AI impersonates a professional when that person is not available ("chat with your therapist any time even if they sleep!")
  3) All they do is based on written or spoken words. There is no substance
I expect that sooner than later a great skepticism for anything non-tangible will develop. Personally, I have been highly distrustful of people who don't build things (even the word "building" is now tainted). I think it will accelerate.
Post reply on HN