Live data from Hacker News

AI Police Reports: Year in Review

eff.org

121–130 of 224 posts

Re: AI Police Reports: Year in Review

#121
post #120

Earlier quoted context omitted.

LLMs may appear to do well on certain programming tasks on which they are trained intensively, but they are incredibly weak. If you try to use an LLM to generate, for example, a story, you will find that it will make unimaginable mistakes. If you ask an LLM to analyze a conversation from the internet it will misrepresent the positions of the participants, often restating things so that they mean something different o…

We do have AI systems that write stories [0]. They work. The quality might not be spectacular but if you've ever gone out and spent time reading fanfiction you'd have to agree there are a lot of rather terrible human writers too (bless them). It still hits this issue that if we want LLMs to compete with the best of humanity then they aren't there yet, but that means defining human intelligence as something that most…

I haven't tried Dramatron, but my experience is that it isn't possible to do sensibly. With regard to the second part

>AI transcription & summary seems to be a strong point of the models so I don't know what exactly you're trying to get to with this one. If you have evidence for that I'd actually be quite interested because humans are so bad at representing what other people said on the internet it seems like it should be an easy win for an AI. Humans typically have some wild interpretations of what other people write that cannot be supported from what was written.

Transcription and summarization is indeed fine, but try posting a longer reddit or HN discussion you've been part of into any model of your choice and ask it to analyze it, and you will see severe errors very soon. It will consistently misrepresent the views expressed and it doesn't really matter what model you go for. They can't do it.

Re: AI Police Reports: Year in Review

#122
post #98
post #45

Earlier quoted context omitted.

> a lot of people seem to see LLMs as smarter than themselves Well, in many cases they might be right..

> ChatGPT (o3): Scored 136 on the Mensa Norway test in April 2025 So yes, most people are right in that assumption, at least by the metric of how we generally measure intelligence.

Yeah I certainly associate LLMs with high intelligence when they provide fake links to fake information, I think, man this thing is SMART

Re: AI Police Reports: Year in Review

#123
post #42

Earlier quoted context omitted.

It's pretty similar to looking something up with a search engine, mashing together some top results + hallucinating a bit, isn't it? The psychological effects of the chat-like interface + the lower friction of posting in said chat again vs reading 6 tabs and redoing your search, seems to be the big killer feature. The main "new" info is often incorrect info. If you could get the full page text of every url on the fir…

> If you could get the full page text of every url on the first page of ddg results and dump it into vim/emacs where you can move/search around quickly, that would probably be similarly as good, and without the hallucinations. Curiously, literally nobody on earth uses this workflow. People must be in complete denial to pretend that LLM (re)search engines can’t be used to trivially save hours or days of work. The accu…

> The accuracy isn’t perfect

The reason why people don't use LLMs to "trivially save hours or days of work" is because LLMs don't do that. People would use a tool that works. This should be evidence that the tools provide no exceptional benefit, why do you think that is not true?

Re: AI Police Reports: Year in Review

#124

What worries me is that _a lot of people seem to see LLMs as smarter than themselves_ and anthropmorphize them into a sort of human-exact intelligence. The worst-case scenario of Utah's law is that when the disclaimer is added that the report is generated by AI, enough jurists begin to associate that with "likely more correct than not".

Maybe it's just my circle, but anecdotally most of the non-CS folks I know have developed a strong anti-AI bias. In a very outspoken way.

If anything, I think they'd consider AI's involvement as a strike against the prosecution if they were on a jury.

Re: AI Police Reports: Year in Review

#125
post #98
post #45

Earlier quoted context omitted.

> a lot of people seem to see LLMs as smarter than themselves Well, in many cases they might be right..

> ChatGPT (o3): Scored 136 on the Mensa Norway test in April 2025 So yes, most people are right in that assumption, at least by the metric of how we generally measure intelligence.

> the metric of how [the uninformed] generally measure intelligence

Re: AI Police Reports: Year in Review

#126

What worries me is that _a lot of people seem to see LLMs as smarter than themselves_ and anthropmorphize them into a sort of human-exact intelligence. The worst-case scenario of Utah's law is that when the disclaimer is added that the report is generated by AI, enough jurists begin to associate that with "likely more correct than not".

Maybe it's just my circle, but anecdotally most of the non-CS folks I know have developed a strong anti-AI bias. In a very outspoken way. If anything, I think they'd consider AI's involvement as a strike against the prosecution if they were on a jury.

Why do people in your circle not like AI? I have similar a experience about friends and family not liking AI, but usually it’s due to water and energy reasons, not because of an issue with the model reasoning

Re: AI Police Reports: Year in Review

#127
post #60

Earlier quoted context omitted.

As far as I can tell from poking people on HN about what "AGI" means, there might be a general belief that the median human is not intelligent. Given that the current batch of models apparently isn't AGI I'm struggling to see a clean test of what AGI might be that a human can pass.

Being an intelligent being is not the same as being considered intelligent relative to the rest of your species. I think we’re just looking to create an intelligence, meaning, having the attributes that make a being intelligent, which mostly are the ability to reason and learn. I think the being might take over from there no? With humans, the speed and ease with which we learn and reason is capped. I think a very dum…

> every resource will be spent in making it smarter

The root motivation on which every resource will be spent is simply and very obviously to make a profit.

Re: AI Police Reports: Year in Review

#128
post #120

Earlier quoted context omitted.

We do have AI systems that write stories [0]. They work. The quality might not be spectacular but if you've ever gone out and spent time reading fanfiction you'd have to agree there are a lot of rather terrible human writers too (bless them). It still hits this issue that if we want LLMs to compete with the best of humanity then they aren't there yet, but that means defining human intelligence as something that most…

I haven't tried Dramatron, but my experience is that it isn't possible to do sensibly. With regard to the second part >AI transcription & summary seems to be a strong point of the models so I don't know what exactly you're trying to get to with this one. If you have evidence for that I'd actually be quite interested because humans are so bad at representing what other people said on the internet it seems like it shou…

I can see why they'd struggle, I'm not sure what you're trying to ask the model to do. What type of analysis are you expecting? If the model is supposed to represent the views expressed that would be a summary. If you aren't asking it for a summary what do you want it to do? Do you literally mean you want the model to perform conversational analysis (ie, https://en.wikipedia.org/wiki/Conversation_analysis#Method)?

Re: AI Police Reports: Year in Review

#129
post #106
post #100

Earlier quoted context omitted.

Court reports should as much be about human sensibility. I have met plenty of high IQ people who were insensitive.

Having listened to some the new AI generated songs on utube, looks like they might be better at being sensitive humans than we are as well..

Where do you imagine they copied those human sensitivities from? The weather?

Re: AI Police Reports: Year in Review

#130
post #60

Earlier quoted context omitted.

As far as I can tell from poking people on HN about what "AGI" means, there might be a general belief that the median human is not intelligent. Given that the current batch of models apparently isn't AGI I'm struggling to see a clean test of what AGI might be that a human can pass.

LLMs may appear to do well on certain programming tasks on which they are trained intensively, but they are incredibly weak. If you try to use an LLM to generate, for example, a story, you will find that it will make unimaginable mistakes. If you ask an LLM to analyze a conversation from the internet it will misrepresent the positions of the participants, often restating things so that they mean something different o…

> We are incredibly far from AGI.

This and we don't actually know what the foundation models are for AGI, we're just assuming LLMs are it.

Post reply on HN