Live data from Hacker News

Anthropic Claude 3.5 can create icalendar files, so I did this

gregsramblings.com

151–160 of 174 posts

Re: Anthropic Claude 3.5 can create icalendar files, so I did this

#151
post #43

Earlier quoted context omitted.

The "off by one" predilection of LLMs is going to lead to this massive erosion of trust in whatever "Truth" is supposed to be, and it's terrifying and going to make for a bumpy couple of years. (Or the complete collapse of objective knowledge, on a long enough time horizon.) It's one thing to ask an LLM when George Washington was born, and have it return "May 20, 2020." It's another thing to ask it, and have it matte…

This is semi-offtopic, but "trust but verify" is an oxymoron. Trusting something means I don't have to verify whether it's correct (I trust that it is), so the saying, in the end, is "don't verify but verify".

My theory that it's a sophism to mollify people who are offended by not being trusted, or to mollify people who think not trusting is rude.

I'd just assume see it not used though.

Re: Anthropic Claude 3.5 can create icalendar files, so I did this

#152
post #3

You just have to double check the results whenever you tell Claude to extract lists and data. 99.9% of it will be correct, but sometimes 1 or 2 records are off. This kind of error is especially hard to notice because you're so impressed that Claude managed to do the extract task at all -- plus the results look wholly plausible upon eyeballing -- that you wouldn't expect anything to be wrong at all. But LLMs can get t…

> Trust but verify.

This is why I don’t bother with GenAI. It’s a lot harder to verify data from a machine trained to generate plausible looking data (that happens to be correct most of the time, sure) than from real data. Ofc this is getting harder as “real data” gets drowned in GenAI output, which drives the animosity against GenAI in my circle.

Re: Anthropic Claude 3.5 can create icalendar files, so I did this

#153
post #137

Earlier quoted context omitted.

The problem is that people could be impressed and use it for things where 0.01% could lead to people getting hurt or even get killed.

I call that problem "Doctor, doctor! It hurts when I do this!" If the risk exists with AI processing this kind of data, it exists with a human processing the data. The fail-safe processes in place for the human output need to be used for the AI output too, obviously - using the AI speeds up the initial process enormously though

The problem is liability.

Who is liable if the AI makes errors?

Re: Anthropic Claude 3.5 can create icalendar files, so I did this

#154

Earlier quoted context omitted.

> An entire civilization can be built with this methodology. It would be a much better one than the one we have now. No, it wouldn't. Trust is an optimization that enables civilization. The extreme end of "verify" is the philosophy behind cryptocurrencies: never trust, always verify. It's interesting because it provides an exchange rate between trust and kilowatt hours you have to burn to not rely on it.

Yes, let's trust VCs and bankers instead, they seem to be great keepers of civilization -- no calamities in sight with them at the helm /s

Possible > impossible.

I'd first trust unicorns shooting rainbows out of their posteriors before the cryptocurrency vision; neither works for fostering civilization, but at least the unicorns aren't proposing an economy based on paying everyone for wasting energy.

Re: Anthropic Claude 3.5 can create icalendar files, so I did this

#155
post #3

You just have to double check the results whenever you tell Claude to extract lists and data. 99.9% of it will be correct, but sometimes 1 or 2 records are off. This kind of error is especially hard to notice because you're so impressed that Claude managed to do the extract task at all -- plus the results look wholly plausible upon eyeballing -- that you wouldn't expect anything to be wrong at all. But LLMs can get t…

The "off by one" predilection of LLMs is going to lead to this massive erosion of trust in whatever "Truth" is supposed to be, and it's terrifying and going to make for a bumpy couple of years. (Or the complete collapse of objective knowledge, on a long enough time horizon.) It's one thing to ask an LLM when George Washington was born, and have it return "May 20, 2020." It's another thing to ask it, and have it matte…

Having also been burned by "better to check Wikipedia first" hallucinations, I find the Anthropic footnotes system to be essential. The "confidence and bravado" is indeed alluring at first, but eventually leads away from using LLMs as search engines.

Re: Anthropic Claude 3.5 can create icalendar files, so I did this

#156
post #3

You just have to double check the results whenever you tell Claude to extract lists and data. 99.9% of it will be correct, but sometimes 1 or 2 records are off. This kind of error is especially hard to notice because you're so impressed that Claude managed to do the extract task at all -- plus the results look wholly plausible upon eyeballing -- that you wouldn't expect anything to be wrong at all. But LLMs can get t…

The "off by one" predilection of LLMs is going to lead to this massive erosion of trust in whatever "Truth" is supposed to be, and it's terrifying and going to make for a bumpy couple of years. (Or the complete collapse of objective knowledge, on a long enough time horizon.) It's one thing to ask an LLM when George Washington was born, and have it return "May 20, 2020." It's another thing to ask it, and have it matte…

You fear that over time, artificially intelligent systems will suffer from increasingly harmful variance due to deteriorating confidence in training data, until there is total model collapse or otherwise systemic real-world harm.

Personally, I believe we will eventually discover mathematical structures which can reliably extract objective truth, a sieve, at least in terms of internal consistency and relationships between objects.

It's unclear philosophically whether this is actually possible, but it might be possible within specific constraints. For example, we could solve accuracy of dates specifically through some sort of automated reference-checking system, and possibly generalize the approach to entire classes of problems. We might even be able to encode this behavior directly into a model. Dates, at least, are purely empirical data, so there is an objective moment or statistically likely range of time which is recorded somewhere, so the problem becomes one of locating, identifying and analyzing the correct sources for a given piece of information, and having a sound-proof method of checking the work, likely not through an LLM.

I think we will come to discover that LLMs/transformers are a highly generalizable and integral component of a complete artificial brain, but we will soon uncover better metamodels which exhibit true executive functioning and self-referential loops which allow them to be trusted for critical tasks, leveraging transformer architecture for tasks which benefit from it, while employing redundancy and other techniques to increase confidence.

Re: Anthropic Claude 3.5 can create icalendar files, so I did this

#157
post #43

Earlier quoted context omitted.

This is semi-offtopic, but "trust but verify" is an oxymoron. Trusting something means I don't have to verify whether it's correct (I trust that it is), so the saying, in the end, is "don't verify but verify".

Yep. https://en.m.wikipedia.org/wiki/Trust,_but_verify

> He said "President Reagan's old adage about 'trust but verify' ... is in need of an update. And we have committed here to a standard that says 'verify and verify'."

Re: Anthropic Claude 3.5 can create icalendar files, so I did this

#158
post #153

Earlier quoted context omitted.

I call that problem "Doctor, doctor! It hurts when I do this!" If the risk exists with AI processing this kind of data, it exists with a human processing the data. The fail-safe processes in place for the human output need to be used for the AI output too, obviously - using the AI speeds up the initial process enormously though

The problem is liability. Who is liable if the AI makes errors?

[deleted]

Re: Anthropic Claude 3.5 can create icalendar files, so I did this

#159
post #153

Earlier quoted context omitted.

I call that problem "Doctor, doctor! It hurts when I do this!" If the risk exists with AI processing this kind of data, it exists with a human processing the data. The fail-safe processes in place for the human output need to be used for the AI output too, obviously - using the AI speeds up the initial process enormously though

The problem is liability. Who is liable if the AI makes errors?

It depends on the richness of the AI

Re: Anthropic Claude 3.5 can create icalendar files, so I did this

#160
post #43

Earlier quoted context omitted.

The "off by one" predilection of LLMs is going to lead to this massive erosion of trust in whatever "Truth" is supposed to be, and it's terrifying and going to make for a bumpy couple of years. (Or the complete collapse of objective knowledge, on a long enough time horizon.) It's one thing to ask an LLM when George Washington was born, and have it return "May 20, 2020." It's another thing to ask it, and have it matte…

This is semi-offtopic, but "trust but verify" is an oxymoron. Trusting something means I don't have to verify whether it's correct (I trust that it is), so the saying, in the end, is "don't verify but verify".

It's a more friendly way of saying "trust noone".
Post reply on HN