Live data from Hacker News

Anthropic Claude 3.5 can create icalendar files, so I did this

gregsramblings.com

81–90 of 174 posts

Re: Anthropic Claude 3.5 can create icalendar files, so I did this

#81

I've been doing similar things with gpt 4o to simplify data entry. Right now, it's kind of useful to use but people are not systematically doing much yet with LLMs. I think that's going to change because it's so obviously useful to do. Any work involving entering data into some form is something that can and will be automated now. Especially if you have the information in some printed or printable way. Just point you…

In my tests, if you operate over diverse forms or documents the information extraction rate is around 90%. It's especially hard for complex forms.

Re: Anthropic Claude 3.5 can create icalendar files, so I did this

#82
post #43

Earlier quoted context omitted.

The "off by one" predilection of LLMs is going to lead to this massive erosion of trust in whatever "Truth" is supposed to be, and it's terrifying and going to make for a bumpy couple of years. (Or the complete collapse of objective knowledge, on a long enough time horizon.) It's one thing to ask an LLM when George Washington was born, and have it return "May 20, 2020." It's another thing to ask it, and have it matte…

This is semi-offtopic, but "trust but verify" is an oxymoron. Trusting something means I don't have to verify whether it's correct (I trust that it is), so the saying, in the end, is "don't verify but verify".

This is not quite true. "Trust" is to give a permission for someone to act on achieving some result. "Verify" means assess the achieved result, and correct aposteriori the probability with which said person is able to achieve the abocementioned result. This is the way Bayesian reasoning works.

Trust has degrees. What you have brought is "unconditional trust". Very rarely works.

Re: Anthropic Claude 3.5 can create icalendar files, so I did this

#83
post #70
post #47

Earlier quoted context omitted.

Yes, which boils down to "verify".

It's possible to trust (or have faith) in my car being able to drive another 50k miles without breaking down. But if I bring it to a mechanic to have the car inspected just in case, does that mean I never had trust/faith in the car to begin with? "I trust my coworkers write good code, but I verify with code reviews" -- doing code reviews doesn't mean you don't trust your coworker. Yet another way to look at it: peopl…

It's a matter of degrees. Absolute trust is a rare thing, but people have given examples of relative trust. Your car won't break down and you can trust it with your kids' life, almost never challenging its trustworthiness, but still you can do checkups or inspections, because some of the bult-in redundancies might be strained. Trusting aircraft but still doing inspections. Trusting your colleagues to do their best but still doing reviews because every fucks up once in a while.

The idea of trusting a next-token-predictor (jesting here) is akin to trusting your System 1 - there's a degree to find where you force yourself to enable System 2 and correct biases.

Re: Anthropic Claude 3.5 can create icalendar files, so I did this

#84
post #64
post #43

Earlier quoted context omitted.

This is semi-offtopic, but "trust but verify" is an oxymoron. Trusting something means I don't have to verify whether it's correct (I trust that it is), so the saying, in the end, is "don't verify but verify".

French armed forces have a better version of this saying. “Trust does not exclude control.” They’re still going to check for explosives under cars that want to park in French embassies.

I think a better translation of "control" in that saying is "checking" or "testing". "Control" in present-day English is a false cognate there.

Re: Anthropic Claude 3.5 can create icalendar files, so I did this

#85
post #3

You just have to double check the results whenever you tell Claude to extract lists and data. 99.9% of it will be correct, but sometimes 1 or 2 records are off. This kind of error is especially hard to notice because you're so impressed that Claude managed to do the extract task at all -- plus the results look wholly plausible upon eyeballing -- that you wouldn't expect anything to be wrong at all. But LLMs can get t…

Can Claude tdd itself ? Lean-Claude

Re: Anthropic Claude 3.5 can create icalendar files, so I did this

#86
And this is where Siri failed in my opinion. Even things like "Create an event that starts on Monday and continues for 90 days" is and always has been a crapshoot.

All the things you didn't want to do manually yourself, Siri couldn't do either. Not much of an assistant then.

Re: Anthropic Claude 3.5 can create icalendar files, so I did this

#87
post #43

Earlier quoted context omitted.

This is semi-offtopic, but "trust but verify" is an oxymoron. Trusting something means I don't have to verify whether it's correct (I trust that it is), so the saying, in the end, is "don't verify but verify".

Is it an oxymoron to generate an asymmetrical cryptographic signature, send it to someone, and that someone verify the signature with the public key? Why not just "trust" them instead? You have a contact and you know them, can't you trust them? This is what "trust but verify" means. It means audit everything you can. Do not really on trust alone. An entire civilization can be built with this methodology. It would be…

> An entire civilization can be built with this methodology. It would be a much better one than the one we have now.

No, it wouldn't. Trust is an optimization that enables civilization. The extreme end of "verify" is the philosophy behind cryptocurrencies: never trust, always verify. It's interesting because it provides an exchange rate between trust and kilowatt hours you have to burn to not rely on it.

Re: Anthropic Claude 3.5 can create icalendar files, so I did this

#88
post #3

You just have to double check the results whenever you tell Claude to extract lists and data. 99.9% of it will be correct, but sometimes 1 or 2 records are off. This kind of error is especially hard to notice because you're so impressed that Claude managed to do the extract task at all -- plus the results look wholly plausible upon eyeballing -- that you wouldn't expect anything to be wrong at all. But LLMs can get t…

The "off by one" predilection of LLMs is going to lead to this massive erosion of trust in whatever "Truth" is supposed to be, and it's terrifying and going to make for a bumpy couple of years. (Or the complete collapse of objective knowledge, on a long enough time horizon.) It's one thing to ask an LLM when George Washington was born, and have it return "May 20, 2020." It's another thing to ask it, and have it matte…

Already quite a while ago I was entertained by a particular British tabloid article, which had been "AI edited". Basically the article was partially correct, but then it went badly wrong because the subject of the article was about recent political events that had happened some years after the point where LLM's training data ended. Because of this, the article contained several AI-generated contextual statements about state of the world that had been true two years ago, but not anymore.

They quietly fixed the article only after I pointed its flaws out to them. I hope more serious journalists don't trust AI so blindly.

Re: Anthropic Claude 3.5 can create icalendar files, so I did this

#89

Earlier quoted context omitted.

To bad that wouldn't work for good music albums. For many years (still, maybe?) albums would come out on Tuesdays. So one day of the week you'd have them all bunched up together. That's how I remember that September 11, 2001 was a Tuesday. Album day, and there was a good one that day, too.

Jay-Z? Or Bob Dylan?

I always think of God Hates Us All by Slayer. Can't quite pinpoint exactly why.

Re: Anthropic Claude 3.5 can create icalendar files, so I did this

#90
post #3

You just have to double check the results whenever you tell Claude to extract lists and data. 99.9% of it will be correct, but sometimes 1 or 2 records are off. This kind of error is especially hard to notice because you're so impressed that Claude managed to do the extract task at all -- plus the results look wholly plausible upon eyeballing -- that you wouldn't expect anything to be wrong at all. But LLMs can get t…

Kind of a noob question, is it possible to design a GAN type network with LLM, where one (or many) LLMs generate outputs, while a few other LLMs validate or discriminate them and thus improving generator LLMs accuracy.
Post reply on HN