Live data from Hacker News

Anthropic Claude 3.5 can create icalendar files, so I did this

gregsramblings.com

131–140 of 174 posts

Re: Anthropic Claude 3.5 can create icalendar files, so I did this

#131
post #3

You just have to double check the results whenever you tell Claude to extract lists and data. 99.9% of it will be correct, but sometimes 1 or 2 records are off. This kind of error is especially hard to notice because you're so impressed that Claude managed to do the extract task at all -- plus the results look wholly plausible upon eyeballing -- that you wouldn't expect anything to be wrong at all. But LLMs can get t…

I wonder if tool-calling to output schema'd json would have a low error rate here. For each field, you could have a description of what is approximately right, and that should anchor the output better than a one-off prompt.

Re: Anthropic Claude 3.5 can create icalendar files, so I did this

#132
post #3

You just have to double check the results whenever you tell Claude to extract lists and data. 99.9% of it will be correct, but sometimes 1 or 2 records are off. This kind of error is especially hard to notice because you're so impressed that Claude managed to do the extract task at all -- plus the results look wholly plausible upon eyeballing -- that you wouldn't expect anything to be wrong at all. But LLMs can get t…

Kind of a noob question, is it possible to design a GAN type network with LLM, where one (or many) LLMs generate outputs, while a few other LLMs validate or discriminate them and thus improving generator LLMs accuracy.

Yes, you can use AI to spot errors in AI output. Done before with good results, but it requires to run 2 different, but equally good, AI models in parallel, which is way more expensive than 1 model.

Re: Anthropic Claude 3.5 can create icalendar files, so I did this

#133

Earlier quoted context omitted.

Kind of a noob question, is it possible to design a GAN type network with LLM, where one (or many) LLMs generate outputs, while a few other LLMs validate or discriminate them and thus improving generator LLMs accuracy.

Yes, you can use AI to spot errors in AI output. Done before with good results, but it requires to run 2 different, but equally good, AI models in parallel, which is way more expensive than 1 model.

"which is way more expensive than 1 model."

In this case "way more" means exactly 2x the cost.

Re: Anthropic Claude 3.5 can create icalendar files, so I did this

#134
post #43

Earlier quoted context omitted.

The "off by one" predilection of LLMs is going to lead to this massive erosion of trust in whatever "Truth" is supposed to be, and it's terrifying and going to make for a bumpy couple of years. (Or the complete collapse of objective knowledge, on a long enough time horizon.) It's one thing to ask an LLM when George Washington was born, and have it return "May 20, 2020." It's another thing to ask it, and have it matte…

This is semi-offtopic, but "trust but verify" is an oxymoron. Trusting something means I don't have to verify whether it's correct (I trust that it is), so the saying, in the end, is "don't verify but verify".

It basically means "trust, but not too much."

Re: Anthropic Claude 3.5 can create icalendar files, so I did this

#135
post #3

You just have to double check the results whenever you tell Claude to extract lists and data. 99.9% of it will be correct, but sometimes 1 or 2 records are off. This kind of error is especially hard to notice because you're so impressed that Claude managed to do the extract task at all -- plus the results look wholly plausible upon eyeballing -- that you wouldn't expect anything to be wrong at all. But LLMs can get t…

The "off by one" predilection of LLMs is going to lead to this massive erosion of trust in whatever "Truth" is supposed to be, and it's terrifying and going to make for a bumpy couple of years. (Or the complete collapse of objective knowledge, on a long enough time horizon.) It's one thing to ask an LLM when George Washington was born, and have it return "May 20, 2020." It's another thing to ask it, and have it matte…

> The "off by one" predilection of LLMs is going to lead to this massive erosion of trust in whatever "Truth" is supposed to be, and it's terrifying and going to make for a bumpy couple of years.

This sounds like searching for truth is a bad thing, but instead is what has triggered every philosophical enquiry in history.

I'm quiet bullish, and think that LLMs will lead to a Renaissance in the concept of truth. Similar to what Wittgenstein did, Plato's cavern or late middle age empiricists.

Re: Anthropic Claude 3.5 can create icalendar files, so I did this

#136
post #114

Earlier quoted context omitted.

> "Trust" is to give a permission for someone to act on achieving some result. This would make the sentence "I asked him to wash the dishes properly, but I don't trust him", as your definition expands this to "I asked him to wash the dishes properly, but I didn't give him permission to achieve this result". If you say "I asked someone to do X but I don't trust them", it means you aren't confident they'll do it proper…

Why could I not say "I trusted him to do the dishes properly, after he was done, I verified, it's a good thing I trusted him to do the dishes properly, my supervision would have been unwarranted and my trust was warranted?" I trusted someone to do their task correctly, after the task was done, I verified my trust was warranted.

What would be different if you didn't trust them to do it correctly?

Re: Anthropic Claude 3.5 can create icalendar files, so I did this

#137
post #3

You just have to double check the results whenever you tell Claude to extract lists and data. 99.9% of it will be correct, but sometimes 1 or 2 records are off. This kind of error is especially hard to notice because you're so impressed that Claude managed to do the extract task at all -- plus the results look wholly plausible upon eyeballing -- that you wouldn't expect anything to be wrong at all. But LLMs can get t…

> 99.9% of it will be correct, but sometimes 1 or 2 records are off. This kind of error is especially hard to notice because you're so impressed that Claude managed to do the extract task at all -- plus the results look wholly plausible upon eyeballing -- that you wouldn't expect anything to be wrong at all. That's far better than I would do on my own. I doubt I'd even be 99% accurate. If it's really 99.9% accurate f…

The problem is that people could be impressed and use it for things where 0.01% could lead to people getting hurt or even get killed.

Re: Anthropic Claude 3.5 can create icalendar files, so I did this

#139

Earlier quoted context omitted.

> to me trust means exactly "I don't need to verify". If you use the slightly weaker definition that trust means you have confidence in someone, then the adage makes sense.

The issue here is that the only value of the adage is in the sleight of hand it lets you perform. If someone asks "don't you trust me?" (ie "do you have to verify what I do/say?"), you can say "trust, but verify!", and kind of make it sound like you do trust them, but also you don't really. The adage doesn't work under any definition of trust other than the one it's conflicting with itself about.

I think I just provided an example where it makes sense.

Specifically: I have confidence in your ability to execute on this task, but I want to check to make sure that everything is correct before we finalize.

Re: Anthropic Claude 3.5 can create icalendar files, so I did this

#140
post #114

Earlier quoted context omitted.

Why could I not say "I trusted him to do the dishes properly, after he was done, I verified, it's a good thing I trusted him to do the dishes properly, my supervision would have been unwarranted and my trust was warranted?" I trusted someone to do their task correctly, after the task was done, I verified my trust was warranted.

What would be different if you didn't trust them to do it correctly?

Instead of sitting in my office doing my work, then, spending a few minutes to verify once they're done, I'd sit in the kitchen next to them checking it as they went, being both distracted AND probably spending more time. I'd much rather trust but verify.
Post reply on HN