Live data from Hacker News

LLMs corrupt your documents when you delegate

arxiv.org

181–190 of 235 posts

Re: LLMs corrupt your documents when you delegate

#181
post #20

I'm suspicious of their results with regards to tool usage. It's unsurprising that round-tripping long content through an LLM results in corruption. Frequent LLM users already know not to do that. They claim that tool use didn't help, which surprised me... but they also said: > To test this, we implemented a basic agentic harness (Yao et al., 2022) with file reading, writing, and code execution tools (Appendix M). We…

I think your argument makes sense but my understanding is that adding the document to the context and spitting it back is prone to corruption in any scenario.

I think this is closely related to other sources saying that even if you have huge context the attention mechanism itself is not back referencing thus any tasks related to bigger contexts are prone to errors.

because I have some preconception of this maybe I am assuming this is what they were saying. Am I missing something ?

Re: LLMs corrupt your documents when you delegate

#182

Earlier quoted context omitted.

I agree with most of what you wrote except for this: >Frequent LLM users already know not to do that. And I think that’s the biggest problem. Amidst the current push to utilize LLMs across orgs and groups there are a large (if even say majority) of people that are using them every day but who have never approached anything as technical as a “harness” before let alone an entire setup. For them the behavior mentioned h…

Exactly. When I use a scissor, I don't want the scissor to not work just because I'm not a "frequent scissor user," and then get told by someone who makes their breakfast with scissors that I'm doing it wrong. Most people will not be "frequent" anything users.

If you make the example any more complicated, it makes sense.

A lathe operator isn’t any good if they don’t frequently operate lathes.

A articulated robot implementer needs frequent experience implementing robots to be any good.

That doesn’t mean lathes or robots are useless. Nor does it mean they have failed as products because they require expertise.

I do think it raises questions as to whether vast swathes of the population will be effective at using LLMs. Are they scissors, or a lathe?

Re: LLMs corrupt your documents when you delegate

#184

Earlier quoted context omitted.

Exactly. When I use a scissor, I don't want the scissor to not work just because I'm not a "frequent scissor user," and then get told by someone who makes their breakfast with scissors that I'm doing it wrong. Most people will not be "frequent" anything users.

If you make the example any more complicated, it makes sense. A lathe operator isn’t any good if they don’t frequently operate lathes. A articulated robot implementer needs frequent experience implementing robots to be any good. That doesn’t mean lathes or robots are useless. Nor does it mean they have failed as products because they require expertise. I do think it raises questions as to whether vast swathes of the…

Everybody seems to want them to be scissors, or at least to treat them as such, but even still the reason everyone can use scissors so well is because they’ve practiced with them, right? You’re probably a lot better at using scissors now than the first time you did it, the functionality is just so simple it’s harder to notice.

To me learning to use LLMs is the same as doing anything else, you have to practice and put in the hours to get good. Maybe some harnesses will eventually allow LLMs to function more as scissors than lathes. This seems to be what Microsoft is trying to do by embedding Copilot in all their products and saying “choose the UI that works best for you”. If that doesn’t end up working we’ll need another paradigm for “non-technical” users to effectively operate computer assistants

Re: LLMs corrupt your documents when you delegate

#185

Earlier quoted context omitted.

Exactly. When I use a scissor, I don't want the scissor to not work just because I'm not a "frequent scissor user," and then get told by someone who makes their breakfast with scissors that I'm doing it wrong. Most people will not be "frequent" anything users.

Most people also understand that, because they're not "frequent" users of a thing, they absolutely suck at using it, and set their expectations accordingly. In particular, they realize that doing anything non-trivial with the thing requires them to spend some learning and practice time, or asking/hiring a "frequent" user to do it for them. So the reasonable response to being told you're holding your scissors wrong is…

I’m interested in the “non-trivial” point as well, this seems to be a common refrain from the anti-LLM tech crowd, “LLMs aren’t good at doing anything non-trivial”, well is that really the case or is it just harder and one needs to put in more practice for more complicated tasks?

I don’t have an example off hand, but I know that it’s easy to dismiss something an LLM does as trivial if your work is extremely marginal. Most devs aren’t creating their own programming languages. I can’t help but think people who hold this opinion also think the work most software professionals do is “trivial” (“you’re just moving strings around, that’s not impressive/trivial”)

Re: LLMs corrupt your documents when you delegate

#187

Earlier quoted context omitted.

> This will change too man. Maybe I am in a bubble but with how fast things are changing, it won’t be too long before the bubble becomes reality. You can’t get mad at an experiment for not happening in the future. > Either way we should be doing experiments on the actual capabilities of AI They simulated common end user behavior >because it helps validate your own negative bias against AI. We’ve gone from “this study…

> You can’t get mad at an experiment for not happening in the future. I’m more getting mad at this sentence not making any sense. I’m disappointed at this experiment for not testing the actual capabilities of an LLM. Comprende? > They simulated common end user behavior Not the way you use it. And not the way it will be used. You love it because you want it to stay this way so you can forever believe AI will never be…

>You love it because you want it to stay this way so you can forever believe AI will never be better than you.

>Bro the reality is unfolding as you speak

>You go pretend you live in that reality where the bullets will never appear.

It’s too late bro, roko’s basilisk was real and it’s already punishing you

Re: LLMs corrupt your documents when you delegate

#188
post #20

I'm suspicious of their results with regards to tool usage. It's unsurprising that round-tripping long content through an LLM results in corruption. Frequent LLM users already know not to do that. They claim that tool use didn't help, which surprised me... but they also said: > To test this, we implemented a basic agentic harness (Yao et al., 2022) with file reading, writing, and code execution tools (Appendix M). We…

[flagged]

Re: LLMs corrupt your documents when you delegate

#189

Earlier quoted context omitted.

I think you’re living in a bubble if you think the average user of AI even knows what a harness is The vast majority of people are literally going to chatGPT, pasting in their document and asking for edits.

This will change too man. Maybe I am in a bubble but with how fast things are changing, it won’t be too long before the bubble becomes reality. Either way we should be doing experiments on the actual capabilities of AI not about the stupidest possible way to use AI because it helps validate your own negative bias against AI. Additionally as software engineers using agentic AI… which HN basically is… this experiment i…

First off, It’s good to study all kinds of things isn’t it? Even if it’s not strictly practical.

Second, and more importantly these AI tools are EVERYWHERE right now. The effects of people using them for work can be seen throughout many industries and workplaces.

So I think studying how these models perform in the vast majority of use cases is not only a good idea, but it’s actually really important.

Even if you’re strictly pro-AI and believe it is the future, a study like this can help you explain to laymen why they need the harnesses you’re so in support of.

Re: LLMs corrupt your documents when you delegate

#190
In my experience there's no longer any good reason to post research papers investigating limitations of LLMs on HN any more because they are always met with one, or all, of the following arguments that have now taken the status of thought-terminating clichés:

1. It's an older model.

2. You're prompting it wrong.

3. That's not what LLMs are for.

4. We knew that already.

It's as if there LLMs have no limitations, which of course goes completely against number 1 in the list above, because if LLMs have no limitations then how are newer models better and why are AI companies constantly releasing new versions?

But the debate has taken on an insidious identitarian character: it's no longer about understanding a technology, its strengths, its limitations, what makes it tick. It's a fractious internet fight between crowds of users who have attached this or that opinion to their very internet persona and will not budge from their entrenched positions.

That is basically the death of curious debate. Obviously there's no point in discussing any research under those conditions: good or bad, flawed or not, we're just not going to get any signal out of the noise on HN anymore.

Post reply on HN