Live data from Hacker News

Handbook.md shows that long policy documents do not reliably govern agents

arxiv.org

191–200 of 237 posts

Re: Handbook.md shows that long policy documents do not reliably govern agents

#191

Earlier quoted context omitted.

So you are essentially saying "you can't, you can only safeguard from effects of LLM eventually ignoring it"

I think things like CI, linters, hooks, etc started as defense against humans eventually ignoring it (low blood sugar, under pressure, never cared in the first place, etc). These seem like they would naturally extend to LLMs as well. (I'm not pro-LLM on balance, but the solutions of today feel similar for humans and LLMs.)

It is almost exactly like the controls from the auditing world, except as applied to AI's instead of people. In fact, having worked at Deloitte in the past is what led me to make the connection immediately

Re: Handbook.md shows that long policy documents do not reliably govern agents

#192
post #176

Earlier quoted context omitted.

We're on the cusp of Kimi K3 becoming usable on sub-10k hardware. https://github.com/gavamedia/deltafin

14 seconds per token? Not tokens per second. Seconds per token? That’s nowhere near the cusp!

It's need it to be an order of magnitude cheaper, but for some queries ("What to discuss at tomorrow's meeting") I can wait 12+ hours.

Re: Handbook.md shows that long policy documents do not reliably govern agents

#193
post #176

Earlier quoted context omitted.

We're on the cusp of Kimi K3 becoming usable on sub-10k hardware. https://github.com/gavamedia/deltafin

14 seconds per token? Not tokens per second. Seconds per token? That’s nowhere near the cusp!

At that rate, it would take me a mere 15 days to generate the number of tokens I typically use in a day.

"Yes, I'm using AI to speed up development. I'll submit that PR in two weeks time!"

Re: Handbook.md shows that long policy documents do not reliably govern agents

#194

Earlier quoted context omitted.

> The parent comment wasn't suggesting that the inherent defects magically disappear that is almost exactly what they said though..? " Want it to go away, almost like magic? Local inference. [...] all of the common LLM defects will go away.

"[...]" ----- edit: I suppose I have to spell this out. "[...]" = "When its under your control, and [you're] no longer being forced to hold it wrong," ≈ "liberates your use cases from all the horribly opaque configuration, shadow prompting, etc." and "Almost like magic" is cliché hyperbole. It also means "not magic" in the same way that "almost like a dog" means "not a dog." It was immediately followed by the explici…

i omitted it because it doesnt change my point at all and it makes the quoted text smaller. not because of "absolute and conscious bad faith".

people can just read the original from the comment. its not like im quoting from a different website and trying to hide something. look 2 comments above mine and it is right there in full. anyone reading my comment would have seen the original comment already...

i have no idea why you took such offense to my comment, but you're reading way too much into it.

Re: Handbook.md shows that long policy documents do not reliably govern agents

#195

Earlier quoted context omitted.

> There are no "attention heads" Yes there quite literally is internally in an LLM.

https://bactra.org/notebooks/nn-attention-and-transformers.h... Just because something is called an "attention head" doesn't mean the terminology makes sense. If I'm reading the article right, hungryhobbit is accurately describing a single attention head. And... what having multiple attention heads means is that you do the "single attention head" thing several times, and average the results. There is no part of hungr…

Off topic, but I'm kinda sad there's only one Cosma Shalizi in the world. That dude is a treasure, I've been a fan for decades now.

Re: Handbook.md shows that long policy documents do not reliably govern agents

#198
post #176

Earlier quoted context omitted.

14 seconds per token? Not tokens per second. Seconds per token? That’s nowhere near the cusp!

It's need it to be an order of magnitude cheaper, but for some queries ("What to discuss at tomorrow's meeting") I can wait 12+ hours.

I'd rather just use my actual human brain to compute the answer at that point. I don't see the value at throughput that is this low.

Re: Handbook.md shows that long policy documents do not reliably govern agents

#199

Earlier quoted context omitted.

> The parent comment wasn't suggesting that the inherent defects magically disappear that is almost exactly what they said though..? " Want it to go away, almost like magic? Local inference. [...] all of the common LLM defects will go away.

"[...]" ----- edit: I suppose I have to spell this out. "[...]" = "When its under your control, and [you're] no longer being forced to hold it wrong," ≈ "liberates your use cases from all the horribly opaque configuration, shadow prompting, etc." and "Almost like magic" is cliché hyperbole. It also means "not magic" in the same way that "almost like a dog" means "not a dog." It was immediately followed by the explici…

> "Almost like magic" is cliché hyperbole. It also means "not magic" in the same way that "almost like a dog" means "not a dog." It was immediately followed by the explicit clarification that was surgically omitted above, in absolute and conscious bad faith.

Incorrect. The omitted text doesn't change the relevant meaning of the quotation in the slightest. Why are you taking such offense?

> Want it to go away, almost like magic? Local inference. When its under your control, and no longer being forced to hold it wrong, all of the common LLM defects will go away.

> "Want it to go away, almost like magic? Local inference. [...] all of the common LLM defects will go away.

There's no difference between these two statements in the relevant context of the article, which is specifically "long policy documents do not reliably govern agents", or the comments above, which are specifically addressing how that issue is not solved by local agents.

Re: Handbook.md shows that long policy documents do not reliably govern agents

#200

Earlier quoted context omitted.

> Well the answer is that VC-backed companies Look, I enjoy local LLMs as much as anyone, but I think there is some motivated reasoning happening in this thread to try to make local LLMs sound like a utopia against those evil VCs. Local LLMs suffer from the same problems.

And I think you didn't understand what I wrote, so let me reiterate: Of course policies are equally ineffective at strictly governing the behavior of local models. Why would that change? The actual problem is that somebody is attempting to misuse them for that in the first place. There are only two reasons for it: - They have no other options OR - They have no idea what they're doing Local models solve the first prob…

> In-context, your reply reads like you're asserting that anthropic's offerings give you all the same control as a local model

It absolutely does not.

Post reply on HN