Live data from Hacker News

Shall I implement it? No

gist.github.com

11–20 of 603 posts

Re: Shall I implement it? No

#12
post #6
post #5

Why is this interesting? Is it a shade of gray from HN's new rule yesterday? https://news.ycombinator.com/item?id=47340079 Personally, the other Ai fail on the front of HN and the US Military killing Iranian school girls are more interesting than someone's poorly harnessed agent not following instructions. These have elements we need to start dealing with yesterday as a society. https://news.ycombinator.com/item?id=4…

Well, imagine this was controlling a weapon. “Should I eliminate the target?” “no” “Got it! Taking aim and firing now.”

It is completely irresponsible to give an LLM direct access to a system. That was true before and remains true now. And unfortunately, that didn't stop people before and it still won't.

Re: Shall I implement it? No

#13
post #7
post #5

Why is this interesting? Is it a shade of gray from HN's new rule yesterday? https://news.ycombinator.com/item?id=47340079 Personally, the other Ai fail on the front of HN and the US Military killing Iranian school girls are more interesting than someone's poorly harnessed agent not following instructions. These have elements we need to start dealing with yesterday as a society. https://news.ycombinator.com/item?id=4…

Opus being a frontier model and this being a superficial failure of the model. As other comments point out this is more of a harness issue, as the model lays out.

Exactly, the words you give it affect the output. You can get hem to say anything, so I find this rather dull

Re: Shall I implement it? No

#16
Claude is quite bad at following instructions compared to other SOTA models.

As in, you tell it "only answer with a number", then it proceeds to tell you "13, I chose that number because..."

Re: Shall I implement it? No

#18
post #4

Yeah this looks like OpenCode. I've never gotten good results with it. Wild that it has 120k stars on GitHub.

Does Claude Code's system prompt have special sauces?

Yes, very much so.

I've been able to get Gemini flash to be nearly as good as pro with the CC prompts. 1/10 the price 1/10 the cycle time. I find waiting 30s for the next turn painful now

https://github.com/Piebald-AI/claude-code-system-prompts

One nice bonus to doing this is that you can remove the guardrail statements that take attention.

Re: Shall I implement it? No

#19
post #9
post #6

Earlier quoted context omitted.

Well, imagine this was controlling a weapon. “Should I eliminate the target?” “no” “Got it! Taking aim and firing now.”

That's why we keep humans in the loop. I've seen stuff like this all the time. It's not unusual thinking text, hence the lack of interestingness

The human in the loop here said “no”, though. Not sure where you’d expect another layer of HITL to resolve this.

Re: Shall I implement it? No

#20
post #5

Why is this interesting? Is it a shade of gray from HN's new rule yesterday? https://news.ycombinator.com/item?id=47340079 Personally, the other Ai fail on the front of HN and the US Military killing Iranian school girls are more interesting than someone's poorly harnessed agent not following instructions. These have elements we need to start dealing with yesterday as a society. https://news.ycombinator.com/item?id=4…

How is this not clear?
Post reply on HN