Live data from Hacker News

AGENTS.md outperforms skills in our agent evals

vercel.com

111–120 of 212 posts

Re: AGENTS.md outperforms skills in our agent evals

#111

Earlier quoted context omitted.

> Obviously directly including context in something like a system prompt will put it in context 100% of the time. How do you suppose skills get announced to the model? It's all in the context in some way. The interesting part here is: Just (relatively naively) compressing stuff in the AGENTS.md seems to work better than however skills are implemented.

Isn't the difference that a skill means you just have to add the script name and explanation to the context instead of the entire script plus the explanation?

Their non-skill based "compressed index" is just similarly "Each line maps a directory path to the doc files it contains" but without "skillification." They didn't load all those things into context directly, just pointers.

They also didn't bother with any more "explanation" beyond "here are paths for docs."

But this straightforward "here are paths for docs" produced better results, and IMO it makes sense since the more extra abstractions you add, the more chance of a given prompt + situational context not connecting with your desired skill.

Re: AGENTS.md outperforms skills in our agent evals

#114
post #86
post #60

Earlier quoted context omitted.

Cheaper? Loading every bit of documentation into context every time, regardless of whether it’s relevant to the task the agent is working on? How? I’d much rather call out the location of relevant docs in Claude.md or Agents.md and tell the agent to read them only when needed.

As they point out in the article, that approach is fragile. Cheaper because it has the right context from the start instead of faffing about trying to find it, which uses tokens and ironically bloats context. It doesn't have to be every bit of documentation, but putting the most salient bits in context makes LLMs perform much more efficiently and accurately in my experience. You can also use the trick of asking an LL…

> Extracting the most useful parts of documentation into a file

Yes, and this file becomes: also documentation. I didn’t mean throw entire unabridged docs at it, I should’ve been more clear. All of my docs for agents are written by agents themselves. Either way once the project becomes sufficiently complex it’s just not going to be feasible to add a useful level of detail of every part of it into context by default, the context window will remain fixed as your project grows. You will have to deal with this limit eventually.

I DO include a broad overview of the project in Agents or Claude.md by default, but have supplemental docs I point the agent to when they’re working on a particular aspect of the project.

Re: AGENTS.md outperforms skills in our agent evals

#116
post #103

[flagged]

This comment instantly set off my LLM alarm bells. Went into the profile, and guess what: next comment (not a one-liner) [0] on a completely different topic was posted 35 seconds later. And includes the classic "aren't just A. They're B.". Why are you doing this? Karma? 8 years old account and first post 3 days ago is a Show HN shilling your "AI agent" SaaS with a boatload of fake comments? [1] Pinging tomhow [0] htt…

Dude I am not AI. Real human. Just started on HN.

Re: AGENTS.md outperforms skills in our agent evals

#117
post #103

Earlier quoted context omitted.

This comment instantly set off my LLM alarm bells. Went into the profile, and guess what: next comment (not a one-liner) [0] on a completely different topic was posted 35 seconds later. And includes the classic "aren't just A. They're B.". Why are you doing this? Karma? 8 years old account and first post 3 days ago is a Show HN shilling your "AI agent" SaaS with a boatload of fake comments? [1] Pinging tomhow [0] htt…

Wow.

[dead]

Re: AGENTS.md outperforms skills in our agent evals

#118
post #103

Earlier quoted context omitted.

This comment instantly set off my LLM alarm bells. Went into the profile, and guess what: next comment (not a one-liner) [0] on a completely different topic was posted 35 seconds later. And includes the classic "aren't just A. They're B.". Why are you doing this? Karma? 8 years old account and first post 3 days ago is a Show HN shilling your "AI agent" SaaS with a boatload of fake comments? [1] Pinging tomhow [0] htt…

Dude I am not AI. Real human. Just started on HN.

[dead]
Post reply on HN