AGENTS.md outperforms skills in our agent evals
1–10 of 212 posts
Re: AGENTS.md outperforms skills in our agent evals
#2Re: AGENTS.md outperforms skills in our agent evals
#3TFA says they added an index to Agents.md that told the agent where to find all documentation and that was a big improvement.
The part I don't understand is that this is exactly how I thought skills work. The short descriptions are given to the model up-front and then it can request the full documentation as it wants. With skills this is called "Progressive disclosure".
Maybe they used more effective short descriptions in the AGENTS.md than they did in their skills?
Re: AGENTS.md outperforms skills in our agent evals
#4This is confusing. TFA says they added an index to Agents.md that told the agent where to find all documentation and that was a big improvement. The part I don't understand is that this is exactly how I thought skills work. The short descriptions are given to the model up-front and then it can request the full documentation as it wants. With skills this is called "Progressive disclosure". Maybe they used more effecti…
Re: AGENTS.md outperforms skills in our agent evals
#5The agent passes the Turing test...
Re: AGENTS.md outperforms skills in our agent evals
#6Skills are new. Models haven't been trained on them yet. Give it 2 months.
Re: AGENTS.md outperforms skills in our agent evals
#7It’s really silly to waste big model tokens on throat clearing steps
Re: AGENTS.md outperforms skills in our agent evals
#8Re: AGENTS.md outperforms skills in our agent evals
#9In a month or three we’ll have the sensible approach, which is smaller cheaper fast models optimized for looking at a query and identifying which skills / context to provide in full to the main model. It’s really silly to waste big model tokens on throat clearing steps
Re: AGENTS.md outperforms skills in our agent evals
#10Isn't it obvious that an agent will do better if he internalizes the knowledge on something instead of having the option to request it? Skills are new. Models haven't been trained on them yet. Give it 2 months.