Live data from Hacker News

Agents.md file isn't the problem. Your lack of Evals is

tessl.io

1–10 of 18 posts

Re: Agents.md file isn't the problem. Your lack of Evals is

#3
post #2

Okay, but how would I write evals for my project's agents file? Any good examples out there?

I wrote https://ai-evals.io (community site) to make the concept approachable no matter what tools you choose to use.

You can learn about them evaluating that site https://github.com/Alexhans/eval-ception and then the pattern should be easy to test on your own thing.

Re: Agents.md file isn't the problem. Your lack of Evals is

#4
post #2

Okay, but how would I write evals for my project's agents file? Any good examples out there?

The agents are smart enough to write the evals too.

It's agents all the way down!

Submit a GitHub repo containing skills to Tessl, and it will generate the evals, run them, and present the results. https://tessl.io/registry/skills/submit

The evals and results are all shown, no login necessary, so you can assess them yourself. e.g. https://tessl.io/registry/skills/github/coreyhaines31/market... (click details to see the eval texts).

Re: Agents.md file isn't the problem. Your lack of Evals is

#7
post #4
post #2

Okay, but how would I write evals for my project's agents file? Any good examples out there?

The agents are smart enough to write the evals too. It's agents all the way down! Submit a GitHub repo containing skills to Tessl, and it will generate the evals, run them, and present the results. https://tessl.io/registry/skills/submit The evals and results are all shown, no login necessary, so you can assess them yourself. e.g. https://tessl.io/registry/skills/github/coreyhaines31/market... (click details to see t…

At first glance this looks like an entire ecosystem full of slop and by running that eval you generate more? I'm looking for something a bit more curated.

Re: Agents.md file isn't the problem. Your lack of Evals is

#8
post #3
post #2

Okay, but how would I write evals for my project's agents file? Any good examples out there?

I wrote https://ai-evals.io (community site) to make the concept approachable no matter what tools you choose to use. You can learn about them evaluating that site https://github.com/Alexhans/eval-ception and then the pattern should be easy to test on your own thing.

Doing an eval on itself is clever but confusing for the reader. How about a tutorial explaining how to do an evals on something more normal?

Re: Agents.md file isn't the problem. Your lack of Evals is

#10
so how would you eval your own claude.md? Each context is unique to the project, team, and personal root claude.md. Do you just take given task and ask it to redo the same one over and over again against a known solution? Do you just keep using it and "feel" whether or not it's working? How is that different from what everyone is already doing?
Post reply on HN