Live data from Hacker News

Agent Skills

addyosmani.com

191–200 of 239 posts

Re: Agent Skills

#191

Cant wait for everyone to realize they've wasted a year + messing with agents and experiencing a feeling of psuedo productivity.

At least we’ll be thankful for all the documentation developers have written in order to feed Claude better context.

Maybe the productivity we were trying to achieve was the friends we made along the way

Re: Agent Skills

#192
post #6

From an SEO/LLMO perspective, the discoverability of these skills will be difficult without a rename: https://agentskills.io/ If Addy reads this, how do you pitch this vs. Superpowers? https://github.com/obra/superpowers

I would love to know how many people are actually using superpowers. I showed up on the agentic dev scene prior to superpowers, and I am getting concerned that >50% of my self-rolled processes are now covered by superpowers. I no longer trust gh stars, can anyone chime in? Is superpowers now truly adopted? If it is truly valuable, why hasn't Boris integrated the concepts yet?

A lover of Superpowers here, using it for about 2+ months now.

It allows you to explore the problem space upfront, it questions your assumptions, asks more probing questions to confirm what it’s found in the code, and by the time you’re ready to implement, it knows exactly what needs to happen.

Jessie should have called it Socrate’s Methods

Re: Agent Skills

#193
post #122
post #68

Snake oil. Good to read for sure. Seems all plausible too. But snake oil nevertheless. Here's why: The slot machine can drop any hard requirement that you specifically in your AGENTS.md, memory.md or your dozens of skill markdowns. Pretty much guaranteed. These harnesses approaches pretend as if LLMs are strict and perfect rule followers and the only problem is not being able to specify enough rules clearly enough. T…

Snake oil may be a bit strong, because snake oil never works (except maybe as placebo?) whereas anything with an LLM, even though stochastic, has a pretty high chance of working. > ... you also realize that promised productivity gains are also snake oil because reading code and building a mental model is way harder than having a mental model and writing it into code. Not really, though it depends on the code; reading…

I think the placebo effect might be a decent comparison. It works most of the time, and you don't worry about it as long as you fully believe in its efficacy. However, once the illusion is shattered, the positive effects are diminished, and you can never fully trust the solution again.

Re: Agent Skills

#194

Cant wait for everyone to realize they've wasted a year + messing with agents and experiencing a feeling of psuedo productivity.

I’m a bit curious with these takes. Arguing in good faith - is the general assumption that people who use AI/agents/harnesses don’t ship features? We’ve been all in Claude Code since ~Septemberish, and have been able to successfully track the boost. Like the features that we ship that get used in production. Both from infrastructure side, and business logic implementations. Frontend and backend. I don’t think people…

I suspect some devs don't want AI to succeed - and it's understandable, as it will fundamentally change the way they work, and possibly put them out of a job as we need less developers.

So they convince themselves AI can't work because they don't want it to.

Re: Agent Skills

#195

Earlier quoted context omitted.

Skills are often invoked imperatively by the user. In cases where they are intended to be used directly by the LLM, it would be included somewhere else in the context. E.g: ``` After implementing the feature, read the testing skill for instructions on how to test. ```

how do you guarantee that the LLM follows an instruction given imperatively by the user? It probably will, but this is not guaranteed behavior. Likewise, _how_ it follows that instruction is non-deterministic. it's turtles all the way down.

You isn't gaurentee it any more than you can guarantee your prompt gives the output you want. Skills are just prompt templates.

Re: Agent Skills

#196
post #125

Earlier quoted context omitted.

I can take on a slightly weaker form in good faith: professionally it’s a non-starter until private, open source inference can be self-hosted and the ROI is clear enough to invest in that. And on the ROI side, trying things out regularly, I haven’t seen the positive ROI in the limited time I’ve dedicated to exploring the tools. I’ve restricted experimenting to 4 hours per month, because spending more than 2.5% of the…

"I studied math 4 hours per month and I can confirm that mathematics is stupid" You can't learn how to use _anything_ by experimenting 4 hours a month.

I think I should also clarify, I work in the training of encoder-decoder transformer models. Before the ChatGPT era I worked on on encoder-only transformer models. I'm not unfamiliar with the literature and general discourse. I just do not use LLMs for programming.

Re: Agent Skills

#197
post #149

Earlier quoted context omitted.

Liability, penalties, trust, and responsibility are means we use to try to influence the application of the processes that do. They do not directly affect reliability. They can be applied just as much to a team using AI as one that does not. Review and oversight does address reliability directly, and hence why we make use of those in processes to improve the reliability of mechanical processes as well, and why they a…

> Liability, penalties, trust, and responsibility are means we use to try to influence the application of the processes that do. They do not directly affect reliability. They can be applied just as much to a team using AI as one that does not. Yes and no. see next point. > You can ask the same thing about all the supporting staff around the experts in your team. I have a good idea of the shape of errors for a human b…

> (there is no singular thread behind my comment. I think we probably have more in agreement than not, and its more a question of finding the precise words to declare the shapes we perceive.)

I moved this up top, because I agree, despite the length of the below:

> However, the current hype cycle has created expectations of reliability from LLMs that drive 'Automated Intelligence' styled workflows.

Because for a lot of things it works. Today. I have a setup doing mostly autonomous software development. I set direction. I don't even write specs. It's not foolproof yet by any means - that is on the edge of what is doable today. Dial it back just a little bit, and I have projects in production that are mostly AI written, that have passed through rigorous reviews from human developers.

The key thing is that you can't "vibecode" that. I'm sure we agree there.

There needs to be a rigorous process behind it, and I think we'll agree on that too.

Those processes are largely the same as the processes required for human developers. Only for human developers we leave a lot of that process "squishy" and under-specified.

We trust our human developers to mostly do the right thing, even though many don't, and to not need written checklists and controls, even though many do.

What is coming out of this is a start of systems that codify processes that are very much feels based with human teams. Partly because we still need to codify them for AI, but also because we can - most people wouldn't want to work in the kind of regimented environment we can enforce on AI.

Sure, there is a lot of hype from people who just want to throw random prompts at an LLM and get finished software out. That is idiocy. Even a super-intelligent future AI can't read minds.

But there are a lot of people building harnesses to wrap these LLMs in process and rigor to squeeze as much reliability as possible from them, and it turns out you can leverage human organisational knowledge to get surprisingly far in that respect.

Re: Agent Skills

#198
post #136

Earlier quoted context omitted.

Because certain aspects (both are error prone) are similar and comparable. The notion that two entities need to be close in abilities for it to be possible to compare them is nonsense. You make the point for me: We managed to put men on the moon despite humans being enormously unreliable and error prone, because we built system around them that allowed for harnessing the good bits and reducing the failures to accepta…

> We are - I am anyway - using our lessons from building reliable systems from unreliable elements to raise the reliability of outputs of LLMs the same way. :) :) :) I could tell immediately you are somehow vested in the "success" of the LLM. So 600 B dollars and five years later, can you tell me how far did you guys get? Apollo programme costed a tiny fraction of that and started putting people on the moon some ~10…

I wish I had used 600B. I've spent a few thousand, and my efforts are very much profitable and earning me a substantial living right now.

Re: Agent Skills

#199

Cant wait for everyone to realize they've wasted a year + messing with agents and experiencing a feeling of psuedo productivity.

I can understand skepticism to a degree, and even fundamentally believing that AI is bad for all sorts of reasons, but I am becoming more and more perplexed at the certainty behind statements like this one. How are you so certain that AI development is this doomed? It just hasn't matched my experience at all, and I wonder what your experience is that has driven you to this level of certainty about the certain doom of…

I dont know any serious engineers thay are doing real work with AI agents. I know some that are building features for web applications and just punching a clock, but I don't think that constitutes real work or provides much value to the world.

I like thinking, solving problems and typing out code myself. Im going to keep putting tons of care into my craft and I promise I'll have more impact than the guy running 3 agents to build the 500th version of some web concept.

Rolex has a much bigger impact on the world than white label mass manufacturers in China.

Re: Agent Skills

#200

Earlier quoted context omitted.

I’m a bit curious with these takes. Arguing in good faith - is the general assumption that people who use AI/agents/harnesses don’t ship features? We’ve been all in Claude Code since ~Septemberish, and have been able to successfully track the boost. Like the features that we ship that get used in production. Both from infrastructure side, and business logic implementations. Frontend and backend. I don’t think people…

You're replying to an account specifically created to post inflammable AI takes (likely a bot anyway). So your attempt > Arguing in good faith will be futile, unfortunately.

[deleted]
Post reply on HN