Live data from Hacker News

What it feels like to work with Mythos

oneusefulthing.org

21–30 of 337 posts

Re: What it feels like to work with Mythos

#21
post #9

> It worked for nine and a half hours. > Again, it wasn’t perfect. As an expert, I was able to spot some errors and omissions (some as a result of the design I had asked for) that I had the AI correct That's the bit that stuck out to me - that's longer than I would expect to work on a problem in a day or even expect to go back & fix the output of something that has a core reward loop of hours. My customers are curren…

My Opus 4.8 regularly works for 10+minutes on a single non-trivial coding request.

Your Opus 4.8? Is it now usual to refer to LLMs like that?

Re: What it feels like to work with Mythos

#23
Anecdote: I fed Fable some models I’ve been hand verifying (basically, I sketch out a scenario for Opus to model, it builds it, I ask it to show me the math, I correct it, we iterate like this, then I double check its code to make sure the math matches the model logic). Fable found almost every error I found, and then had some interesting suggestions for additional variables.

It also burned through my usage quota like a late-90s Hummer.

Re: What it feels like to work with Mythos

#25
post #9

> It worked for nine and a half hours. > Again, it wasn’t perfect. As an expert, I was able to spot some errors and omissions (some as a result of the design I had asked for) that I had the AI correct That's the bit that stuck out to me - that's longer than I would expect to work on a problem in a day or even expect to go back & fix the output of something that has a core reward loop of hours. My customers are curren…

> At the same time, it is very dissonant to see the industry heading towards hour+ long workflows with an agent.

At this point, pay me significantly more, and I'll do it.

Re: What it feels like to work with Mythos

#26

Earlier quoted context omitted.

My Opus 4.8 regularly works for 10+minutes on a single non-trivial coding request.

Your Opus 4.8? Is it now usual to refer to LLMs like that?

Isn't it common to refer to all software like that? "Let my look at my JIRA", "I can't find anything using my Outlook's search function", "My Powerpoint is acting up today", "My browser just crashed" are all sentences I might say during a normal work day

Re: What it feels like to work with Mythos

#28
post #11
post #9

> It worked for nine and a half hours. > Again, it wasn’t perfect. As an expert, I was able to spot some errors and omissions (some as a result of the design I had asked for) that I had the AI correct That's the bit that stuck out to me - that's longer than I would expect to work on a problem in a day or even expect to go back & fix the output of something that has a core reward loop of hours. My customers are curren…

In Claude's defense (and I cannot believe I'm defending it), I know no single dev who could create what it did (Concord), from a 19-page design document, in 9.5 working hours. We're gonna go back to the days where our bosses ask why we're just sitting around, but instead of saying "compiling," we'll just say, "waiting for Claude."

This. I get told things like "you can't build all that on your own?" I've had Claude poop out full feature web apps in under 30 minutes, to a spec. Was it perfect? No, but sometimes even in a simple setup phase you can burn 15 minutes to some obscure setup step that's failing. I cannot just code nonstop at 900WPM or whatever ridiculous speed, and poop out an entire full feature web app, with maybe a few bugs here or there. If you can, come show me, I'll gladly have you race against my Claude prompting capabilities.

Will Claude's code be perfect in one shot? Probably not, will it get you 80 to 90% of the way there with your chosen design patterns in under a few hours? Absolutely.

Post reply on HN