Live data from Hacker News

What it feels like to work with Mythos

oneusefulthing.org

91–100 of 337 posts

Re: What it feels like to work with Mythos

#91

Earlier quoted context omitted.

[flagged]

I'm happy to discuss arguments if you want to add any?

The only thing they’ve overtaken is arguably batteries, and even that is questionable if the quality is as good as Korean manufacturers. I think it’s more likely that the Chinese chip industry overtaking competitors will remain like nuclear fusion, forever “just 5 years away”

Re: What it feels like to work with Mythos

#92
post #38

What I find fascinating that there is so little substance in this article about the quality of produced code and the medium. Is the code documented and tested? Is it understandable and extendable? Is it secure? What language, framework, database was used? Author mentions judgement and taste - well, is the code tasteful? Will the model rearchitecture the entire thing if I ask it to add new functionality, spending anot…

I’m starting to realize that LLMs are really good at building low-stakes projects. Your questions mostly presume that the stakes are higher. The software will last a long time; the requirements will evolve; we can’t tolerate mistakes; etc. The trick to getting good at using LLMs for software is to learn how to make _all_ projects low-stakes.

You don't need LLM for that. You make _all_ projects low-stakes by working on green field project using (insert buzzword soup of the day) and leaving for a new green field opportunity (that requires experience with buzzword soup of the day) before the project ships.

Re: What it feels like to work with Mythos

#95

Earlier quoted context omitted.

now for the best question: whats your ROI here?

Humans are very expensive, so the equation almost always falls against them. It's not just salary, but also safety/labor regulation, legal risk, vacations, sick time, personal conflicts, HR, benefits. Even when automation is more expensive on paper, it's generally still cheaper

That's the beauty of these AI advancements. You, a human, will have to compete against a model for the same job.

If you get $100,000 per year as a SWE, and Anthropic offers a coding model for $100,000 per year (but working 24/7), then you'll have to give up all of those addons that make the fully burdened cost of the employee. Say goodbye to vacation, sick time, benefits, etc.

Re: What it feels like to work with Mythos

#96

Reading the first few paragraphs of what he calls "the most sophisticated academic social science paper I have yet seen from an AI" does not impress as much as I hoped. "Posterior beliefs about market demand are purely referencedependent: holding dollars raised constant, they track only performance relative to the founder’s self-chosen goal—jumping half a standard deviation at the threshold, responding steeply for th…

[dead]

Re: What it feels like to work with Mythos

#98
post #9

> It worked for nine and a half hours. > Again, it wasn’t perfect. As an expert, I was able to spot some errors and omissions (some as a result of the design I had asked for) that I had the AI correct That's the bit that stuck out to me - that's longer than I would expect to work on a problem in a day or even expect to go back & fix the output of something that has a core reward loop of hours. My customers are curren…

> At the same time, it is very dissonant to see the industry heading towards hour+ long workflows with an agent. At this point, pay me significantly more, and I'll do it.

> pay me significantly more

Ha ha, that's how you negotiate yourself out of a job!

Re: What it feels like to work with Mythos

#99
I'm using Fable this afternoon and it's definitely a step up from Opus 4.8, finding and fixing things Opus 4.8 was blind to even perceiving. The next 13 days are going to be fun IMO. And Opus 4.8 was less annoying than Opus 4.7 FWIW.

Edit: A couple hours in and I just got my first gaslighting attempt from the model. Good times!

Re: What it feels like to work with Mythos

#100
post #78

I am… underwhelmed by the artifacts in the post. I don’t see why working longer is a pro. The results don’t seem much better than you’d get from putting Opus in a long loop.

> The results don’t seem much better than you’d get from putting Opus in a long loop.

Care to share the results you got from Opus working on the same prompt? It should be easy to compare quality.

Post reply on HN