> Numbers like that buy a model a real migration effort. Such a silly choice of words. I wish the human directing the LLM writing the article put some effort into rewriting the worst examples of LLM style. > But it did extremely well, and the promise was immediate and specific: builds finishing in less than half the wall-clock time, at 27% lower cost, scoring at or above our incumbent on completed work. The way the L…
Can we get over the detective work about if the text was written by LLM or not in 2026 already ? This is a lost cause, and we could instead focus on substance over syntax.
Migrating a production AI agent to GPT-5.6: 2.2x faster, 27% cheaper
91–100 of 144 posts
Re: Migrating a production AI agent to GPT-5.6: 2.2x faster, 27% cheaper
#92> Ploy’s agent builds and edits real marketing websites. It plans a page, reads the codebase, writes components, generates imagery, screenshots its own work, and decides when it’s done. That job description sets a very high bar for a model, and we test every frontier release against it. For the four months Opus held the default slot (first Opus 4.7, then 4.8), nothing we tested beat it. Well, unlike OP I haven't run…
Re: Migrating a production AI agent to GPT-5.6: 2.2x faster, 27% cheaper
#93My experience mirrors this: services like OpenRouter that promise “failover” are pretty much useless except for sandbox testing because models in production are not really interchangeable. Any production harness doing serious agentic work in production is dependent on more model-specific quirks than you would expect. And even if another model works without errors, performance and efficiency is a whole different story…
Re: Migrating a production AI agent to GPT-5.6: 2.2x faster, 27% cheaper
#94Earlier quoted context omitted.
You are reading it wrong. Ask your LLM to read and summarize based on your style preferences. Better yet don’t read anything at all, just tell your agent to convert it to a skill file for it’s future reference..
Unironically, I think in the future we will have the option to run filters in our browsers that can reword articles to your preferred writing style. Like user stylesheets for text. I also assume it already exists out there, I've just been to lazy so far to look for it. It's a relatively obvious application of LLMs.
Re: Migrating a production AI agent to GPT-5.6: 2.2x faster, 27% cheaper
#95Re: Migrating a production AI agent to GPT-5.6: 2.2x faster, 27% cheaper
#96> Numbers like that buy a model a real migration effort. Such a silly choice of words. I wish the human directing the LLM writing the article put some effort into rewriting the worst examples of LLM style. > But it did extremely well, and the promise was immediate and specific: builds finishing in less than half the wall-clock time, at 27% lower cost, scoring at or above our incumbent on completed work. The way the L…
Re: Migrating a production AI agent to GPT-5.6: 2.2x faster, 27% cheaper
#97> The fix that worked is a schema transform at the provider boundary. For OpenAI-family models only, we rewrite every optional property to be required but nullable, using anyOf: [T, null], which gives the model an explicit way to say “not using this.” I admit, I've only used a bastardized form of MCP, but this smells... wrong? It's not clear to me why the Typescript type definitions would have any influence on (what…
Re: Migrating a production AI agent to GPT-5.6: 2.2x faster, 27% cheaper
#98Earlier quoted context omitted.
Can we get over the detective work about if the text was written by LLM or not in 2026 already ? This is a lost cause, and we could instead focus on substance over syntax.
To me it's a useful signal not to read an article that someone didn't bother to write. Which is a shame as real insights are buried inside some of these articles, which if the author bothered to write in his own words could have reached an audience that would have appreciated them. Writing is one of the areas where I want no LLM involvement.
The number of things that make it to the top of HN/Reddit/wherever now that are devoid a human's touch is exhausting. Whether it's a site that's got that Claude frontend smell, or a repo that's got a burst of 10 claude commits before getting shared and abandoned, or a series of blog posts that were written by LLMs... it's all, at this point, a flag for me that the human behind the LLM doesn't really want to engage with others or share; in some ways it dehumanises their entire (supposed) audience.
IDK. Maybe having Claude contribute writing about something novel to the general blogosphere is useful in some dimension, but it usually gives me no confidence in the truth of the post.
Re: Migrating a production AI agent to GPT-5.6: 2.2x faster, 27% cheaper
#99 As of today, Ploy’s agent runs on GPT-5.6 Sol, the flagship tier of the model family OpenAI released this morning.
Wait a moment, did they make the switch based on half a days of playing with Sol? Are these companies ran by teenagers?Re: Migrating a production AI agent to GPT-5.6: 2.2x faster, 27% cheaper
#100Earlier quoted context omitted.
You are reading it wrong. Ask your LLM to read and summarize based on your style preferences. Better yet don’t read anything at all, just tell your agent to convert it to a skill file for it’s future reference..
this is probably sarcasm, but might actually try this. the visceral negative reaction i have to llm writing makes me instantly want to close the tab
I would have downvoted the sarcasm (it doesn’t contribute to the conversation), but I believe it actually is the author’s opinion.