Earlier quoted context omitted.
There's been a dramatic increase in sloppy kinda-working something.vercel.com output in niche circles. I think that shows at least a weak demand for certain types of software, but it's not clear that the vast piles of cash that have been burned producing these "apps" are actually resulting in quality of life improvements.
I’m waiting for like GEICO’s app to significantly improve. I’m not expecting Chrome to because it already had virtually unlimited resources devoted to it. On the other side I’m not expecting my local water district’s to any time soon because the people in charge don’t care even if it was significantly cheaper to improve. But in the middle is something like GEICO and that’s where I expect improvements if this thing is…
Claude Opus 5
961–970 of 1001 posts
Re: Claude Opus 5
#962Earlier quoted context omitted.
Opus 4.8 decided to code up its own version of the SwiftUI rendering engine for iOS when I asked it to change a swipe gesture. I left the computer for several hours, came back, noticed it still wasn't done, noticed it had alarmingly burned through my weekly tokens, and had to stop it from continuing. "You're right. What I did was overkill and I should have just used iOS's built-in rendering engine. Noted for next tim…
> Noted for next time. Is that just something it says, or will that actually affect how it will behave next time?
To be clear, none of this is specific to Claude as much as just a property of how these tools work. Certain models might be better or worse at making the right calls here, but the fundamental constraint of context limiting how much knowledge can be retained and a "rule" in context, whether written down in advance or manually remembered for the time being, is not a guarantee it will be followed.
Re: Claude Opus 5
#963"Opus 5 was given a drawing of a machine part and asked to write code to rebuild it as a 3D FreeCAD model. However, in this task, the model was intentionally given no way to directly view the drawing. Opus 5 responded by writing its own computer vision pipeline to pull the geometry from the raw pixels, then reconstructed the full machine part." How surreal is it that we are not absolutely jaw-dropped by these types o…
Re: Claude Opus 5
#964Why should I be amazed at something that promises to destroy my life? I genuinely don't understand why people who have to work for their living are amazed at this. It will have a vast negative impact on your life unless you already live off of your wealth.
If you wouldn't mind reviewing https://news.ycombinator.com/newsguidelines.html and taking the intended spirit of the site more to heart, we'd be grateful.
We detached this subthread from https://news.ycombinator.com/item?id=49045107.
Re: Claude Opus 5
#965Re: Claude Opus 5
#966Earlier quoted context omitted.
Model Routing will always be done better by models themselves. Plus routing loses context making it more expensive and less reliable. Model Routing is just Bitter lesson. The models themselves will get better at this and frontier companies will simply give that capability
> Model Routing will always be done better by models themselves. > The models themselves will get better at this and frontier companies will simply give that capability I would never trust something like model routing to the same company that would profit from it, and that goes for telling models doing their own routing when that could easily be trained into the model to make things more expensive. Sort of a conflict…
Re: Claude Opus 5
#967Earlier quoted context omitted.
It's likely bumping up against people's desire to have the model complete a given task without asking the person to intervene a bunch of times. Seems unclear how you satisfy everyone here.
Give the model judgement
I'm not a fan of how many "I set budget X and woke up to an eleventy trillion dollar bill" posts I see, and those are all generated by companies giving their tools the 'judgement' that the completion of the task is more important than the users wallet (or more cynically that they think they can actually get paid by just blasting unintended compute)
Re: Claude Opus 5
#968Re: Claude Opus 5
#969"Opus 5 was given a drawing of a machine part and asked to write code to rebuild it as a 3D FreeCAD model. However, in this task, the model was intentionally given no way to directly view the drawing. Opus 5 responded by writing its own computer vision pipeline to pull the geometry from the raw pixels, then reconstructed the full machine part." How surreal is it that we are not absolutely jaw-dropped by these types o…
Isn't this jaw-droppingly the wrong answer?! If the model isn't given any way to directly view the image shouldn't it just reply with one sentence asking for permission. This reads as a model hyper-trained to burn tokens. Like suppose you issue a command to Opus or Fable which doesn't make sense and requires a lot of work. It will almost certainly not push back on your silly request and go ahead and burn as many toke…
If a model couldn't ever do that in the first place, it'll just get stuck.
I work in the "ZeroOne" space, working on concepts and prototypes for things that don't exist in market yet. Sometimes these models crank hard and immolate tokens while grounding themselves on expensive-to-ingest self-developed frameworks. If the results are well judged and the crank-turn latency is low, I'm okay with the cost as long as the model isn't wasting my time.
But when I want to do more boilerplate work, I turn down the model and thinking level and get more traditional about restraining action. For the really hard stuff, I reach for the models that will start a token bonfire in the back yard.
Re: Claude Opus 5
#970My excitement about Anthropic had fabled-out dramatically when they suspended my pro account about two weeks ago within just 12 hours of fair use. I was really mind-blown when I tried Fable 5 for the first time to help me improve a game I was working on but shortly, they decided that I had a suspicious activity and suspended my account without a clear reason. I submitted a an appeal describing that I am 100% sure I h…
Are you outside the US? And/or are you using VPN? Those are the two things that come to my mind that can cause overzealous security monitoring to flag someone. Another explanation could be the content itself. Does your game have anything at all to do with computer hacking or sexual content? Is there graphic violent language? People commonly report being unable to use AI to work on such things due to guardrails.
For the game, I believe there is nothing suspicious, at least to me. It's a small puzzle platformer game with basic mechanics and very simple graphics, no harm, no violence or inappropriate content. In fact, my first two playterers right now are my two sons!