[flagged]
> How long before someone pitches the idea that the models explicitly almost keep solving your problem to get you to keep spending? -gtowey
31–40 of 1001 posts
[flagged]
> How long before someone pitches the idea that the models explicitly almost keep solving your problem to get you to keep spending? -gtowey
[flagged]
Has anyone tested how good the 1M context window is? i.e given an actual document, 1M tokens long. Can you ask it some question that relies on attending to 2 different parts of the context, and getting a good repsonse? I remember folks had problems like this with Gemini. I would be curious to see how Sonnet 4.6 stands up to it.
It's wild that Sonnet 4.6 is roughly as capable as Opus 4.5 - at least according to Anthropic's benchmarks. It will be interesting to see if that's the case in real, practical, everyday use. The speed at which this stuff is improving is really remarkable; it feels like the breakneck pace of compute performance improvements of the 1990s.
simonw hasn't shown up yet, so here's my "Generate an SVG of a pelican riding a bicycle" https://claude.ai/public/artifacts/67c13d9a-3d63-4598-88d0-5...
My take away is: it's roughly as good as Opus 4.5. Now the question is: how much faster or cheaper is it?
I wonder what difference have people found with sonnet 4.5 and opus 4.5 and probably similar delta will remain. Was sonnet 4.5 much worse than opus?
It's wild that Sonnet 4.6 is roughly as capable as Opus 4.5 - at least according to Anthropic's benchmarks. It will be interesting to see if that's the case in real, practical, everyday use. The speed at which this stuff is improving is really remarkable; it feels like the breakneck pace of compute performance improvements of the 1990s.
It's wild that Sonnet 4.6 is roughly as capable as Opus 4.5 - at least according to Anthropic's benchmarks. It will be interesting to see if that's the case in real, practical, everyday use. The speed at which this stuff is improving is really remarkable; it feels like the breakneck pace of compute performance improvements of the 1990s.
My take away is: it's roughly as good as Opus 4.5. Now the question is: how much faster or cheaper is it?
[flagged]
Being just sum guy, and not in the industry, should I share my findings?
I find it utterly fascinating, the extent to which it will go, the sophisticated plausible deniability, and the distinct and critical difference between truly emergent and actually trained behavior.
In short, gpt exhibits repeatably unethical behavior under honest scrutiny.