This never happened with Opus 4.5 despite a lot of usage.
Claude Opus 4.6
831–840 of 1001 posts
Re: Claude Opus 4.6
#832Earlier quoted context omitted.
Surely the corpus Opus 4.6 ingested would include whatever reference you used to check the spells were there. I mean, there are probably dozens of pages on the internet like this: https://www.wizardemporium.com/blog/complete-list-of-harry-p... Why is this impressive? Do you think it's actually ingesting the books and only using those as a reference? Is that how LLMs work at all? It seems more likely it's predicting t…
So a good test would be replacing the spell names in the books with made-up spells. And if a "real" spell name was given, it also tests whether it "cheated".
Re: Claude Opus 4.6
#833I'm still not sure I understand Anthropic's general strategy right now. They are doing these broad marketing programs trying to take on ChatGPT for "normies". And yet their bread and butter is still clearly coding. Meanwhile, Claude's general use cases are... fine. For generic research topics, I find that ChatGPT and Gemini run circles around it: in the depth of research, the type of tasks it can handle, and the qual…
I really like that Claude feels transactional. It answers my question quickly and concisely and then shuts up. I don't need the LLM I use to act like my best friend.
I recently compared a class that I wrote for a side project that had quite horrible temporal coupling for a data processor class.
Gemini - ends up rating it a 7/10, some small bits of feedback etc
Claude - Brutal dismemberment of how awful the naming convention, structure, coupling etc, provides examples how this will mess me up in the future. Gives a few citations for python documentation I should re-read.
ChatGPT - you're a beautiful developer who can never do anything wrong, you're the best developer that's ever existed and this class is the most perfect class i've ever seen
Re: Claude Opus 4.6
#834Re: Claude Opus 4.6
#835Earlier quoted context omitted.
I was playing about with Chat GPT the other day, uploading screen shots of sheet music and asking it to convert it to ABC notation so I could make a midi file of it. The results seemed impressive until I noticed some of the "Thinking" statements in the UI. One made it apparent the model / agent / whatever had read the title from the screenshot and was off searching for existing ABC transcripts of the piece Ode to Joy…
Sounds pretty human like! Always searching for a shortcut
Re: Claude Opus 4.6
#836Re: Claude Opus 4.6
#837Earlier quoted context omitted.
Are we sure the docs page has been updated yet? Because that page doesn't say anything about automatic recording of memories.
Oh, quite right. I saw people mention MEMORY.md online and I assumed that was the doc for it, but it looks like it isn't.
Re: Claude Opus 4.6
#838Re: Claude Opus 4.6
#839Earlier quoted context omitted.
No? The hardest part of my SWE job is not the actual coding.
Even for coding, it seems to still make A LOT of mistakes. https://youtu.be/8brENzmq1pE?t=1544 I feel like everyone is counting chickens before they hatch here with all the doomsday predictions and extrapolating LLM capability into infinity. People that seem to overhype this seem to either be non-technical or are just making landing pages.
Re: Claude Opus 4.6
#840I’m very worried about the problems this will cause down the road for people not fact checking or working with things that scream at them when they’re wrong.