I've heard rumors this might be Sonnet 5 rebranded as Opus 4.6. But why? Profit? WDYT?
Calling it part of the Sonnet line would not provide the same level of blind buy in as calling it part of the Opus line does
931–940 of 1001 posts
I've heard rumors this might be Sonnet 5 rebranded as Opus 4.6. But why? Profit? WDYT?
Calling it part of the Sonnet line would not provide the same level of blind buy in as calling it part of the Opus line does
Earlier quoted context omitted.
Insane to think that a relatively simple CLI tool has so many open issues...
sips coffee… ahh yes, let me find that classic Dropbox rsync comment
Tools like https://github.com/badlogic/pi-mono implement most of the functionality Claude Code has, even adding loads of stuff Claude doesn't have and can actually scroll without flickering inside terminal, all built by a single guy as a side project. I guess we can't ask that much from a 250B USD company.
Be careful with the coffee.
I asked > Can you find an academic article that _looks_ legitimate -- looks like a real journal, by researchers with what look like real academic affiliations, has been cited hundreds or thousands of times -- but is obviously nonsense, e.g. has glaring typos in the abstract, is clearly garbled or nonsensical? It pointed me to a bunch of hoaxes. I clarified: > no, I'm not looking for a hoax, or a deliberate comment on…
Well, if there are papers that match your criteria, it's hallucinating the "no".
Earlier quoted context omitted.
Can you be more specific than this? does it vary in time from launch of a model to the next few months, beyond tinkering and optimization?
Yeah, happy to be more specific. No intention of making any technically true but misleading statements. The following are true: - In our API, we don't change model weights or model behavior over time (e.g., by time of day, or weeks/months after release) - Tiny caveats include: there is a bit of non-determinism in batched non-associative math that can vary by batch / hardware, bugs or API downtime can obviously change…
I feel like you need to be making a bigger statement about this. If you go onto various parts of the Net (Reddit, the bird site etc) half the posts about AI are seemingly conspiracy theories that AI companies are watering down their products after release week.
Earlier quoted context omitted.
Paying $10 per request doesn't have me jumping at the opportunity to try it!
The only way to not go bankrupt is to use a Claude Code Max subscription…
Earlier quoted context omitted.
Even for coding, it seems to still make A LOT of mistakes. https://youtu.be/8brENzmq1pE?t=1544 I feel like everyone is counting chickens before they hatch here with all the doomsday predictions and extrapolating LLM capability into infinity. People that seem to overhype this seem to either be non-technical or are just making landing pages.
Waiting until the moment they get good enough is not a smart thing to do either. If you are a farmer and know it is going to snow, at some point in the next 5 months, you make plans NOW, you don't wait until the temperatures drop and you see the snow falling. Right now, people are waiting for the snowfall before moving their proverbial chickens indoors
It's hard to tell with these releases if Anthropic's astroturfing campaign has come to HN or not but I feel like it probably has
the top 5 comments on this thread are from accounts that are around 10 years old each. What gives you any reason to believe this is an astroturfing campaign?
I asked > Can you find an academic article that _looks_ legitimate -- looks like a real journal, by researchers with what look like real academic affiliations, has been cited hundreds or thousands of times -- but is obviously nonsense, e.g. has glaring typos in the abstract, is clearly garbled or nonsensical? It pointed me to a bunch of hoaxes. I clarified: > no, I'm not looking for a hoax, or a deliberate comment on…
Just tested the new Opus 4.6 (1M context) on a fun needle-in-a-haystack challenge: finding every spell in all Harry Potter books. All 7 books come to ~1.75M tokens, so they don't quite fit yet. (At this rate of progress, mid-April should do it ) For now you can fit the first 4 books (~733K tokens). Results: Opus 4.6 found 49 out of 50 officially documented spells across those 4 books. The only miss was "Slugulus Eruc…
It feels like a very odd test because it's such an unreasonable way to answer this with an LLM. Nothing about the task requires more than a very localized understanding. It's not like a codebase or corporate documentation, where there's a lot of interconectedness and context that's important. It also doesn't seem to poke at the gap between human and AI intelligence.
Why are people excited? What am I missing?
Earlier quoted context omitted.
Well, if there are papers that match your criteria, it's hallucinating the "no".
And there are: https://en.wikipedia.org/wiki/Sokal_affair
The Sokal paper was a hoax so it doesn’t meet the criteria.