Earlier quoted context omitted.
I was playing about with Chat GPT the other day, uploading screen shots of sheet music and asking it to convert it to ABC notation so I could make a midi file of it. The results seemed impressive until I noticed some of the "Thinking" statements in the UI. One made it apparent the model / agent / whatever had read the title from the screenshot and was off searching for existing ABC transcripts of the piece Ode to Joy…
Yes I have found that grok for example actually suddenly becomes quite sane when you tell it to stop querying the internet And just rethink the conversation data and answer the question. It's weird, it's like many agents are now in a phase of constantly getting more information and never just thinking with what they've got.
Claude Opus 4.6
811–820 of 1001 posts
Re: Claude Opus 4.6
#812Say I am just an average coder doing a days work with Claude. How much will that cost?
Re: Claude Opus 4.6
#813Re: Claude Opus 4.6
#814Earlier quoted context omitted.
I have not see any reporting or evidence at all that Anthropic or OpenAI is able to make money on inference yet. > Turns out there was a lot of low-hanging fruit in terms of inference optimization that hadn't been plucked yet. That does not mean the frontier labs are pricing their APIs to cover their costs yet. It can both be true that it has gotten cheaper for them to provide inference and that they still are subsid…
> I have not see any reporting or evidence at all that Anthropic or OpenAI is able to make money on inference yet. Anthropic planning an IPO this year is a broad meta-indicator that internally they believe they'll be able to reach break-even sometime next year on delivering a competitive model. Of course, their belief could turn out to be wrong but it doesn't make much sense to do an IPO if you don't think you're clo…
VC firms, even ones the size of Softbank, also literally just don't have enough capital to fund the planned next-generation gigawatt-scale data centers.
Re: Claude Opus 4.6
#815Earlier quoted context omitted.
> This is obviously not true, you can use real data and common sense. It isn't "common sense" at all. You're comparing several companies losing money, to one another, and suggesting that they're obviously making money because one is under-cutting another more aggressively. LLM/AI ventures are all currently under-water with massive VC or similar money flowing in, they also all need training data from users, so it is v…
Doing some math in my head, buying the GPUs at retail price, it would take probably around half a year to make the money back, probably more depending how expensive electricity is in the area you're serving from. So I don't know where this "losing money" rhetoric is coming from. It's probably harder to source the actual GPUs than making money off them.
https://www.dbresearch.com/PROD/RI-PROD/PROD0000000000611818...
Re: Claude Opus 4.6
#816I feel like I can't even try this on the Pro plan because Anthropic has conditioned me to understand that even chatting lightly with the Opus model blows up usage and locks me out. So if I would normally use Sonnet 4.5 for a day's worth of work but I wake up and ask Opus a couple of questions, I might as well just forget about doing anything with Claude for the rest of the day lol. But so far I haven't had this issue…
Re: Claude Opus 4.6
#8175.2 (and presumably 5.3) is really smart though and feels like it has higher "raw" intelligence.
Opus feels like a better model to talk to, and does a much better job at non-coding tasks especially in the Claude Desktop app.
Here's an example prompt where Opus in Claude put in a lot more effort and did a better job than GPT5.2 Thinking in ChatGPT:
`find all the pure software / saas stocks on the nyse/nasdaq with at least $10B of market cap. and give me a breakdown of their performance over the last 2 years, 1 year and 6 months. Also find their TTM and forward PE`
Opus usage limits are a bummer though and I am conditioned to reach for Codex/ChatGPT for most trivial stuff.
Works out in Anthropic's favor, as long as I'm subscribed to them.
Re: Claude Opus 4.6
#818Just tested the new Opus 4.6 (1M context) on a fun needle-in-a-haystack challenge: finding every spell in all Harry Potter books. All 7 books come to ~1.75M tokens, so they don't quite fit yet. (At this rate of progress, mid-April should do it ) For now you can fit the first 4 books (~733K tokens). Results: Opus 4.6 found 49 out of 50 officially documented spells across those 4 books. The only miss was "Slugulus Eruc…
https://www.wizardemporium.com/blog/complete-list-of-harry-p...
Why is this impressive?
Do you think it's actually ingesting the books and only using those as a reference? Is that how LLMs work at all? It seems more likely it's predicting these spell names from all the other references it has found on the internet, including lists of spells.
Re: Claude Opus 4.6
#819Earlier quoted context omitted.
Can you be more specific than this? does it vary in time from launch of a model to the next few months, beyond tinkering and optimization?
Yeah, happy to be more specific. No intention of making any technically true but misleading statements. The following are true: - In our API, we don't change model weights or model behavior over time (e.g., by time of day, or weeks/months after release) - Tiny caveats include: there is a bit of non-determinism in batched non-associative math that can vary by batch / hardware, bugs or API downtime can obviously change…
Maybe a dumb question but does this mean model quality may vary based on which hardware your request gets routed to?
Re: Claude Opus 4.6
#820Earlier quoted context omitted.
Your game is amazing! I wish there was a "Reset" button to go back to the original position. Where are you in Poland?
Thanks :) Click "Level" -> "Try again" Originally from Wrocław, but don't live in Poland anymore
BUT, I meant a button to restart after a few moves. Anyways, cool!