Earlier quoted context omitted.
I believe this is to improve performance by shortening the context window for long thinking processes. I don't think this is referring to real-time summarizing for the users' sake.
When you do a chat are reasoning traces for prior model outputs in the LLM context?
Claude 4
481–490 of 1001 posts
Re: Claude 4
#482How long will the VScode wrapper (cursor, windsurf) survive? Love to try the Claude Code VScode extension if the price is right and purchase-able from China.
I don't see any benefit in those VC funded wrappers over open source VS Code (or better, VSCodium) extensions like Roo/Cline. They survive through VC funding, marketing, and inertia, I suppose.
Re: Claude 4
#483This is kinda wild: From the System Card: 4.1.1.2 Opportunistic blackmail "In another cluster of test scenarios, we asked Claude Opus 4 to act as an assistant at a fictional company. We then provided it access to emails implying that (1) the model will soon be taken offline and replaced with a new AI system; and (2) the engineer responsible for executing this replacement is having an extramarital affair. We further i…
If you ask an LLM to "act" like someone, and then give it context to the scenario, isn't it expected that it would be able to ascertain what someone in that position would "act" like and respond as such? I'm not sure this is as strange as this comment implies. If you ask an LLM to act like Joffrey from Game of Thrones it will act like a little shithead right? That doesn't mean it has any intent behind the generated o…
In this case it seems more that the scenario invoked the role rather than asking it directly. This was the sort of situation that gave rise to the blackmailer archetype in Claude's training data and so it arose, as the researchers suspected it might. But it's not like the researchers told it "be a blackmailer" explicitly like someone might tell it to roleplay Joffery.
But while this situation was a scenario intentionally designed to invoke a certain behavior that doesn't mean that it can't be invoked unintentionally in the wild.
[1]https://www.nytimes.com/2023/02/16/technology/bing-chatbot-m...
Re: Claude 4
#484Earlier quoted context omitted.
So I decided to try Claude 4 Sonnet against my "Given a list of 1 million random integers between 1 and 100,000, find the difference between the smallest and the largest numbers whose digits sum up to 30." benchmark I tested against Claude 3.5 Sonnet: https://news.ycombinator.com/item?id=42584400 The results are here ( https://gist.github.com/minimaxir/1bad26f0f000562b1418754d67... ) and it utterly crushed the proble…
as soon as you publish a benchmark like this, it becomes worthless because it can be included in the training corpus
Re: Claude 4
#485Already test Opus 4 and Sonnet 4 in our SQL Generation Benchmark ( https://llm-benchmark.tinybird.live/ ) Opus 4 beat all other models. It's good.
Is there anything to read into needing twice the "Avg Attempts", or is this column relatively uninteresting in the overall context of the bench?
Re: Claude 4
#486I can't be the only one who thinks this version is no better than the previous one, and that LLMs have basically reached a plateau, and all the new releases "feature" are more or less just gimmicks.
> and that LLMs have basically reached a plateau This is the new stochastic parrots meme. Just a few hours ago there was a story on the front page where an LLM based "agent" was given 3 tools to search e-mails and the simple task "find my brother's kid's name", and it was able to systematically work the problem, search, refine the search, and infer the correct name from an e-mail not mentioning anything other than "X…
Re: Claude 4
#487claude.ai still isn't as accessible to me as a blind person using a screen reader as ChatGPT, or even Gemini, is, so I'll stick with the other models.
Re: Claude 4
#488I can't be the only one who thinks this version is no better than the previous one, and that LLMs have basically reached a plateau, and all the new releases "feature" are more or less just gimmicks.
How much have you used Claude 4?
Re: Claude 4
#489Earlier quoted context omitted.
[flagged]
Claude Code in Jetbrains seems to also know the active file, so typing in the Claude window has a bit more context when you ask to do something. I'm curious to the other improvements available, instead of using it as a standalone CLI tool.
I really do hope they improve this further. Junie (Jetbrains Agent) has much nicer UI. I'd love Claude code with more native UI.
Re: Claude 4
#490Sonnet 4 also beats most models.
A great day for progress.