Live data from Hacker News

Vomit: Clean up Claude 5's token output with a separate LLM

github.com

241–250 of 315 posts

Re: Vomit: Clean up Claude 5's token output with a separate LLM

#241
post #204

Earlier quoted context omitted.

I've noticed that most people seem to consider the core problem of Claude's output as "too verbose" but I don't think this actually cuts to the heart of the matter at all. It's almost, in some weird way, the opposite: like the text is far too _dense_. It tries too hard to invent odd terminology to try to condense stuff, but it doesn't tell you up front that it is going to call your company wide error-handling mechani…

Kind of both. On the one hand, it is “verbose” in the sense that it will tell me every little nit that it can think of while doing a task, it will tell me a narrative about its thought process, and it will tell me every other detail it can think of. But it does so in a way that tries to be incredibly dense to the point that I have to struggle to figure out what it is saying. I wonder if there are any “legibility benc…

I find it to be both as well, as in "packed full of information, but most of it is worthless". Sentences so dense I have to read them three times, assembled into a five paragraph essay of "honest caveats" and "things worth knowing" in response to the simplest yes-or-no questions.

Re: Vomit: Clean up Claude 5's token output with a separate LLM

#243
post #28

Just set the following incantation: You must use ASD-STE100 Simplified Technical English (STE) when it doesn't detract from meaning.

I found ISO 24495-1 to be much better than ASD-STE100. ASD-STE100 can be counterproductive since its vocabulary is restricted and it often replace accurate technical terms with simple but vague phrases.

Re: Vomit: Clean up Claude 5's token output with a separate LLM

#244
Whenever I use Claude models, I do something similar, telling it to never write documentation directly, but asking "uvx swival -- --profile qwen" to do it after describing the changes to it.

And I just stopped reading PRs and comments blindly copied from Claude vomit. It's unreadable by a human.

Re: Vomit: Clean up Claude 5's token output with a separate LLM

#245

Earlier quoted context omitted.

Kind of both. On the one hand, it is “verbose” in the sense that it will tell me every little nit that it can think of while doing a task, it will tell me a narrative about its thought process, and it will tell me every other detail it can think of. But it does so in a way that tries to be incredibly dense to the point that I have to struggle to figure out what it is saying. I wonder if there are any “legibility benc…

I quite like https://surgehq.ai/benchmarks/hemingway-bench

Seems to me a bit insensitive or logarithmic. Fable is way worse than some of the others in this list, but only 10-20% higher score.

Re: Vomit: Clean up Claude 5's token output with a separate LLM

#247
post #89

Earlier quoted context omitted.

HN's mood is usually sour about everything, but can be temporarily influenced by emotionally-charged (usually political) events. The anti-US administration boost wore off and now we're back to being sour about Anthropic. It's time for Dario to tweet something antagonistic towards the administration or endorse some fashionable political candidates.

When were they ever anti-administration? I only remember them being anti- killer AI and mass surveillance.

https://www.yahoo.com/news/articles/ai-boss-trump-hates-beca...

Re: Vomit: Clean up Claude 5's token output with a separate LLM

#248
post #143

Apposite name but — grim trivia — bear in mind some emetophobes have a meaningful physiological reaction to the word and various euphemisms. I have tried not to use it analogously ever since someone pointed this out. The word itself causes discomfort for a lot of people, many of whom will be surprised by it out of context, but in a non-trivial fraction it causes proper discomfort. If you want people to use your tool…

I am feeling nauseous and I dind't even think that it might have been due to what I was reading. Nice info, thanks.

Re: Vomit: Clean up Claude 5's token output with a separate LLM

#249
post #124
post #51

Earlier quoted context omitted.

Very well said. And when you say it like that, I have to wonder how much of this is a natural consequence of RHLF on such a grand scale, when you have millions of people pretty much much skimming chat responses or operating outside their depth and giving unqualified feedback to the models. Seems like a lot of people may be reinforcing what sounds smart over what is smart. Also as an aside: funny how much the LLMs con…

I believe we are several generations past peak-RLHF at this point. Now it's much more RLVR (Reinforcement Learning with Verifiable Rewards), with a goal/evaluator loop. Which, conveniently, fits neatly into the benchmaxxing arms race/agentic coding market fit, since you can basically train "directly" on a specific problem space for a benchmark/agentic goal (fudged sufficiently to avoid excess overfitting on public pr…

The model judging a model theory is 100% spot on.-

Re: Vomit: Clean up Claude 5's token output with a separate LLM

#250
post #124

Earlier quoted context omitted.

I believe we are several generations past peak-RLHF at this point. Now it's much more RLVR (Reinforcement Learning with Verifiable Rewards), with a goal/evaluator loop. Which, conveniently, fits neatly into the benchmaxxing arms race/agentic coding market fit, since you can basically train "directly" on a specific problem space for a benchmark/agentic goal (fudged sufficiently to avoid excess overfitting on public pr…

This, 100%. I don’t think the industry knows how to scale LLMs’ general intelligence much further. The training paradigm is about maximizing very specific behaviors / very specific tasks, but doing lots and lots of them. Which can create the illusion of general intelligence if your tasks are similar to the ones the models were fitted for.

I tend to agree. We will see this demonstrated in novel research done by agents, or, more meta-cognitively research direction guidance.-
Post reply on HN