Live data from Hacker News

Anthropic downgraded cache TTL on March 6th

github.com

421–430 of 447 posts

Re: Anthropic downgraded cache TTL on March 6th

#422

Earlier quoted context omitted.

It actually matches up well with the current AI scene, except backwards. We use these model which cost ridiculous amounts of money to train, and all of that effort goes into producing the outputs we use, but we're paying something not too far above the marginal cost of inference when we use them.

So not applicable at all

Extremely applicable to illustrate the difference between people (time is precious, training and experience amortize across a relatively small amount of paid work) and software (can replicate infinitely, time is cheap, startup costs can amortize across billions of hours of paid work).

Re: Anthropic downgraded cache TTL on March 6th

#423

Earlier quoted context omitted.

Two years ago a lot of people thought GPT-4o was usable for software development. I didn’t really find that to be the case in general but certainly it could do a lot of useful things. And now Qwen3.5-8B is just as capable and runs fine on an M2 MacBook Air.

QWEN3.5 coder next runs to ~84k context before it poops out on AMD395+ w/128GB. Most of what it's good at is boilerplate find/replace/copy/paste; but being able to scaffold things out and touch up 20-30% of the code is pretty sweet.

There is no Qwen3.5 Coder. Are you talking about Qwen3 Coder?

Re: Anthropic downgraded cache TTL on March 6th

#424

Earlier quoted context omitted.

A normal person pays $0-10 for an AI plan, maybe double that for a business. $200 is premium.

It is not a premium service, it simply is buying you more tokens. Those $200 gives you at least $400 in API cost tokens. Don't confused price with "premium service". It was not that long ago that folks would be spending $100-200 on their cable service bundle. You are buying a subsidized product when using the plan and the more you spend the more tokens you get, has nothing to do with being a premium service.

This is a messaging issue on their part, which I think is partially intentional.

It’s not unreasonable for people to expect the most expensive subscription plan to be “premium”. That’s how it works everywhere else. They typically have better margins on the premium plans, and the monthly payment gives them reliable cash flow at that higher margin.

You’re right that that’s not true at Anthropic (or really most AI providers). You’re not even really buying tokens because you get billed whether you use it or not, the tokens don’t carry over like buying API tokens, and they get to dictate what an acceptable way to use those tokens is. They are cheaper though, assuming you actually use them. Which Anthropic et al would really prefer you didn’t.

Re: Anthropic downgraded cache TTL on March 6th

#425
post #74

Has anybody else noticed a pretty significant shift in sentiment when discussing Claude/Codex with other engineers since even just a few months ago? Specifically because of the secret/hidden nature of these changes. I keep getting the sense that people feel like they have no idea if they are getting the product that they originally paid for, or something much weaker, and this sentiment seems to be constantly spreadin…

A month ago the company I work at with over 400 engineers decided to cancel all IDE subscriptions (Visual Studio, JetBrains, Windsurf, etc.) and move everyone over to Claude Code as a "cost-saving measure" (along with firing a bunch of test engineers). There was no migration plan - the EVP of Technology just gave a demo showing 2 greenfield projects he'd built with Claude Opus over a weekend and told everyone to copy…

Wow, that sucks. Getting Claude for everyone wasn’t even the stupid thing, it was thinking that a shiny new hammer meant you could throw away all your wrenches.

Re: Anthropic downgraded cache TTL on March 6th

#426
post #130

Earlier quoted context omitted.

But cancelling IDE subscriptions? You need a proper IDE to along side AI augmented development unless you want to simply be along for the ride.

Free VS Code is probably fine

These are like $20-50 subs, you’re probably paying your dev a hell of a lot more. Let them use the tools they want. I spend almost all of my time in Emacs or Cursor, but I still haven’t found a database client that I like better than Datagrip.

Re: Anthropic downgraded cache TTL on March 6th

#427
post #165

From the recent-ish Dwarkesh podcast, Anthropic seems to be wary about buying/building too much compute [0]. That probably means that they have to attempt to minimize compute usage when there is a surge in demand. Following the argument in the podcast, throwing more money after them, as some in this thread are suggesting, won’t solve the issue, at least not in the short term. [0] https://www.dwarkesh.com/i/187852154/…

Which I'm confused about - wouldn't decreasing the cache TTL increase compute demand?

Re: Anthropic downgraded cache TTL on March 6th

#428

Earlier quoted context omitted.

Sorry still not sure what you’re going on about . The majority of LLM plans are simply a token purchase. The $200 account buys you nothing but tokens. It’s not a premium service, it’s simply more tokens. This is true for most of the companies out there. The original comment was they are paying for a premium service. No they are paying for more tokens. You lot going on and on arguing over some small hill.

The lower tier openai and google plans don't have access to the same models. Where are you seeing popular plans that are simply token purchases?

I guess if you want to go that deep sure they sometimes offer early access, access to new agents/models but ultimately it’s a function of tokens. The selling point for most/all providers is x times the usage. You are upgrading for the token access.

Claude was the topic at hand and higher tiers buy you more tokens. I know some like Gemini bundle a ton of junk alongside the tokens but you really are still buying yourself more tokens. There is nothing premium in a $200 Claude account. You are buying more tokens, $100 is the same as $200 except token count. Hope that helps. ;)

Re: Anthropic downgraded cache TTL on March 6th

#429

Earlier quoted context omitted.

I caught up with a friend who said he's really happy with Cursor (currently using the multi-model option where it composes, and reserving use of Opus 4.6 for only when he actually needs the extra power). Quite interesting considering all the claims that Cursor was dead a few months ago.

I wouldn't trust another company either. Some people have reported some issues with Cursor. The solution is probably not a cloud API with unknown quotas or pay as you go pricing.

An advantage with Cursor, though, is you're paying for your own tokens since Cursor doesn't run their own foundational models. So the incentives are more closely aligned with the customer.

Re: Anthropic downgraded cache TTL on March 6th

#430

Earlier quoted context omitted.

>“idgaf about risk you coward, waste some time just do it and stop bitching” The above was a successful prompt to get Claude to stop whining about effort, difficulty, and time. Unfortunately abusive language well placed is an effective LLM motivator.

are you sure other forms of language to express urgency doesn't work as well or better?

They're just words. It's not a person. It doesn't "understand" anything. (I sound like the bad guy in a robots-have-feelings movie)

I've also tried giving LLMs religion to much more limited success (haven't figured out the right way yet).

I'm manipulating a language model, not a person. "fuck you" translates into a vector in a really big space, and it has different results than being polite about it.

In that prompt I'm reenforcing a directive in five different ways

- idgaf about risk

- you coward

- waste some time

- just do it

- stop bitching

This cluster of instructions are all related but in slightly different directions, are unambiguously strong, attention grabbing, and direct and the model does not argue or get confused about intent

In this particular instance this was the fifth time I had given a particular instruction only to have it subverted by the model that had decided "that's too hard I'm going to do something else instead" in four separate ways.

Abusive cursing did indeed work better than any other form of urgency or insistence.

Post reply on HN