Live data from Hacker News

Uber's $1,500/month AI limit is a useful signal for AI tool pricing

simonwillison.net

641–650 of 819 posts

Re: Uber's $1,500/month AI limit is a useful signal for AI tool pricing

#641

Plenty of comparisons here between salaries and token costs. All fair but very much assumes that salaries are rational. Why do we pay some engineers 10x as much for the same role just because they are in a different location? The WFH discussion surfaced some of that. If money is cheap, all sorts of funny things are happening. Is it worth to spend 1500 USD on AI? I don’t know. Is it worth paying engineers 300k USD ins…

Why even pay them at all? Just lock them in a cell and give them a bowl of rice.

Re: Uber's $1,500/month AI limit is a useful signal for AI tool pricing

#642

Earlier quoted context omitted.

One aspect Paul Kedrosky mentioned recently is the concept of „duration mismatch“. The price per token goes down over time (either because the AI vendor reduces due to competition pressure, or because customers are now incentivized to use older cheaper models). But datacenters are financed through debt, with the assumption their revenue increases over time. Quoting him: „[AI vendors are] paying for a fixed cost with…

If you have a good model router, you can route to older, cheaper models that run on older hardware, for simpler tasks. That helps labs extend the economic life of their hardware investments. They will likely fight it at first though as they see it as reducing ASP. This is why I'm building role-model, a routing protocol and a router runtime: https://role-model.dev/

Running cheaper models on newer hardware is always going to beat running them on older hardware.

Re: Uber's $1,500/month AI limit is a useful signal for AI tool pricing

#643

Earlier quoted context omitted.

Using a shittier model is just more work for the user, I’m not sure why anyone does it, unless they’re playing with it like a toy.

Local privacy respecting inference can be worth it. I use a local model to log everything I do all week to automate my timesheet. I also have it do a bunch of other data tasks. I won't say that larger SOTA models wouldn't do these tasks better than a local model but PII is a concern and my employer wouldn't approve of me just setting tokens on fire everyday to do what I could do myself.

> I use a local model to log everything I do all week to automate my timesheet.

Isn’t that just more work than logging it yourself?

Re: Uber's $1,500/month AI limit is a useful signal for AI tool pricing

#644
post #472

Earlier quoted context omitted.

Have you tried adding this information to claude.md so it knows? I also think your excuse is bad. "The code is legacy fucked so I'll just legacy fuck it some more because I can't be bothered to make an effort"

This is a spicy take, unless the business is willing to face some down time, and I am hired to do exactly what you said, I’d never touch any line of code unless I absolutely have to. Different environments don’t help as much. We tend to obsess over software quality when it’s the least important thing for a business. It’s just a means to an end.

> Least important thing for a business

- Takes weeks or months to get simple features out the door, and when they're out they're buggy as hell and the bugs never get fixed. Sound familiar?

> I’d never touch any line of code unless I absolutely have to

And this is how legacy code is made. Years of everyone "never touching anything they don't have to" leads to a giant steaming pile of shit.

> unless the business is willing to face some down time

How does a simple refactor cause downtime? I do this kind of stuff all the time and pretty much never cause any downtime. In the very rare cases that prod downtime does occur it's generally not because of some simple code refactor, and we have it back up in no time by just rolling it back. Unless it's not related to the code at all, in which case it also wasn't a refactor that caused it.

Re: Uber's $1,500/month AI limit is a useful signal for AI tool pricing

#645
How are people using so many tokens? I'm on the $200/month enterprise plan for Claude Code (because it's a better deal than the API pricing) and I don't come close to the limits.

If you use stuff like opusplan and /advisor so you use Sonnet for most of the work and only Opus for the really complex stuff then it's quite easy to keep costs low without affecting performance.

Re: Uber's $1,500/month AI limit is a useful signal for AI tool pricing

#646

How are people using so many tokens? I'm on the $200/month enterprise plan for Claude Code (because it's a better deal than the API pricing) and I don't come close to the limits. If you use stuff like opusplan and /advisor so you use Sonnet for most of the work and only Opus for the really complex stuff then it's quite easy to keep costs low without affecting performance.

All new/renewing enterprise contracts with Claude Enterprise and ChatGPT Enterprise no longer offer usage-based subscriptions, but instead will charge API pricing for all tokens consumed, and as you've said, the subs are better deals than raw API pricing.

Re: Uber's $1,500/month AI limit is a useful signal for AI tool pricing

#647

$1500/mo is $18,000/seat/annum. Maybe Microsoft and Nvidia are on to something. 128 GB machines that can run local LLMs are a bargain even if priced $5-8k. Yes, tok/s is not quite there, but that's probably OK since the bottleneck really isn't the code; it's WTF did Uber build with all of that spend? How did it meaningfully impact their revenue in a positive direction?

Your last question is really important. What did they accomplish with all that spend? I suspect there’s some mass delusion with respect to actual accomplishments as a result of LLM use. Sure, things are moving faster, but does it matter?

Never confuse movement with action.

Re: Uber's $1,500/month AI limit is a useful signal for AI tool pricing

#648

$1500/mo is $18,000/seat/annum. Maybe Microsoft and Nvidia are on to something. 128 GB machines that can run local LLMs are a bargain even if priced $5-8k. Yes, tok/s is not quite there, but that's probably OK since the bottleneck really isn't the code; it's WTF did Uber build with all of that spend? How did it meaningfully impact their revenue in a positive direction?

>WTF did Uber build with all of that spend? How did it meaningfully impact their revenue in a positive direction? Uber (and quite a few bay area companies and startups) can afford to spend that money. There is no expectation of profit, Uber lost ~62B and growing: https://uberlosses.com/

As much as I love to hate on Uber, that website is from 2022. Uber has been profitable since 2023.

It's profit margin seems to have stabilized around 10%.

The real economic crime is losing at least $40bn over 10 years scaling a business that ended up having retail profit margins (i.e. low profit margins).

Re: Uber's $1,500/month AI limit is a useful signal for AI tool pricing

#649

Earlier quoted context omitted.

Cro.ai seems to be: https://crof.ai/ Of course though they are not necessarily a viable solution for companies with security requirements etc. given it is just a single person project, but they still serve as a proof it can be done.

This costs more.

Not as far as I can tell. Are we seeing different things?

For deepseek-v4-pro:

- $0.350 in, $0.003000 cache, $0.80 out https://crof.ai/pricing

- $0.435 in, $0.003625 cache, $0.87 out https://api-docs.deepseek.com/quick_start/pricing

Re: Uber's $1,500/month AI limit is a useful signal for AI tool pricing

#650
post #274

Earlier quoted context omitted.

It's called AWS. Bedrock is right there. Price or data policy is never the issue. The models themselves are the problem -- most large US companies are not going to touch them. Source: directly involved in these discussions. You can downvote as much as you'd like but you can't ignore the facts.

> The models themselves are the problem -- most large US companies are not going to touch them. Can you expand on this?

Some suits with no understanding of how LLMs work are scared that the models might hack them, or believe that they'd have to send data to China because they do not know that open models can be run on your own infra.
Post reply on HN