Live data from Hacker News

The beginning of scarcity in AI

tomtunguz.com

221–230 of 239 posts

Re: The beginning of scarcity in AI

#221

This isn't the first time they've dealt with scarcity, there's been supply chain scarcity four times since 2000. Post-dotcom boom, CDMA scarcity, HDD/flash scarcity, Pandemic scarcity. The scarcity isn't long-term. Like all manufactured products, they'll ramp up production and flood the market with hardware, people will buy too much, market will drop. Boom and bust. We're also still in the bubble. Eventually markets…

> In 5 years consumer chips and model inference will be so good you won't need a server for SOTA. Naw man, you crazy. If you tell me that in 5 years, consumer chips will be so good that I can run GPT-5.4-level AI on my phone, I'd find that plausible (I buy cheap phones). If you're telling me that in 5 years we won't need _servers_ because our _phones and/or desktops_ will be powerful enough to run the biggest newest…

The thing is SOTA has a plateau. All LLMs work on the same principle: input goes in for training, reinforced by humans. There is only so much input (all recorded human knowledge), only so many human tweaks, that can produce only so much increased signal-to-noise in output. The machine can't read your mind, and there is no one truthful answer to most questions, so there will always be a limit on how accurate or correct or whatever any response will get. So at some point, you just can't make a better response. The agent harness, prompts, etc, are the only way to get better, and that's gonna be open source.

Add to that the algorithmic improvements on inference that's making inference faster with more context and higher quality. TurboQuant is just one example, more methods are coming out all the time. So the inference is getting more efficient.

At the same time, hardware can kind of keep getting infinitely better. Even if you can't make it smaller, you can make it more energy efficient, improve multitasking, more GPU cores/RAM or iGPUs, pack in more chips, improve cooling, use new materials... the sky's the limit.

Add all 3 together and at some point you will get Opus 4.7 on a phone with 40 t/s. At that point there's no way I'm paying for inference on a server. You can do RAG on-device, and image/video/voice is done by multi-modals. I want my agent chats replicated, but that's Google Drive. I want the agent to search the web, but that's Google Search. So eventually we're back to just doing what we do today (pre-AI) only with more automation.

The really advanced shit will come in 10 years, when we finally crack real memory and learning. That will absolutely be locked up in the cloud. But that's not an LLM, it's something else entirely. (slight caveat that WW3 will delay progress by 10-20 years)

Re: The beginning of scarcity in AI

#223
post #128

Earlier quoted context omitted.

Yup. Also regardless of price they need to spend more and more as the project collapses under the inevitable incidental complexity of 30k lines of code a day. It's similar to how if you know what you're doing you can manage a simple VPS and scale a lot more cost effectively than something like vercel. In a saturated market margins are everything. You can't necessarily afford to be giving all your margins to anthropic…

I also can’t wait for the time when few know how to code. Just like how many folks don’t know html from css when the homebrew website went away. Their might always be llms, but the dependence is an interesting topic.

A time when few know how to code?

I think that was about 10 years ago…

Re: The beginning of scarcity in AI

#224
post #26
post #14

Constraints can lead to innovation. Just two things that I think will get dramatically better now that companies have incentive to focus on them: * harness design * small models (both local and not) I think there is tremendous low hanging fruit in both areas still.

China already operates like this. Low cost specialized models are the name of the game. Cheaper to train, easy to deploy. The US has a problem of too much money leading to wasteful spending. If we go back to the 80s/90s, remember OS/2 vs Windows. OS/2 had more resources, more money behind it, more developers, and they built a bigger system that took more resources to run. Mac vs Lisa. Mac team had constraints, Lisa t…

It has been a very bad bet that hardware will not evolve to exceed the performance requirements of today's software tomorrow, just as it is a bad bet that tomorrow someone will rewrite today's software to be slower.

Re: The beginning of scarcity in AI

#225
post #26

Earlier quoted context omitted.

China already operates like this. Low cost specialized models are the name of the game. Cheaper to train, easy to deploy. The US has a problem of too much money leading to wasteful spending. If we go back to the 80s/90s, remember OS/2 vs Windows. OS/2 had more resources, more money behind it, more developers, and they built a bigger system that took more resources to run. Mac vs Lisa. Mac team had constraints, Lisa t…

It has been a very bad bet that hardware will not evolve to exceed the performance requirements of today's software tomorrow, just as it is a bad bet that tomorrow someone will rewrite today's software to be slower.

Eh, but then as hardware evolves, the software will also follow suit. We’ve had an explosion of compute performance and yet software is crawling for the same tasks we did a decade ago.

Better hardware ensures that software that is “finished” today will run at acceptable levels of performance in the future, and nothing more.

I think we won’t see software performance improve until real constraints are put on the teams writing it and leaders who prioritize performance as a North Star for their product roadmap. Good luck selling that to VCs though.

Re: The beginning of scarcity in AI

#226
post #209
post #164

Earlier quoted context omitted.

The decline of independent thoughts for one. As people become reliant on LLMs to do their thinking for them and solve all problems that they stumble upon, they become a shell of their previous self. Sadly, this is already happening.

There is no decline. Human assets were always too expensive to process some additional information. We are simply processing lot more of low signal data. Actually some of our analysts are empowered by the tools at their disposal. Their jobs are safe and necessary. Others were let go. Clients are happy to get fuller picture of their universe, which drives more informed decissions . Everybody wins.

Are you being satirical?

Re: The beginning of scarcity in AI

#227

Earlier quoted context omitted.

When you go to the command line and type “Claude”, there is an LLM, and everything else is the harness

I'm having an hard time getting my mind to see this. > Users should re-tune their prompts and harnesses accordingly. I read this in the press release and my mind thought it meant test harness. Then there was a blog post about long running harnesses with a section about testing which lead me to a little more confusion. Yes, the word 'harness' is consistently used in the context as a wrapper around the LLM model not as…

Some people also call evaluations "tests". There are unexpected things that come along with new models, like the model in a workflow you'd set up suddenly starts calling a tool and never stops or decides to no longer call a particular tool, so running your existing evaluations to catch regressions like this and potentially updating the prompts is considered "testing" your prompts and harnesses.

Re: The beginning of scarcity in AI

#228
post #14

Constraints can lead to innovation. Just two things that I think will get dramatically better now that companies have incentive to focus on them: * harness design * small models (both local and not) I think there is tremendous low hanging fruit in both areas still.

[dead]

Re: The beginning of scarcity in AI

#229

Earlier quoted context omitted.

You misunderstand. "I built a ship to go to the Indies and bring back tea." "Bro, the ship cost 100,000 pounds sterling and only brought back 50,000 pounds of tea. I don't care if you paid 12,500 pounds for the tea itself, you're losing money." There is a very rational reason labs are spending everything they can get for more compute right now. The tea (inference) pays 60%+ margins. And that is rising. And that numbe…

60%+ margins according to numbers which are not published publicly and have not AFAICT been audited. Could they be accurate? Sure, I think people who claim this is impossible are overconfident. But I would encourage anyone who assumes they must be right to read a history of the Worldcom scandal. It's really quite easy for a person who wants to be making money (or an LLM who's been instructed to "run the accounts make…

Any materially false public statement by one of the foundation lab CEOs is a huge foot fault. I'm not saying they would never lie, but it would be a very, very dumb thing to do. That public information can be relied on by their private (very powerful) investors. I think if you're hearing these numbers ballparked in public settings, they are, as a prior, directionally accurate.

Re: The beginning of scarcity in AI

#230

Earlier quoted context omitted.

60%+ margins according to numbers which are not published publicly and have not AFAICT been audited. Could they be accurate? Sure, I think people who claim this is impossible are overconfident. But I would encourage anyone who assumes they must be right to read a history of the Worldcom scandal. It's really quite easy for a person who wants to be making money (or an LLM who's been instructed to "run the accounts make…

Any materially false public statement by one of the foundation lab CEOs is a huge foot fault. I'm not saying they would never lie, but it would be a very, very dumb thing to do. That public information can be relied on by their private (very powerful) investors. I think if you're hearing these numbers ballparked in public settings, they are, as a prior, directionally accurate.

I agree, although I would emphasize that Worldcom is a great example of a CEO doing that very dumb thing. But I am not hearing these numbers ballparked in public settings. As far as I can tell, all the numbers people discuss for OpenAI or Anthropic margins come from anonymous leaks of internal documents.
Post reply on HN