Live data from Hacker News

The sigmoids won't save you

astralcodexten.com

191–200 of 297 posts

Re: The sigmoids won't save you

#191
post #155

If you want a model, here's one: LLMs have never demonstrated the ability to go obviously beyond interpolating their training data. It takes an army of paid data producers solving homework problems to give ChatGPT the ability to do your homework. All vibecoded apps that turned out to be successful could put on a geological soil chart with other apps, probably on GitHub somewhere, on the corners. The prediction? They…

Just to play devils advocate: are we sure humans have demonstrated the ability to go beyond their training data? Like.. are we sure-sure about that?

Planes, trains, and automobiles are not natural features of the environment.

Re: The sigmoids won't save you

#192

Earlier quoted context omitted.

Mind you, he is only personally invested insofar as he's staked his reputation on it. Throughout his writing, he expresses the same point over and over again: desperately wants AI to slow down, advocates for politics that would slow it down, and most likely nothing would bring him greater peace than to see a sigmoid curve appear.

How convenient; when AGI doesn’t appear in 1-2 years his reputation is pristine because he slowed it down.

What do you want? This sounds like you have something against people making a claim in public, at all, on any topic of importance.

Re: The sigmoids won't save you

#193

> then what is their model? My mental model has been 3D computer graphics: doubling the polygon count had huge returns early on but delivered diminishing returns over time. Ultimately, you can't make something look more realistic than real. I don't know what the future holds, but the answer to the question "can LLMs be more realistic than real" will determine much about whether or not you think the curve will level o…

In 3D graphics there's diminishing returns on investment for the technology itself, but the real limit is one of economics. How to create all those assets required by the rendering tech, and make your money back. Preferably while also keeping your customers interested long term, by not becoming risk averse.

In the same fashion, LLMs have to pay for themselves to keep the trendlines going. In a whole-systems -sense, mind, not "$2000/month is cheaper than hiring a developer" while the rest of the economy collapses.

Re: The sigmoids won't save you

#194

Earlier quoted context omitted.

It's not that easy to assess diminishing returns with saturated benchmarks where asymptoting to 100% is mathematically baked in. I could point to the number of Erdos proofs being solved by AI going from 0 to many very recently as evidence for acceleration.

That is not evidence of acceleration, just of some measurable improvement compared to a previous model. After all, humans have made these breakthroughs since before recorded history—that never by itself implied accelerating intelligence.

What would be evidence of acceleration? What would be evidence of diminishing returns? Both questions are hard to answer because it's difficult to avoid constructing a metric where the conclusion is already baked in.

Re: The sigmoids won't save you

#195
Births is, sans miscalculation, a number that tracks exact events.

Is the "capability" number on these LLM strengh graphs as tangible?

I think it would be interesting to visit a reality that obeys arbitrary abstractions, but I would personally never go there.

Re: The sigmoids won't save you

#196

I felt the better takeaway from this was that it's impossible to know for certainty how long this will or will not continue regardless of the data or models you're using, because if you (or anyone else) could predict that accurately they'd be one of the richest people on the planet. I don't know when (or if) AI will implode or succeed with any degree of provable certainty, because that's not my area of expertise. Rat…

I think his agenda here is to point out that your probability distribution for AI outcomes should be broad (what you said), but most importantly: this means you must take seriously the possibility that we are gonna get superintelligence quite soon. Basically a lot of people say "but isn't it also pretty likely that we DON'T get superintelligence?" And, yes, it is. But superintelligence being even a remotely plausible…

That's literally the singularity though - the point past which predictions are meaningless.

My "plan" is hope for a benevolent intelligence that establishes a post-human government and then enjoy poat-scarcity society doing wood working or something.

Billionaires should probably be more worried.

Re: The sigmoids won't save you

#198
post #65

Earlier quoted context omitted.

IIRC that graph tracks capabilities as time_to_solve a task for humans (i.e. the model can now handle tasks that usually take a human ~8h). Which, depending on what tasks you look at, could be a reasonable finding. I could see Opus 4.6 handling tasks that take ~8h for humans, and that 5.1 couldn't previously handle (with 5.1 being "limited" at 4h tasks let's say). It is a bit arbitrary, but I think this is what they'…

I don't know why people are so impressed by 8h. I trained an LLM to write the whole Harry Potter series, and that took JK Rowling like 17 years. For my next point on the graph, I'll train the LLM to write the Bible, something that took humans >1500 years.

Have you used the models, out of interest? They routinely do things autonomously that are not in the training set that would take me 8h, and I wouldn't say I'm slow. The profile of tasks they can do this way is jagged, and maintaining architectural coherence ("months, not hours") is still beyond them, but they're perfectly capable of writing plans and sticking to them.

Re: The sigmoids won't save you

#199

Earlier quoted context omitted.

While this is very fun as a mathematical exercise, it's completely irrelevant as a real tool for getting a better understanding of unknown processes in the real world. The law only applies for certain types of processes, and is completely wrong for other types (e.g. a human who has lived 50 years may live 50 more, but one who has lived 100 years will certainly not live 100 more). So the question becomes: what type of…

But if you met an alien who said they'd been alive for 100 years you wouldn't assume they're on the verge of dropping dead: you would assume they live longer. It's a rough rule for when you don't have other information, and if you're arguing against it you need to specify what other information you're using to make that argument.

But it only applies for when you have a single data point: it's more likely to be from the middle of the distribution then the edge.

So meeting exactly 1 100 year old alien makes it decent odds that's somewhere near the middle of their lifespan.

Because if you grabbed one random human, chances are you'd find someone roughly middle aged.

Re: The sigmoids won't save you

#200
post #37

Lindy’s Law is an absolute gem, that I'm keeping. If we don't understand the fundamental limits to any particular kind of trend, our default assumption should be that it will continue for about as long as it has gone on already. We can, in fact, easily put a confidence interval on this. With 90% odds we're not in the first 5% of the trend, or the last 5% of the trend. Therefore it will probably go on between 1/19th l…

I feel like Lindy's law doesn't work for things whose observation is partly controlled by the thing itself. For example, take something like a fad or trend; they don't have a hard end date like human lifespan, so it should follow Lindy's law. However, the likelihood, on average across the population, that you observe a trend is going to be higher at the end of a trend lifecycle than at the beginning. This is baked in…

Well it only works when there is no information at all apart from the past frequency.

It's the solution to the tank problem. You know that the enemy number their tanks as they're produced. You capture a tank and know its number, N. What's the best guess about how many tanks the enemy has produced so far? As a pure mathematical model with no other details, the best guess is 2N. Of course in reality you have some ideas about how long it takes to make a tank, how many resources the enemy has etc.

Analogously you have information about the way trends develop.

Post reply on HN