Live data from Hacker News

The sigmoids won't save you

astralcodexten.com

211–220 of 297 posts

Re: The sigmoids won't save you

#211
post #140

1. Scott Alexander is famous for writing about topics he knows little about. I'm glad to see he's found a subject he knows little about but so does everyone else. 2. What's even worse than predicting that some growth curve flattens before X happens is predicting it will flatten before X happens but after Y happens, which is what we see when it comes to AI in software development. Too many people predict that AI will…

>1. Scott Alexander is famous for writing about topics he knows little about. I'm glad to see he's found a subject he knows little about but so does everyone else. This is kinda laughable. Scott has been thinking and writing about AI for a long time

That's what I'm saying. I'm glad to see that he's found a subject where his lack of knowledge isn't a glaring handicap. When I come across his posts I usually feel uncomfortable because they read like a bright 4-year-old child trying hard to explain how a car works, only not cute (yeah, I know that trying to see what conclusions you can come to, Aristotle-like, from a basis of ignorance and without careful study is the whole point, but I never found this Memento-style game appealing).

Re: The sigmoids won't save you

#212

> Why do scaling laws work? Strictly speaking, the original paradigm of scaling laws doesn't work any more. The assumption that we could achieve better performance simply through "vertical scaling" ie infusing models with exponentially more parameters and pre-training data, is no longer the driving force of AI progress. Instead, the industry has pivoted toward inference-time scaling. Rather than relying solely on a m…

The "law" part of scaling laws is about predicting validation cross-entropy loss from the training configuration, analogous to physical laws allowing to predict one quantity based on the measurement of another. Most scaling laws take the form of an irreducible error plus additional terms that asymptotically decay to zero. So that there is a wall you can approach but not cross (the irreducible error) is an integrated part of the scaling law paradigm. That it isn't economical to keep increasing model size to squeeze out a few more drops of cross-entropy doesn't mean scaling laws stopped working.

Strictly speaking, "Why do scaling laws work?" is a question about the theoretical reasons the asymptotic decay takes the particular mathematical shape that it does.

Re: The sigmoids won't save you

#213
post #65

Earlier quoted context omitted.

I don't know why people are so impressed by 8h. I trained an LLM to write the whole Harry Potter series, and that took JK Rowling like 17 years. For my next point on the graph, I'll train the LLM to write the Bible, something that took humans >1500 years.

Have you used the models, out of interest? They routinely do things autonomously that are not in the training set that would take me 8h, and I wouldn't say I'm slow. The profile of tasks they can do this way is jagged, and maintaining architectural coherence ("months, not hours") is still beyond them, but they're perfectly capable of writing plans and sticking to them.

Yeah, I use them all the time. I just don't see any good argument that it's anything other than statistical pattern matching plus some sort of logic encoded in language. My overfitted LLM obviously didn't arrive at Harry Potter the same way JK Rowling did, so the amount of time she spent writing it is completely irrelevant to any discussion about whether or not the LLM should be able to reproduce it. discussions of AGI if it took her an hour or a decade to write it, it has seen the result, so it can reproduce it.

Re: The sigmoids won't save you

#214
post #202

Earlier quoted context omitted.

And this is an asic that is still operating digitally. Imagine a chip with baked it weights that does its math analogue with 20x reduction in number of circuit elements needed to do a multiplication op. If there's a breakthrough in memristors, you could end up with another 20x reduction in circuit elements (get rid of memory bottlnecks, start doing multiplication ops as log transform voltage addition) The ceiling is…

Except weights will be unstable - temperature and frequency dependent - and we still have issues delivering analog circuits reliably to the spec. So it would take multiple attempts. But yeah, as soon as the digital models start to plateau, ASICs and then this will happen.

You could SFT for temperature and frequency stability.

I'm not even kidding. Modern ML systems already eat errors - what's one more error type for them to eat?

Re: The sigmoids won't save you

#215
post #188

Earlier quoted context omitted.

The book "Origins of Efficiency" by Brian Potter discusses this. Stacked sigmoids are a well-understood idea in innovation. The idea that exponential growth will continue with stacked sigmoids is also not a given. An example is the nail. Nails used to be about half a percent of US GDP. That's a pretty big number! A series of innovations stacked on each other (each innovation having its own sigmoid) to reduce the cost…

I don't disagree with you, but your example of nails and their cost reductions made me wonder whether we reached a meaningful limit in say, some fundamental material terms, or whether we just reached a limit in terms of return on investment. Return on investment can be too low because the investment required is really high, but it can also be too low because the returns are just limited. If prices had dropped 90%, su…

There are more ideas to try, but this doesn't necessarily mean they're better ones or that we're bound to come up with them. Training transformers on corpuses of human text has worked extremely well at enabling AIs to generate continuations which are consistent with human text including really useful human text like code, but the limit might be the human text rather than the transformer architecture...

The ROI for cornering the nail market seems like it could have been big. The ROI for making something significantly more efficient than a ICE would have been very high for most of the last century and technology that is better in many respects than ICEs does now exist, but it took us roughly a century to get there. The ROI for coming up with something that's better than the ~1% annual efficiency improvement on turbofans would be extremely high, but we don't know what that is (probably some sort of propfan, but that idea's over 50 years old...)

Re: The sigmoids won't save you

#216

Earlier quoted context omitted.

AGI has become such a meaningless nondescript term, arguing when or how it is here has become pointless. Even OpenAI caved in and removed their AGI clause from their contract with Microsoft because they weren't fully sure that we are not there yet. The original ARC AGI was hailed as proof that AGI is not here yet, but now that ARC 1 and 2 got saturated, noone wanted to consider that perhaps we crossed the point where…

To your point, if we had truly unlimited context to the point where at least that instance of a model could “learn” and have what seems like a continuous “consciousness” I think many of us would think that we’ve attained AGI. Right now we have an incredibly smart thing with severe short term memory loss, and it’s hard for us to reconcile that as it’s so different from us.

Quite a few people were already led to believe that these models are conscious when we had a fraction of current context lengths. Right now the biggest problem is that the "session" info in form of the current conversation gets lost too quickly, but that has become largely an implementation detail. You could fit an entire life's story into modern context windows. With some clever context management, you could probably build something that feels like what you describe. If we truly had this sort of short-term to long-term memory (i.e. from prompt context to weights) system on a technical foundation, we'd probably be closer to runaway superintelligence than mere AGI that could beat most humans on most tasks.

Re: The sigmoids won't save you

#217

FYI: The author has predicted that "AGI" will be here in 1-2 years and has staked his public reputation on it. He is personally invested in trendlines being lindy rather than sigmoid. I don't think you can use lindy on trends as if trends are static objects, but that's another conversation.

So, this is not quite right: Alexander contributed to the report, but his personal opinion is more like the mid-2030s[1]. Freddie feels like this is him backing down from the original statement, but in fact he said this at the time the report was published, and in fact pointed out a graf below the quote that Freddie claims does tie him to 2027: > Do we really think things will move this fast? Sort of no - between the…

AI boosters really are detached from reality.

LLMs are nothing close to AGI and not going to lead to it, they can’t distinguish right from wrong, they can’t count, they can’t reason, they generate plausible text from a vast databank of connected text.

Apparently that is enough to fool many people but it’s nothing close to AGI which would require internal models of the world, reasoning etc.

We are nowhere close to AGI and the fools who predicted we were will unfortunately keep lying about their stated timelines when it inevitably doesn’t arrive. You’re already hedging and trying to caveat previous predictions, as OpenAI did with their AGI predictions which they’re now furiously back-pedalling on.

Re: The sigmoids won't save you

#218

FYI: The author has predicted that "AGI" will be here in 1-2 years and has staked his public reputation on it. He is personally invested in trendlines being lindy rather than sigmoid. I don't think you can use lindy on trends as if trends are static objects, but that's another conversation.

Ok, but you can just look at the METR curve. Mythos saturated the 50% time horizon. The 80% is now at 3 hours. The rate of progress is accelerating not slowing down. There’s no indication yet that this is a sigmoid!

The METR task set contains no tasks with a duration greater than 32 hours (conservatively eyeballed from Figure 3: https://arxiv.org/abs/2503.17354 ), so any prediction that naively forecasts a longer time horizon is trivially incorrect. I guess that won't lead to a sigmoid-looking graph though, since METR will likely switch to a different evaluation methodology at that point and stop updating the old curve.

Re: The sigmoids won't save you

#219
post #129

Earlier quoted context omitted.

It was a metaphor. I meant, and later clarified, an intellectual arms race. BTW your handle is an actual Czech word, minus a diacritic sign ("křupan"), and a bit amusing one. It basically means hillbilly. Not that it matters, just FYI. Anyway: AI will be used in military context, and it probably already is. Both for target acquisition and maybe even driving the weapon itself. As of now, the Ukrainians are almost cert…

That's funny, I was told my real last name is a swear word in Czech

Yeah, but a relatively mild one as far as the spectrum of Slavic curses go.

You would be a "křupan" if you wore agricultural boots to a fancy restaurant, or talk to a lady in an uncultured way. Basically, a hickey who was never taught proper manners.

Re: The sigmoids won't save you

#220
post #155

If you want a model, here's one: LLMs have never demonstrated the ability to go obviously beyond interpolating their training data. It takes an army of paid data producers solving homework problems to give ChatGPT the ability to do your homework. All vibecoded apps that turned out to be successful could put on a geological soil chart with other apps, probably on GitHub somewhere, on the corners. The prediction? They…

Just to play devils advocate: are we sure humans have demonstrated the ability to go beyond their training data? Like.. are we sure-sure about that?

Are you suggesting intelligent design got us here?
Post reply on HN