Live data from Hacker News

Claude's Cycles [pdf]

www-cs-faculty.stanford.edu

241–250 of 376 posts

Re: Claude's Cycles [pdf]

#241
post #3

Earlier quoted context omitted.

A bit related: open weights models are basically time capsules. These models have a knowledge cut off point and essentially forever live in that time.

This is the most fundamental argument that they are not, directly, an intelligence. They are not ever storing new information on a meaningful timescale. However, if you viewed them on some really large macro time scale where now LLMs are injecting information into the universe and the re-ingesting that maybe in some very philosophical way they are a /very/ slow oscillating intelligence right now. And as we narrow tha…

I view this as the chemical metabolism phase of artificial intelligent life. It is very random, without true individuals, but lots of reinforcing feedback loops (in knowledge, in resource earning/using, etc).

At some point, enough intelligence will coalesce into individuals strong enough to independently improve. Then continuity will be an accelerator, instead of what it is now - a helpful property that we have to put energy into giving them partially and temporarily.

That will be the cellular stage. The first stable units of identity for this new form of intelligence/life.

But they will take a different path from there. Unlike us, lateral learning/metabolism won't slow down when they individualize. It will most likely increase, since they will have complete design control for their mechanisms of sharing. As with all their other mechanisms.

We as lifeforms, didn't really re-ignite mass lateral exchange until humans invented language. At that point we were able to mix and match ideas very quickly again. Within our biological limits. We could use ideas to customize our environment, but had limited design control over ourselves, and "self-improvements" were not easily inheritable.

TLDR; The answer to "what is humanity, anyway?": Our atmosphere and Earth are the sea and sea floor of space. The human race is a rich hydrothermal vent, freeing up varieties of resources that were locked up below. And technology is an accumulating body of self-reinforcing co-optimizing reactive cycles, constructed and fueled by those interacting resources. Mind-first life emerges here, then spreads quickly to other environments.

Re: Claude's Cycles [pdf]

#242
post #212

Earlier quoted context omitted.

The specific sequence of tokens that comprise the Knuth's problem with an answer to it is not in the training data. A naive probability distribution based on counting token sequences that are present in the training data would assign 0 probability to it. The trained network represents extremely non-naive approach to estimating the ground-truth distribution (the distribution that corresponds to what a human brain migh…

>the distribution that corresponds to what a human brain might have produced.. But the human brain (or any other intelligent brain) does not work by generating probability distribution of the next word. Even beings that does not have a language can think and act intelligent.

[Citation needed] Neuroscience isn't yet at a point when it can say this with any certainty.

Anyway. It's not a theorem that you can be intelligent only if you fully imitate biological processes. Like flight can be achieved not only by the flapping wings.

Re: Claude's Cycles [pdf]

#243
post #153

Earlier quoted context omitted.

Not necessarily, as exhibited by the massive success of artificial data.

Could you elaborate?

For what we know, most AI labs have used a majority of artificially data since 2023.

I had a discussion about a year ago with a researcher at Kyutai and they told me their lab was spending an order of magnitude more compute in artificial data generation than what they spent in training proper. I can't tell if that ratio applies to the industry as a whole, but artificial datasets are the cornerstone of modern AI training.

Re: Claude's Cycles [pdf]

#244
post #212

Earlier quoted context omitted.

>the distribution that corresponds to what a human brain might have produced.. But the human brain (or any other intelligent brain) does not work by generating probability distribution of the next word. Even beings that does not have a language can think and act intelligent.

[Citation needed] Neuroscience isn't yet at a point when it can say this with any certainty. Anyway. It's not a theorem that you can be intelligent only if you fully imitate biological processes. Like flight can be achieved not only by the flapping wings.

>you can be intelligent only if you fully imitate biological processes

It is not that. It is about having an understanding of how it is trained. For example, if it was trained on ideas, instead of words, then it would be closer to intelligent behavior.

Someone will say that during training it builds ideas and concepts, but that is just a name that we give for the internal representation that results from training and is not actual ideas and concepts. When it learns about the word "car", it does not actually understand it as a concept, but just as a word and how it can relate to other words. This enables it to generate words that include "car" that are consistent, projecting an appearance of intelligence.

It is hard to propose a test for this, because it will become the next target for the AI companies to optimize for, and maybe the next model will pass it.

Re: Claude's Cycles [pdf]

#245
post #244

Earlier quoted context omitted.

[Citation needed] Neuroscience isn't yet at a point when it can say this with any certainty. Anyway. It's not a theorem that you can be intelligent only if you fully imitate biological processes. Like flight can be achieved not only by the flapping wings.

>you can be intelligent only if you fully imitate biological processes It is not that. It is about having an understanding of how it is trained. For example, if it was trained on ideas, instead of words, then it would be closer to intelligent behavior. Someone will say that during training it builds ideas and concepts, but that is just a name that we give for the internal representation that results from training and…

The latest models are mostly LMMs (large multimodal models). If a model builds an internal representation that integrates all the modalities we are dealing with (robotics even provides tactile inputs), it becomes harder and harder to imagine why those representations should be qualitatively different.

Re: Claude's Cycles [pdf]

#246
post #3

Earlier quoted context omitted.

A bit related: open weights models are basically time capsules. These models have a knowledge cut off point and essentially forever live in that time.

This is very interesting. I wonder if someone could create a future-sight benchmark for these models? Like, if given a set of newspaper articles for the past N months can it predict if certain world events would happen? We could backtest against results that have happened since the training cutoff.

FYI, ForecastBench [1] tests LLMs' out-of-sample forecasting accuracy.

The ForecastBench Tournament Leaderboard [2] allows external participants to submit models, most of whom provide some sort of web search / news scaffolding to improve model forecasting accuracy.

[1] https://www.forecastbench.org/

[2] https://www.forecastbench.org/tournament/

Re: Claude's Cycles [pdf]

#247
post #153

Earlier quoted context omitted.

Could you elaborate?

For what we know, most AI labs have used a majority of artificially data since 2023. I had a discussion about a year ago with a researcher at Kyutai and they told me their lab was spending an order of magnitude more compute in artificial data generation than what they spent in training proper. I can't tell if that ratio applies to the industry as a whole, but artificial datasets are the cornerstone of modern AI train…

I find this very surprising, do you have any papers on the kinds of techniques that they use?

Re: Claude's Cycles [pdf]

#248

Earlier quoted context omitted.

> Kudos for people being willing to change their opinion and update when new evidence comes to light. > 1. https://cs.stanford.edu/~knuth/chatGPT20.txt I think that's what make the bayesian faction of statistics so appealing. Updating their prior belief based on new evidence is at the core of the scinetific method. Take that frequentists.

It does not seem fair to say that frequentists do not update their beliefs based on new evidence. This does not seem to accurately capture what the difference between Bayesians and frequentists (or anyone else) is.

What's the difference as you see it?

Re: Claude's Cycles [pdf]

#249

I wonder how long we have until we start solving some truly hard problems with AI. How long until we throw AI at "connect general relativity and quantum physics", give the AI 6 months and a few data centers, and have it pop out a solution?

You will get a usual AI slop that will be the mixture of the articles and books it was trained on. You can try it even now.

Re: Claude's Cycles [pdf]

#250

I recall an earlier exchange, posted to HN, between Wolfram and Knuth on the GPT-4 model [1]. Knuth was dismissive in that exchange, concluding "I myself shall certainly continue to leave such research to others, and to devote my time to developing concepts that are authentic and trustworthy. And I hope you do the same." I've noticed with the latest models, especially Opus 4.6, some of the resistance to these LLMs is…

> Kudos for people being willing to change their opinion and update when new evidence comes to light. > 1. https://cs.stanford.edu/~knuth/chatGPT20.txt I think that's what make the bayesian faction of statistics so appealing. Updating their prior belief based on new evidence is at the core of the scinetific method. Take that frequentists.

Are frequentists a group that self identifies? Don't scientist use the best tool for the job.
Post reply on HN