Live data from Hacker News

TimeCapsuleLLM: LLM trained only on data from 1800-1875

github.com

281–290 of 334 posts

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#281
post #81

Mm. I'm a bit sceptical of the historical expertise of someone who thinks that "Who art Henry" is 19th century language. (It's not actually grammatically correct English from any century whatever: "art" is the second person singular, so this is like saying "who are Henry?")

As a reader of a lot of 17th, 18th, and 19th century Christian books, this was my thought exactly.

That text was from v0, the responses improved from there.

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#284
post #224

Earlier quoted context omitted.

AlphaGo is not an LLM

And? Do the arguments differ for LLM vs the other models? I guess the arguments sometimes mention languages. But I feel like the core of the arguments are pretty much the same regardless?

The discussion is about training an LLM on old text and then asking it about new concepts.

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#285

Earlier quoted context omitted.

Are you saying it wouldn't be able to converse using english of the time?

That's not what they are saying. SOTA models include much more than just language, and the scale of training data is related to its "intelligence". Restricting the corpus in time => less training data => less intelligence => less ability to "discover" new concepts not in its training data

Could always train them on data up to 2015ish and then see if you can rediscover LLMs. There's plenty of data.

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#287
post #81

Mm. I'm a bit sceptical of the historical expertise of someone who thinks that "Who art Henry" is 19th century language. (It's not actually grammatically correct English from any century whatever: "art" is the second person singular, so this is like saying "who are Henry?")

As a reader of a lot of 17th, 18th, and 19th century Christian books, this was my thought exactly.

What kind of Christian books do you read?Jonathan Edwards, John Bunyan, J.C. Ryle, C.H. Spurgeon?

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#288
post #244

I’m sure I’m not the only one, but it seriously bothers me, the high ranking discussion and comments under this post about whether or not a model trained on data from this time period (or any other constrained period) could synthesize it and postulate “new” scientific ideas that we now accept as true in the future. The answer is a resounding “no”. Sorry for being so blunt, but that is the answer that is a consensus a…

I think it's pretty likely the answer is no, but the idea here is that you could actually test that assertion. I'm also pessimistic about it but that doesn't mean it wouldn't be a little interesting to try.

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#289

Earlier quoted context omitted.

As a reader of a lot of 17th, 18th, and 19th century Christian books, this was my thought exactly.

That text was from v0, the responses improved from there.

That text was from the example prompt, not from the models response

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#290

Would be interesting to train a cutting edge model with a cut off date of say 1900 and then prompt it about QM and relativity with some added context. If the model comes up with anything even remotely correct it would be quite a strong evidence that LLMs are a path to something bigger if not then I think it is time to go back to the drawing board.

That’s how p-hacking works (or doesn’t work). This is analogous to shooting an arrow and then drawing a target around where it lands.

Yes, I don't understand how such an experiment could work. You either:

A). contaminate the model with your own knowledge of relativity, leading it on to "discover" what you know, or

B). you will try to simulate a blind operation but without the "competent human physicist knowledgeable up to the the 1900 scientific frontier" component prompting the LLM, because no such person is alive today nor can you simulate them (if you could, then by definition you can use that simulated Einstein to discover relativity, so the problem is moot).

So in both cases you would prove nothing about what a smart and knowledgeable scientist can achieve today from a frontier LLM.

Post reply on HN