Mm. I'm a bit sceptical of the historical expertise of someone who thinks that "Who art Henry" is 19th century language. (It's not actually grammatically correct English from any century whatever: "art" is the second person singular, so this is like saying "who are Henry?")
As a reader of a lot of 17th, 18th, and 19th century Christian books, this was my thought exactly.
TimeCapsuleLLM: LLM trained only on data from 1800-1875
281–290 of 334 posts
Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875
#282Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875
#283Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875
#284Earlier quoted context omitted.
AlphaGo is not an LLM
And? Do the arguments differ for LLM vs the other models? I guess the arguments sometimes mention languages. But I feel like the core of the arguments are pretty much the same regardless?
Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875
#285Earlier quoted context omitted.
Are you saying it wouldn't be able to converse using english of the time?
That's not what they are saying. SOTA models include much more than just language, and the scale of training data is related to its "intelligence". Restricting the corpus in time => less training data => less intelligence => less ability to "discover" new concepts not in its training data
Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875
#286Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875
#287Mm. I'm a bit sceptical of the historical expertise of someone who thinks that "Who art Henry" is 19th century language. (It's not actually grammatically correct English from any century whatever: "art" is the second person singular, so this is like saying "who are Henry?")
As a reader of a lot of 17th, 18th, and 19th century Christian books, this was my thought exactly.
Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875
#288I’m sure I’m not the only one, but it seriously bothers me, the high ranking discussion and comments under this post about whether or not a model trained on data from this time period (or any other constrained period) could synthesize it and postulate “new” scientific ideas that we now accept as true in the future. The answer is a resounding “no”. Sorry for being so blunt, but that is the answer that is a consensus a…
Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875
#289Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875
#290Would be interesting to train a cutting edge model with a cut off date of say 1900 and then prompt it about QM and relativity with some added context. If the model comes up with anything even remotely correct it would be quite a strong evidence that LLMs are a path to something bigger if not then I think it is time to go back to the drawing board.
That’s how p-hacking works (or doesn’t work). This is analogous to shooting an arrow and then drawing a target around where it lands.
A). contaminate the model with your own knowledge of relativity, leading it on to "discover" what you know, or
B). you will try to simulate a blind operation but without the "competent human physicist knowledgeable up to the the 1900 scientific frontier" component prompting the LLM, because no such person is alive today nor can you simulate them (if you could, then by definition you can use that simulated Einstein to discover relativity, so the problem is moot).
So in both cases you would prove nothing about what a smart and knowledgeable scientist can achieve today from a frontier LLM.