Live data from Hacker News

TimeCapsuleLLM: LLM trained only on data from 1800-1875

github.com

11–20 of 334 posts

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#11
post #9

Would be interesting to train a cutting edge model with a cut off date of say 1900 and then prompt it about QM and relativity with some added context. If the model comes up with anything even remotely correct it would be quite a strong evidence that LLMs are a path to something bigger if not then I think it is time to go back to the drawing board.

I suppose the vast majority of training data used for cutting edge models was created after 1900.

Ofc they are because their primary goal is to be useful and to be useful they need to always be relevant.

But considering that Special Relativity was published in 1905 which means all its building blocks were already floating in the ether by 1900 it would be a very interesting experiment to train something on Claude/Gemini scale and then say give in the field equations and ask it to build a theory around them.

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#14
Could this be an experiment to show how likely LLMs are to lead to AGI, or at least intelligence well beyond our current level?

If you could only give it texts and info and concepts up to Year X, well before Discovery Y, could we then see if it could prompt its way to that discovery?

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#16
post #9

Earlier quoted context omitted.

I suppose the vast majority of training data used for cutting edge models was created after 1900.

Ofc they are because their primary goal is to be useful and to be useful they need to always be relevant. But considering that Special Relativity was published in 1905 which means all its building blocks were already floating in the ether by 1900 it would be a very interesting experiment to train something on Claude/Gemini scale and then say give in the field equations and ask it to build a theory around them.

How can you train a Claude/Gemini scale model if you’re limited to <10% of the training data?

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#17
Oh I have really been thinking long about this. The intelligence that we have in these models represent a time.

Now if I train a foundation models with docs from library of Alexandria and only those texts of that period, I would have a chance to get a rudimentary insight on what the world was like at that time.

And maybe time shift further more.

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#18

Would be interesting to train a cutting edge model with a cut off date of say 1900 and then prompt it about QM and relativity with some added context. If the model comes up with anything even remotely correct it would be quite a strong evidence that LLMs are a path to something bigger if not then I think it is time to go back to the drawing board.

Looking at the training data I don't think it will know anything.[0] Doubt On the Connexion of the Physical Sciences (1834) is going to have much about QM. While the cut-off is 1900, it seems much of the texts a much closer to 1800 than 1900.

[0] https://github.com/haykgrigo3/TimeCapsuleLLM/blob/main/Copy%...

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#19
post #14

Could this be an experiment to show how likely LLMs are to lead to AGI, or at least intelligence well beyond our current level? If you could only give it texts and info and concepts up to Year X, well before Discovery Y, could we then see if it could prompt its way to that discovery?

I think not if only for the fact that the quantity of old data isn't enough to train anywhere near a SoTA model, until we change some fundamentals of LLM architecture

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#20
post #9

Earlier quoted context omitted.

I suppose the vast majority of training data used for cutting edge models was created after 1900.

Ofc they are because their primary goal is to be useful and to be useful they need to always be relevant. But considering that Special Relativity was published in 1905 which means all its building blocks were already floating in the ether by 1900 it would be a very interesting experiment to train something on Claude/Gemini scale and then say give in the field equations and ask it to build a theory around them.

His point is that we can't train a Gemini 3/Claude 4.5 etc model because we don't have the data to match the training scale of those models. There aren't trillions of tokens of digitized pre-1900s text.
Post reply on HN