Live data from Hacker News

TimeCapsuleLLM: LLM trained only on data from 1800-1875

github.com

231–240 of 334 posts

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#231

Would be interesting to train a cutting edge model with a cut off date of say 1900 and then prompt it about QM and relativity with some added context. If the model comes up with anything even remotely correct it would be quite a strong evidence that LLMs are a path to something bigger if not then I think it is time to go back to the drawing board.

We've thought of doing this sort of exercise at work but mostly hit the wall of data becoming a lot more scare the further back in time we go. Particularly high quality science data - even going pre 1970 (and that's already a stretch) you lose a lot of information. There's a triple whammy of data still existing, being accessible in any format, and that format being suitable for training an LLM. Then there's the complications of wanting additional model capabilities that won't leak data causally.

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#232

Earlier quoted context omitted.

This alone still wouldn't be a clear demonstration that AGI is around the corner. It's quite possible a LLM could've done Einstein's job, if Einstein's job was truly just synthesising already available information into a coherent new whole. (I couldn't say, I don't know enough of the physics landscape of the day to claim either way.) It's still unclear whether this process could be merely continued, seeded only with…

This does make me think about Kuhn's concept of scientific revolutions and paradigms, and that paradigms are incommensurate with one another. Since new paradigms can't be proven or disproven by the rules of the old paradigm, if an LLM could independently discover paradigm shifts similar to moving from Newtonian gravity to general relativity, then we have empirical evidence of an LLM performing a feature of general in…

His concept sounds odd. There will always be many hints of something yet to be discovered, simply by the nature of anything worth discovering having an influence on other things.

For instance spectroscopy enables one to look at the spectra emitted by another 'thing', perhaps the sun, and it turns out that there's little streaks within the spectra the correspond directly to various elements. This is how we're able to determine the elemental composition of things like the sun.

That connection between elements and the patterns in their spectra was discovered in the early 1800s. And those patterns are caused by quantum mechanical interactions and so it was perhaps one of the first big hints of quantum mechanics, yet it'd still be a century before we got to relativity, let alone quantum mechanics.

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#233

Earlier quoted context omitted.

You would find things in there that were already close to QM and relativity. The Michelson-Morley experiment was 1887 and Lorentz transformations came along in 1889. The photoelectric effect (which Einstein explained in terms of photons in 1905) was also discovered in 1887. William Clifford (who _died_ in 1889) had notions that foreshadowed general relativity: "Riemann, and more specifically Clifford, conjectured tha…

I presume that's what the parent post is trying to get at? Seeing if, given the cutting edge scientific knowledge of the day, the LLM is able to synthesis all it into a workable theory of QM by making the necessary connections and (quantum...) leaps Standing on the shoulders of giants, as it were

I think it's not productive to just have the LLM site like Mycroft in his armchair and from there, return you an excellent expert opinion.

THat's not how science works.

The LLM would have to propose experiments (which would have to be simulated), and then develop its theories from that.

Maybe there had been enough facts around to suggest a number of hypotheses, but the LLM in its curent form won't be able to confirm them.

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#234

Earlier quoted context omitted.

Who said anything of a minimum bar? "If so", not "Only if so".

I think the problem is the formulation "If so, AGI can't be far behind". I think that if a model were advanced enough such that it could do Einstein's job, that's it; that's AGI. Would it be ASI? Not necessarily, but that's another matter.

The phone in your pocket can perform arithmetic many orders of magnitude faster than any human, even the fringe autistic savant type. Yet it's still obviously not intelligent.

Excellence at any given task is not indicative of intelligence. I think we set these sort of false goalposts because we want something that sounds achievable but is just out of reach at one moment in time. For instance at one time it was believed that a computer playing chess at the level of a human would be proof of intelligence. Of course it sounds naive now, but it was genuinely believed. It ultimately not being so is not us moving the goalposts, so much as us setting artificially low goalposts to begin with.

So for instance what we're speaking of here is logical processing across natural language, yet human intelligence predates natural language. It poses a bit of a logical problem to then define intelligence as the logical processing of natural language.

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#235

Earlier quoted context omitted.

I think the problem is the formulation "If so, AGI can't be far behind". I think that if a model were advanced enough such that it could do Einstein's job, that's it; that's AGI. Would it be ASI? Not necessarily, but that's another matter.

The phone in your pocket can perform arithmetic many orders of magnitude faster than any human, even the fringe autistic savant type. Yet it's still obviously not intelligent. Excellence at any given task is not indicative of intelligence. I think we set these sort of false goalposts because we want something that sounds achievable but is just out of reach at one moment in time. For instance at one time it was believ…

[dead]

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#236

Earlier quoted context omitted.

They do not manipulate concepts. There is no representation of a concept for them to manipulate. It may, however, turn out that in doing what they do, they are effectively manipulating concepts, and this is what I was alluding to: by building the model, even though your approach was through tokenization and whatever term you want to use for the network, you end up accidentally building something that implicitly manip…

>They do not manipulate concepts. There is no representation of a concept for them to manipulate. Yes, they do. And of course there is. And there's plenty of research on the matter. >It may, however, turn out that in doing what they do, they are effectively manipulating concepts There is no effectively here. Text is what goes in and what comes out, but it's by no means what they manipulate internally. >Nevertheless "…

please pass on a link to a solid research paper that supports the idea that to "find the next probable token", LLM's manipulate concepts ... just one will do.

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#237

Would be interesting to train a cutting edge model with a cut off date of say 1900 and then prompt it about QM and relativity with some added context. If the model comes up with anything even remotely correct it would be quite a strong evidence that LLMs are a path to something bigger if not then I think it is time to go back to the drawing board.

I think it would raise some interesting questions, but if it did yield anything noteworthy, the biggest question would be why that LLM is capable of pioneering scientific advancements and none of the modern ones are.

Or maybe, LLMs are pioneering scientific advancements - people are using LLMs to read papers, choose what problems to work on, come up with experiments, analyze results, and draft papers, etc., at this very moment. Except they eventually stick their human names on the cover so we almost never know.

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#238
post #130
post #88

Earlier quoted context omitted.

Can you elaborate on this? After skimming the README, I understand that "Who art Henry" is the prompt. What should be the correct 19th century prompt?

Who art thou? (Well, not 19th century...)

The problem is the subjunctive mood of the word "art".

"Art thou" should be translated into modern English as "are you to be", and so works better with things (what are you going to be), or people who are alive, and have a future (who are you going to be?).

Those are probably the contexts you are thinking of.

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#239
I wonder how representative this is of life in those days. Most written communication was official back then. Books, newspapers. Plays. All very formal and staged. There's not much real life interaction between common people in that. In fact I would imagine a lot of people were illiterate.

With the internet and pervasive text communication and audio video recording we have the unique ability to make an LLM mimic daily life but I doubt that would be possible for those days.

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#240

Earlier quoted context omitted.

>They do not manipulate concepts. There is no representation of a concept for them to manipulate. Yes, they do. And of course there is. And there's plenty of research on the matter. >It may, however, turn out that in doing what they do, they are effectively manipulating concepts There is no effectively here. Text is what goes in and what comes out, but it's by no means what they manipulate internally. >Nevertheless "…

please pass on a link to a solid research paper that supports the idea that to "find the next probable token", LLM's manipulate concepts ... just one will do.

Revealing emergent human-like conceptual representations from language prediction - https://www.pnas.org/doi/10.1073/pnas.2512514122

Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic Task - https://openreview.net/forum?id=DeG07_TcZvT

On the Biology of a Large Language Model - https://transformer-circuits.pub/2025/attribution-graphs/bio...

Emergent Introspective Awareness in Large Language Models - https://transformer-circuits.pub/2025/introspection/index.ht...

Post reply on HN