Live data from Hacker News

TimeCapsuleLLM: LLM trained only on data from 1800-1875

github.com

81–90 of 334 posts

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#81
Mm. I'm a bit sceptical of the historical expertise of someone who thinks that "Who art Henry" is 19th century language. (It's not actually grammatically correct English from any century whatever: "art" is the second person singular, so this is like saying "who are Henry?")

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#82
post #81

Mm. I'm a bit sceptical of the historical expertise of someone who thinks that "Who art Henry" is 19th century language. (It's not actually grammatically correct English from any century whatever: "art" is the second person singular, so this is like saying "who are Henry?")

As a reader of a lot of 17th, 18th, and 19th century Christian books, this was my thought exactly.

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#84
post #78

Earlier quoted context omitted.

It doesn’t need to know about QM or reactivity just about the building blocks that led to them. Which were more than around in the year 1900. In fact you don’t want it to know about them explicitly just have enough background knowledge that you can manage the rest via context.

LLMs are models that predict tokens. They don't think, they don't build with blocks. They would never be able to synthesize knowledge about QM.

You realize parent said "This would be an interesting way to test proposition X" and you responded with "X is false because I say say", right?

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#85

If the output of this is even somewhat coherent, it would disprove the argument that mass amounts of copyrighted works are required to train an LLM. Unfortunately that does not appear to be the case here.

Take a look at The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text (https://arxiv.org/pdf/2506.05209). They build a reasonable 7B parameter model using only open-licensed data.

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#86
post #28

Earlier quoted context omitted.

> Could this be an experiment to show how likely LLMs are to lead to AGI, or at least intelligence well beyond our current level? You'd have to be specific what you mean by AGI: all three letters mean a different thing to different people, and sometimes use the whole means something not present in the letters. > If you could only give it texts and info and concepts up to Year X, well before Discovery Y, could we then…

> You'd have to be specific what you mean by AGI Well, they obviously can't. AGI is not science, it's religion. It has all the trappings of religion: prophets, sacred texts, origin myth, end-of-days myth and most importantly, a means to escape death. Science? Well, the only measure to "general intelligence" would be to compare to the only one which is the human one but we have absolutely no means by which to describe…

> And no, the "brain is a computer" is not a scientific description, it's a metaphor.

Disagree. A brain is turing complete, no? Isn't that the definition of a computer? Sure, it may be reductive to say "the brain is just a computer".

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#88
post #81

Mm. I'm a bit sceptical of the historical expertise of someone who thinks that "Who art Henry" is 19th century language. (It's not actually grammatically correct English from any century whatever: "art" is the second person singular, so this is like saying "who are Henry?")

Can you elaborate on this? After skimming the README, I understand that "Who art Henry" is the prompt. What should be the correct 19th century prompt?

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#89
post #14

Could this be an experiment to show how likely LLMs are to lead to AGI, or at least intelligence well beyond our current level? If you could only give it texts and info and concepts up to Year X, well before Discovery Y, could we then see if it could prompt its way to that discovery?

It'd be difficult to prove that you hadn't leaked information to the model. The big gotcha of LLMs is that you train them on BIG corpuses of data, which means it's hard to say "X isn't in this corpus", or "this corpus only contains Y". You could TRY to assemble a set of training data that only contains text from before a certain date, but it'd be tricky as heck to be SURE about it. Ways data might leak to the model t…

There's also severe selection effects: what documents have been preserved, printed, and scanned because they turned out to be on the right track towards relativity?

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#90

Earlier quoted context omitted.

Yeah but... we still might not know if it could do that because we were really close by 1900 or if the LLM is very smart.

What's the bar here? Does anyone say "we don't know if Einstein could do this because we were really close or because he was really smart?" I by no means believe LLMs are general intelligence, and I've seen them produce a lot of garbage, but if they could produce these revolutionary theories from only <= year 1900 information and a prompt that is not ridiculously leading, that would be a really compelling demonstrati…

> Does anyone say "we don't know if Einstein could do this because we were really close or because he was really smart?"

Kind of, how long would it have realistically taken for someone else (also really smart) to come up with the same thing if Einstein wouldn't have been there?

Post reply on HN