Live data from Hacker News

TimeCapsuleLLM: LLM trained only on data from 1800-1875

github.com

121–130 of 334 posts

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#121

Earlier quoted context omitted.

This is definitely wrong, most AI researchers DO NOT agree with LeCun. Most ML researchers think AGI is imminent.

their employment and business opportunities depend on the hype, so they will continue to 'think' that (on xitter) despite the current SOTA of transformers-based models being 3 year old GPT4, and no revolutionary new architecture in sight.

You're going to be in for a very rude awakening.

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#122

Earlier quoted context omitted.

Yann LeCun spoke explicitly on this idea recently and he asserts definitively that the LLM would not be able to add anything useful in that scenario. My understanding is that other AI researchers generally agree with him, and that it's mostly the hype beasts like Altman that think there is some "magic" in the weights that is actually intelligent. Their payday depends on it, so it is understandable. My opinion is that…

This is definitely wrong, most AI researchers DO NOT agree with LeCun. Most ML researchers think AGI is imminent.

Do you have poll of ML researchers that shows this?

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#123

Would be interesting to train a cutting edge model with a cut off date of say 1900 and then prompt it about QM and relativity with some added context. If the model comes up with anything even remotely correct it would be quite a strong evidence that LLMs are a path to something bigger if not then I think it is time to go back to the drawing board.

I'm trying to work towards that goal by training a model on mostly German science texts up to 1904 (before the world wars German was the lingua franca of most sciences).

Training data for a base model isn't that hard to come by, even though you have to OCR most of it yourself because the publicly available OCRed versions are commonly unusably bad. But training a model large enough to be useful is a major issue. Training a 700M parameter model at home is very doable (and is what this TimeCapsuleLLM is), but to get that kind of reasoning you need something closer to a 70B model. Also a lot of the "smarts" of a model gets injected in fine tuning and RL, but any of the available fine tuning datasets would obviously contaminate the model with 2026 knowledge.

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#124
post #29

Earlier quoted context omitted.

It doesn’t need to know about QM or reactivity just about the building blocks that led to them. Which were more than around in the year 1900. In fact you don’t want it to know about them explicitly just have enough background knowledge that you can manage the rest via context.

I was vague. My point is that I don't think the building blocks are in the data. Its mainly tertiary and popular sources. Maybe if you had the writings of Victorian scientists, both public and private correspondence.

Probably a lot of it exists but in archives, private collections etc. Would be great if it will all end up digitized as well.

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#125

Earlier quoted context omitted.

But that's not the OP's challenge, he said "if the model comes up with anything even remotely correct ." The point is there were things already "remotely correct" out there in 1900. If the LLM finds them, it wouldn't "be quite a strong evidence that LLMs are a path to something bigger."

It's not the comment which is illogical, it's your (mis)interpretation of it. What I (and seemingly others) took it to mean is basically could an LLM do Einstein's job ? Could it weave together all those loose threads into a coherent new way of understanding the physical world? If so, AGI can't be far behind.

AGI is human level intelligence, and the minimum bar is Einstein?

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#126

Earlier quoted context omitted.

Yeah but... we still might not know if it could do that because we were really close by 1900 or if the LLM is very smart.

What's the bar here? Does anyone say "we don't know if Einstein could do this because we were really close or because he was really smart?" I by no means believe LLMs are general intelligence, and I've seen them produce a lot of garbage, but if they could produce these revolutionary theories from only <= year 1900 information and a prompt that is not ridiculously leading, that would be a really compelling demonstrati…

> Does anyone say "we don't know if Einstein could do this because we were really close or because he was really smart?"

It turns out my reading is somewhat topical. I've been reading Rhodes' "The Making of the Atomic Bomb" and of the things he takes great pains to argue (I was not quite anticipating how much I'd be trying to recall my high school science classes to make sense of his account of various experiments) is that the development toward the atomic bomb was more or less inexorable and if at any point someone said "this is too far; let's stop here" there would be others to take his place. So, maybe, to answer your question.

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#127

Would be interesting to train a cutting edge model with a cut off date of say 1900 and then prompt it about QM and relativity with some added context. If the model comes up with anything even remotely correct it would be quite a strong evidence that LLMs are a path to something bigger if not then I think it is time to go back to the drawing board.

Yann LeCun spoke explicitly on this idea recently and he asserts definitively that the LLM would not be able to add anything useful in that scenario. My understanding is that other AI researchers generally agree with him, and that it's mostly the hype beasts like Altman that think there is some "magic" in the weights that is actually intelligent. Their payday depends on it, so it is understandable. My opinion is that…

Preface: Most of my understand of how LLMs actually work comes from 3blue1brown's videos, so I could easily be wrong here.

I mostly agree with you, especially about distrusting the self-interested hype beasts.

While I don't think the models are actually "intelligent", I also wonder if there are insights to be gained by looking at how concepts get encoded by the models. It's not really that the models will add something "new", but more that there might be connections between things that we haven't noticed, especially because academic disciplines are so insular these days.

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#128

Earlier quoted context omitted.

Yann LeCun spoke explicitly on this idea recently and he asserts definitively that the LLM would not be able to add anything useful in that scenario. My understanding is that other AI researchers generally agree with him, and that it's mostly the hype beasts like Altman that think there is some "magic" in the weights that is actually intelligent. Their payday depends on it, so it is understandable. My opinion is that…

This is definitely wrong, most AI researchers DO NOT agree with LeCun. Most ML researchers think AGI is imminent.

Well, can you point us to their research then? Please.

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#129

Would be interesting to train a cutting edge model with a cut off date of say 1900 and then prompt it about QM and relativity with some added context. If the model comes up with anything even remotely correct it would be quite a strong evidence that LLMs are a path to something bigger if not then I think it is time to go back to the drawing board.

Yann LeCun spoke explicitly on this idea recently and he asserts definitively that the LLM would not be able to add anything useful in that scenario. My understanding is that other AI researchers generally agree with him, and that it's mostly the hype beasts like Altman that think there is some "magic" in the weights that is actually intelligent. Their payday depends on it, so it is understandable. My opinion is that…

How about this for an evaluation: Have this (trained-on-older-corpus) LLM propose experiments. We "play the role of nature" and inform it of the results of the experiments. It can then try to deduce the natural laws.

If we did this (to a good enough level of detail), would it be able to derive relativity? How large of an AI model would it have to be to successfully derive relativity (if it only had access to everything published up to 1904)?

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#130
post #88
post #81

Mm. I'm a bit sceptical of the historical expertise of someone who thinks that "Who art Henry" is 19th century language. (It's not actually grammatically correct English from any century whatever: "art" is the second person singular, so this is like saying "who are Henry?")

Can you elaborate on this? After skimming the README, I understand that "Who art Henry" is the prompt. What should be the correct 19th century prompt?

Who art thou?

(Well, not 19th century...)

Post reply on HN