Suppose two models with similar parameters trained the same way on 1800-1875 and 1800-2025 data. Running both models, we get probability distributions across tokens, let's call the distributions 1875' and 2025'. We also get a probability distribution finite difference (2025' - 1875'). What would we get if we sampled from 1.1*(2025' - 1875') + 1875'? I don't think this would actually be a decent approximation of 2040'…
What if it's just genAlpha slang?
TimeCapsuleLLM: LLM trained only on data from 1800-1875
131–140 of 334 posts
Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875
#132Mm. I'm a bit sceptical of the historical expertise of someone who thinks that "Who art Henry" is 19th century language. (It's not actually grammatically correct English from any century whatever: "art" is the second person singular, so this is like saying "who are Henry?")
Can you elaborate on this? After skimming the README, I understand that "Who art Henry" is the prompt. What should be the correct 19th century prompt?
Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875
#133Would be interesting to train a cutting edge model with a cut off date of say 1900 and then prompt it about QM and relativity with some added context. If the model comes up with anything even remotely correct it would be quite a strong evidence that LLMs are a path to something bigger if not then I think it is time to go back to the drawing board.
I'm trying to work towards that goal by training a model on mostly German science texts up to 1904 (before the world wars German was the lingua franca of most sciences). Training data for a base model isn't that hard to come by, even though you have to OCR most of it yourself because the publicly available OCRed versions are commonly unusably bad. But training a model large enough to be useful is a major issue. Train…
Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875
#134Could this be an experiment to show how likely LLMs are to lead to AGI, or at least intelligence well beyond our current level? If you could only give it texts and info and concepts up to Year X, well before Discovery Y, could we then see if it could prompt its way to that discovery?
Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875
#135Earlier quoted context omitted.
> And no, the "brain is a computer" is not a scientific description, it's a metaphor. Disagree. A brain is turing complete, no? Isn't that the definition of a computer? Sure, it may be reductive to say "the brain is just a computer".
Not even close. Turing complete does not apply to the brain plain and simple. That's something to do with algorithms and your brain is not a computer as I have mentioned. It does not store information. It doesn't process information. It just doesn't work that way. https://aeon.co/essays/your-brain-does-not-process-informati...
Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875
#136Very cool concept though, but it definitely has some bias.
Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875
#137Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875
#138Earlier quoted context omitted.
LLMs are models that predict tokens. They don't think, they don't build with blocks. They would never be able to synthesize knowledge about QM.
I am a deep LLM skeptic. But I think there are also some questions about the role of language in human thought that leave the door just slightly ajar on the issue of whether or not manipulating the tokens of language might be more central to human cognition than we've tended to think. If it turned out that this was true, then it is possible that "a model predicting tokens" has more power than that description would s…
I'm convinced of this. I think it's because we've always looked at the most advanced forms of human languaging (like philosophy) to understand ourselves. But human language must have evolved from forms of communication found in other species, especially highly intelligent ones. It's to be expected that the building blocks of it is based on things like imitation, playful variation, pattern-matching, harnessing capabilities brains have been developing long before language, only now in the emerging world of sounds, calls, vocalizations.
Ironically, the other crucial ingredient for AGI which LLMs don't have, but we do, is exactly that animal nature which we always try to shove under the rug, over-attributing our success to the stochastic parrot part of us, and ignoring the gut instinct, the intuitive, spontaneous insight into things which a lot of the great scientists and artists of the past have talked about.
Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875
#139Earlier quoted context omitted.
It'd be difficult to prove that you hadn't leaked information to the model. The big gotcha of LLMs is that you train them on BIG corpuses of data, which means it's hard to say "X isn't in this corpus", or "this corpus only contains Y". You could TRY to assemble a set of training data that only contains text from before a certain date, but it'd be tricky as heck to be SURE about it. Ways data might leak to the model t…
There's also severe selection effects: what documents have been preserved, printed, and scanned because they turned out to be on the right track towards relativity?
Especially for London there is a huge chunk of recorded parliament debates.
More interesting for dialoge seems training on recorded correspondence in form of letters anyway.
And that corpus script just looks odd to say the least, just oversample by X?
Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875
#140Earlier quoted context omitted.
It's not the comment which is illogical, it's your (mis)interpretation of it. What I (and seemingly others) took it to mean is basically could an LLM do Einstein's job ? Could it weave together all those loose threads into a coherent new way of understanding the physical world? If so, AGI can't be far behind.
AGI is human level intelligence, and the minimum bar is Einstein?