Live data from Hacker News

TimeCapsuleLLM: LLM trained only on data from 1800-1875

github.com

141–150 of 334 posts

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#141
I think it would be very cute to train a model exclusively in pre-information age documents, and then try to teach it what a computer is and get it to write some programs. That said, this doesn't look like it's nearly there yet, with the output looking closer to Markov chain than ChatGPT quality.

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#142

Very interesting but the slight issue I see here is one of data: the information that is recorded and in the training data here is heavily skewed to those intelligent/recognized enough to have recorded it and had it preserved - much less than the current status quo of "everyone can trivially document their thoughts and life" diorama of information we have today to train LLMs on. I suspect that a frontier model today…

Biases exposed through artificial constraints help to make visible the hidden/obscured/forgotten biases of state-of-the-art systems.

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#143
post #103

Earlier quoted context omitted.

What's the bar here? Does anyone say "we don't know if Einstein could do this because we were really close or because he was really smart?" I by no means believe LLMs are general intelligence, and I've seen them produce a lot of garbage, but if they could produce these revolutionary theories from only <= year 1900 information and a prompt that is not ridiculously leading, that would be a really compelling demonstrati…

> Does anyone say "we don't know if Einstein could do this because we were really close or because he was really smart? Yes. It is certainly a question if Einstein is one of the smartest guy ever lived or all of his discoveries were already in the Zeitgeist, and would have been discovered by someone else in ~5 years.

Both can be true?

Einstein was smart and put several disjointed things together. It's amazing that one person could do so much, from explaining the Brownian motion to explaining the photoeffect.

But I think that all these would have happened within _years_ anyway.

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#144

Earlier quoted context omitted.

It's not the comment which is illogical, it's your (mis)interpretation of it. What I (and seemingly others) took it to mean is basically could an LLM do Einstein's job ? Could it weave together all those loose threads into a coherent new way of understanding the physical world? If so, AGI can't be far behind.

This alone still wouldn't be a clear demonstration that AGI is around the corner. It's quite possible a LLM could've done Einstein's job, if Einstein's job was truly just synthesising already available information into a coherent new whole. (I couldn't say, I don't know enough of the physics landscape of the day to claim either way.) It's still unclear whether this process could be merely continued, seeded only with…

Einstein is chosen in such contexts because he's the paradigmatic paradigm-shifter. Basically, what you're saying is: "I don't know enough history of science to confirm this incredibly high opinion on Einstein's achievements. It could just be that everyone's been wrong about him, and if I'd really get down and dirty, and learn the facts at hand, I might even prove it." Einstein is chosen to avoid exactly this kind of nit-picking.

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#145

Would be interesting to train a cutting edge model with a cut off date of say 1900 and then prompt it about QM and relativity with some added context. If the model comes up with anything even remotely correct it would be quite a strong evidence that LLMs are a path to something bigger if not then I think it is time to go back to the drawing board.

I'm trying to work towards that goal by training a model on mostly German science texts up to 1904 (before the world wars German was the lingua franca of most sciences). Training data for a base model isn't that hard to come by, even though you have to OCR most of it yourself because the publicly available OCRed versions are commonly unusably bad. But training a model large enough to be useful is a major issue. Train…

Can we follow along with your work / results somewhere?

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#146

Earlier quoted context omitted.

But that's not the OP's challenge, he said "if the model comes up with anything even remotely correct ." The point is there were things already "remotely correct" out there in 1900. If the LLM finds them, it wouldn't "be quite a strong evidence that LLMs are a path to something bigger."

It's not the comment which is illogical, it's your (mis)interpretation of it. What I (and seemingly others) took it to mean is basically could an LLM do Einstein's job ? Could it weave together all those loose threads into a coherent new way of understanding the physical world? If so, AGI can't be far behind.

[deleted]

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#147

Very interesting but the slight issue I see here is one of data: the information that is recorded and in the training data here is heavily skewed to those intelligent/recognized enough to have recorded it and had it preserved - much less than the current status quo of "everyone can trivially document their thoughts and life" diorama of information we have today to train LLMs on. I suspect that a frontier model today…

> but it definitely has some bias.

to be frank though, I think this a better way than all people's thoughts all of the time.

I think the "crowd" of information makes the end output of an LLM worse rather than better. Specifically in our inability to know really what kind of Bias we're dealing with.

Currently to me it feels really muddy knowing how information is biased, beyond just the hallucination and factual incosistencies.

But as far as I can tell, "correctness of the content aside", sometimes frontier LLMs respond like freshman college students, other times they respond with the rigor of a mathematics PHD canidate, and sometimes like a marketing hit piece.

This dataset has a consistency which I think is actually a really useful feature. I agree that having many perspectives in the dataset is good, but as an end user being able to rely on some level of consistency with an AI model is something I really think is missing.

Maybe more succinctly I want frontier LLM's to have a known and specific response style and bias which I can rely on, because there already is a lot of noise.

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#149

Would be interesting to train a cutting edge model with a cut off date of say 1900 and then prompt it about QM and relativity with some added context. If the model comes up with anything even remotely correct it would be quite a strong evidence that LLMs are a path to something bigger if not then I think it is time to go back to the drawing board.

I think it would raise some interesting questions, but if it did yield anything noteworthy, the biggest question would be why that LLM is capable of pioneering scientific advancements and none of the modern ones are.

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#150

Earlier quoted context omitted.

This is definitely wrong, most AI researchers DO NOT agree with LeCun. Most ML researchers think AGI is imminent.

The guy who built chatgpt literally said we're 20 years away? Not sure how to interpret that as almost imminent.

> The guy who built chatgpt literally said we're 20 years away?

20 years away in 2026, still 20 years away in 2027, etc etc.

Whatever Altman's hyping, that's the translation.

Post reply on HN