Live data from Hacker News

TimeCapsuleLLM: LLM trained only on data from 1800-1875

github.com

61–70 of 334 posts

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#61
post #22

This kind of technique seems like a good way to test model performance against benchmarks. I'm too skeptical that new models are taking popular benchmark solutions into their training data. So-- how does e.g. ChatGPT's underlying architecture perform on SWE-bench if trained only on data prior to 2024.

> are taking popular benchmark solutions into their training data

That happened in the past, and the "naive" way of doing it is usually easy to spot. There are, however, many ways in which testing data can leak into models, even without data contamination. However this doesn't matter much, as any model that only does well in benchmarks but is bad in real-world usage will be quickly sussed out by people actually using them. There are also lots and lots of weird, not very popular benchmarks out there, and the outliers are quickly identified.

> perform on SWE-bench if trained only on data prior to 2024.

There's a benchmark called swe-REbench, that takes issues from real-world repos, published ~ monthly. They perform tests and you can select the period and check their performance. This is fool-proof for open models, but a bit unknown for API-based models.

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#62
It's interesting that it's trained off only historic text.

Back in the pre-LLM days, someone trained a Markov chain off the King James Bible and a programming book: https://www.tumblr.com/kingjamesprogramming

I'd love to see an LLM equivalent, but I don't think that's enough data to train from scratch. Could a LoRA or similar be used in a way to get speech style to strictly follow a few megabytes worth of training data?

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#63

Would be interesting to train a cutting edge model with a cut off date of say 1900 and then prompt it about QM and relativity with some added context. If the model comes up with anything even remotely correct it would be quite a strong evidence that LLMs are a path to something bigger if not then I think it is time to go back to the drawing board.

Yann LeCun spoke explicitly on this idea recently and he asserts definitively that the LLM would not be able to add anything useful in that scenario. My understanding is that other AI researchers generally agree with him, and that it's mostly the hype beasts like Altman that think there is some "magic" in the weights that is actually intelligent. Their payday depends on it, so it is understandable. My opinion is that…

There is some ability for it to make novel connections but it's pretty small. You can see this yourself having it build novel systems.

It largely cannot imaginr anything beyond the usual but there is a small part that it can. This is similar to in context learning, it's weak but it is there.

It would be incredible if meta learning/continual learning found a way to train exactly for novel learning path. But that's literally AGI so maybe 20yrs from now? Or never..

You can see this on CL benchmarks. There is SOME signal but it's crazy low. When I was traing CL models i found that signal was in the single % points. Some could easily argue it was zero but I really do believe there is a very small amount in there.

This is also why any novel work or findings is done via MASSIVE compute budgets. They find RL enviroments that can extract that small amount out. Is it random chance? Maybe, hard to say.

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#64
Fascinating idea. There was another "time-locked" LLM project that popped up on HN recently[1]. Their model output is really polished but the team is trying to figure out how to avoid abuse and misrepresentation of their goals. We think it would be cool to talk to someone from 100+ years ago but haven't seriously considered the many ways in which it would be uncool. Interesting times!

[1] https://news.ycombinator.com/item?id=46319826

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#65

Earlier quoted context omitted.

> I would have a chance to get a rudimentary insight on what the world was like at that time Congratulations, you've reinvented the history book (just with more energy consumption and less guarantee of accuracy)

History books, especially those from classical antiquity, are notoriously not guaranteed to be accurate either.

Do you expect something exclusively trained on them to be any better?

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#66

Would be interesting to train a cutting edge model with a cut off date of say 1900 and then prompt it about QM and relativity with some added context. If the model comes up with anything even remotely correct it would be quite a strong evidence that LLMs are a path to something bigger if not then I think it is time to go back to the drawing board.

Chemistry would be a great space to explore. The last quarter of the 19th century had a ton of advancements in chemistry. It'd be interesting the see if an LLM could propose fruitful hypotheses, made predictions of the science of thermodynamics.

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#67
post #14

Could this be an experiment to show how likely LLMs are to lead to AGI, or at least intelligence well beyond our current level? If you could only give it texts and info and concepts up to Year X, well before Discovery Y, could we then see if it could prompt its way to that discovery?

It'd be difficult to prove that you hadn't leaked information to the model. The big gotcha of LLMs is that you train them on BIG corpuses of data, which means it's hard to say "X isn't in this corpus", or "this corpus only contains Y". You could TRY to assemble a set of training data that only contains text from before a certain date, but it'd be tricky as heck to be SURE about it.

Ways data might leak to the model that come to mind: misfiled/mislabled documents, footnotes, annotations, document metadata.

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#68
post #28
post #14

Could this be an experiment to show how likely LLMs are to lead to AGI, or at least intelligence well beyond our current level? If you could only give it texts and info and concepts up to Year X, well before Discovery Y, could we then see if it could prompt its way to that discovery?

> Could this be an experiment to show how likely LLMs are to lead to AGI, or at least intelligence well beyond our current level? You'd have to be specific what you mean by AGI: all three letters mean a different thing to different people, and sometimes use the whole means something not present in the letters. > If you could only give it texts and info and concepts up to Year X, well before Discovery Y, could we then…

> You'd have to be specific what you mean by AGI

Well, they obviously can't. AGI is not science, it's religion. It has all the trappings of religion: prophets, sacred texts, origin myth, end-of-days myth and most importantly, a means to escape death. Science? Well, the only measure to "general intelligence" would be to compare to the only one which is the human one but we have absolutely no means by which to describe it. We do not know where to start. This is why you scrape the surface of any AGI definition you only find circular definitions.

And no, the "brain is a computer" is not a scientific description, it's a metaphor.

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#69

Earlier quoted context omitted.

History books, especially those from classical antiquity, are notoriously not guaranteed to be accurate either.

Do you expect something exclusively trained on them to be any better?

To a large extent, yes. A model trained on many different accounts of an event is likely going to give a more faithful picture of that event than any one author.

This isn't super relevant to us because very few histories from this era survived, but presumably there was sufficient material in the Library of Alexandria to cover events from multiple angles and "zero out" the different personal/political/religious biases coloring the individual accounts.

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#70
Suppose two models with similar parameters trained the same way on 1800-1875 and 1800-2025 data. Running both models, we get probability distributions across tokens, let's call the distributions 1875' and 2025'. We also get a probability distribution finite difference (2025' - 1875'). What would we get if we sampled from 1.1*(2025' - 1875') + 1875'? I don't think this would actually be a decent approximation of 2040', but it would be a fun experiment to see. (Interpolation rather than extrapolation seems just as unlikely to be useful and less likely to be amusing, but what do I know.)
Post reply on HN