Live data from Hacker News

TimeCapsuleLLM: LLM trained only on data from 1800-1875

github.com

201–210 of 334 posts

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#201

Earlier quoted context omitted.

But that's not the OP's challenge, he said "if the model comes up with anything even remotely correct ." The point is there were things already "remotely correct" out there in 1900. If the LLM finds them, it wouldn't "be quite a strong evidence that LLMs are a path to something bigger."

It's not the comment which is illogical, it's your (mis)interpretation of it. What I (and seemingly others) took it to mean is basically could an LLM do Einstein's job ? Could it weave together all those loose threads into a coherent new way of understanding the physical world? If so, AGI can't be far behind.

Einstein is not AGI, and neither the other way around.

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#202
post #170

Earlier quoted context omitted.

I am a deep LLM skeptic. But I think there are also some questions about the role of language in human thought that leave the door just slightly ajar on the issue of whether or not manipulating the tokens of language might be more central to human cognition than we've tended to think. If it turned out that this was true, then it is possible that "a model predicting tokens" has more power than that description would s…

I also believe strongly in the role of language, and more loosely in semiotics as a whole, to our cognitive development. To the extent that I think there are some meaningful ideas within the mountain of gibberish from Lacan, who was the first to really tie our conception of ourselves with our symbolic understanding of the world. Unfortunately, none of that has anything to do with what LLMs are doing. The LLM is not t…

>Unfortunately, none of that has anything to do with what LLMs are doing. The LLM is not thinking about concepts and then translating that into language. It is imitating what it looks like to read people doing so and nothing more.

'Language' is only the initial and final layers of a Large Language Model. Manipulating concepts is exactly what they do, and it's unfortunate the most obstinate seem to be the most ignorant.

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#204

Earlier quoted context omitted.

This alone still wouldn't be a clear demonstration that AGI is around the corner. It's quite possible a LLM could've done Einstein's job, if Einstein's job was truly just synthesising already available information into a coherent new whole. (I couldn't say, I don't know enough of the physics landscape of the day to claim either way.) It's still unclear whether this process could be merely continued, seeded only with…

Einstein is chosen in such contexts because he's the paradigmatic paradigm-shifter. Basically, what you're saying is: "I don't know enough history of science to confirm this incredibly high opinion on Einstein's achievements. It could just be that everyone's been wrong about him, and if I'd really get down and dirty, and learn the facts at hand, I might even prove it." Einstein is chosen to avoid exactly this kind of…

No, by saying this, I am not downplaying Einstein's sizeable achievements nor trying to imply everyone was wrong about him. His was an impressive breadth of knowledge and mathematical prowess and there's no denying this.

However, what I'm saying is not mere nitpicking either. It is precisely because of my belief in Einstein's extraordinary abilities that I find it unconvincing that an LLM being able to recombine the extant written physics-related building blocks of 1900, with its practically infinite reading speed, necessarily demonstrates comparable capabilities to Einstein.

The essence of the question is this: would Einstein, having been granted eternal youth and a neverending source of data on physical phenomena, be able to innovate forever? Would an LLM?

My position is that even if an LLM is able to synthesise special relativity given 1900 knowledge, this doesn't necessarily mean that a positive answer to the first question implies a positive answer to the second.

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#205
post #161

Earlier quoted context omitted.

This alone still wouldn't be a clear demonstration that AGI is around the corner. It's quite possible a LLM could've done Einstein's job, if Einstein's job was truly just synthesising already available information into a coherent new whole. (I couldn't say, I don't know enough of the physics landscape of the day to claim either way.) It's still unclear whether this process could be merely continued, seeded only with…

I mean, "the pieces were already there" is true of everything? Einstein was synthesizing existing math and existing data is your point right? But the whole question is whether or not something can do that synthesis! And the "anyone who read all the right papers" thing - nobody actually reads all the papers. That's the bottleneck. LLMs don't have it. They will continue to not have it. Humans will continue to not be ab…

> I mean, "the pieces were already there" is true of everything? Einstein was synthesizing existing math and existing data is your point right?

If it's true of everything, then surely having an LLM work iteratively on the pieces, along with being provided additional physical data, will lead to the discovery of everything?

If the answer is "no", then surely something is still missing.

> And the "anyone who read all the right papers" thing - nobody actually reads all the papers. That's the bottleneck. LLMs don't have it. They will continue to not have it. Humans will continue to not be able to read faster than LLMs.

I agree with this. This is a definitive advantage of LLMs.

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#206
I'm wondering in what ways is this similar/different to https://github.com/DGoettlich/history-llms?

I saw TimeCapsuleLLM a few months ago, and I'm a big fan of the concept but I feel like the execution really isn't that great. I wish you:

- Released the full, actual dataset (untokenized, why did you pretokenize the small dataset release?)

- Created a reproducible run script so I can try it out myself

- Actually did data curation to remove artifacts in your dataset

- Post-trained the model so it could have some amount of chat-ability

- Released a web demo so that we could try it out (the model is tiny! Easily can run in the web browser without a server)

I may sit down and roll a better iteration myself.

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#207

It's interesting that it's trained off only historic text. Back in the pre-LLM days, someone trained a Markov chain off the King James Bible and a programming book: https://www.tumblr.com/kingjamesprogramming I'd love to see an LLM equivalent, but I don't think that's enough data to train from scratch. Could a LoRA or similar be used in a way to get speech style to strictly follow a few megabytes worth of training da…

Yup that'd be very interesting. Notably missing from this project's list is the KJV (1611 was in use at the time.) The first random newspaper that I pulled up from a search for "london newspaper 1950" has sermon references on the front page so it seems like an important missing piece.

Somewhat missing the cutoff of 1875 is the revised NT of the KJV. Work on it started in 1870 but likely wasn't used widely before 1881.

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#208

Earlier quoted context omitted.

You would find things in there that were already close to QM and relativity. The Michelson-Morley experiment was 1887 and Lorentz transformations came along in 1889. The photoelectric effect (which Einstein explained in terms of photons in 1905) was also discovered in 1887. William Clifford (who _died_ in 1889) had notions that foreshadowed general relativity: "Riemann, and more specifically Clifford, conjectured tha…

I agree, but it's important to note that QM has no clear formulation until 2025/6, it's like 20 years more of work than SR.

2025/6?

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#209
post #86

Earlier quoted context omitted.

> You'd have to be specific what you mean by AGI Well, they obviously can't. AGI is not science, it's religion. It has all the trappings of religion: prophets, sacred texts, origin myth, end-of-days myth and most importantly, a means to escape death. Science? Well, the only measure to "general intelligence" would be to compare to the only one which is the human one but we have absolutely no means by which to describe…

> And no, the "brain is a computer" is not a scientific description, it's a metaphor. Disagree. A brain is turing complete, no? Isn't that the definition of a computer? Sure, it may be reductive to say "the brain is just a computer".

probably not actually turing complete right? for one it is not infinite so

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#210

Would be interesting to train a cutting edge model with a cut off date of say 1900 and then prompt it about QM and relativity with some added context. If the model comes up with anything even remotely correct it would be quite a strong evidence that LLMs are a path to something bigger if not then I think it is time to go back to the drawing board.

You would find things in there that were already close to QM and relativity. The Michelson-Morley experiment was 1887 and Lorentz transformations came along in 1889. The photoelectric effect (which Einstein explained in terms of photons in 1905) was also discovered in 1887. William Clifford (who _died_ in 1889) had notions that foreshadowed general relativity: "Riemann, and more specifically Clifford, conjectured tha…

With LLMs the synthesis cycles could happen at a much higher frequency. Decades condensed to weeks or days?

I imagine possible buffers on that conjecture synthesis being epxerimentation and acceptance by the scientific community. AIs can come up with new ideas every day but Nature won't publish those ideas for years.

Post reply on HN