Live data from Hacker News

TimeCapsuleLLM: LLM trained only on data from 1800-1875

github.com

221–230 of 334 posts

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#221

Earlier quoted context omitted.

You would find things in there that were already close to QM and relativity. The Michelson-Morley experiment was 1887 and Lorentz transformations came along in 1889. The photoelectric effect (which Einstein explained in terms of photons in 1905) was also discovered in 1887. William Clifford (who _died_ in 1889) had notions that foreshadowed general relativity: "Riemann, and more specifically Clifford, conjectured tha…

If (as you seem to be suggesting) relativity was effectively lying there on the table waiting for Einstein to just pick it up, how come it blindsided most, if not quite all, of the greatest minds of his generation?

That's the case with all scientific discoveries - pieces of prior work get accumulated, until it eventually becomes obvious[0] how they connect, at which point someone[1] connects the dots, making a discovery... and putting it on the table, for the cycle to repeat anew. This is, in a nutshell, the history of all scientific and technological progress. Accumulation of tiny increments.

--

[0] - To people who happen to have the right background and skill set, and are in the right place.

[1] - Almost always multiple someones, independently, within short time of each other. People usually remember only one or two because, for better or worse, history is much like patent law: first to file wins.

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#222
post #78

Earlier quoted context omitted.

LLMs are models that predict tokens. They don't think, they don't build with blocks. They would never be able to synthesize knowledge about QM.

I am a deep LLM skeptic. But I think there are also some questions about the role of language in human thought that leave the door just slightly ajar on the issue of whether or not manipulating the tokens of language might be more central to human cognition than we've tended to think. If it turned out that this was true, then it is possible that "a model predicting tokens" has more power than that description would s…

If anything, I feel that current breed of multimodal LLMs demonstrate that language is not fundamental - tokens are, or rather their mutual association in high-dimensional latent space. Language as we recognize it, sequences of characters and words, are just a special case. Multimodal models manage to turn audio, video and text into tokens in the same space - they do not route through text when consuming or generating images.

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#223

It's interesting that it's trained off only historic text. Back in the pre-LLM days, someone trained a Markov chain off the King James Bible and a programming book: https://www.tumblr.com/kingjamesprogramming I'd love to see an LLM equivalent, but I don't think that's enough data to train from scratch. Could a LoRA or similar be used in a way to get speech style to strictly follow a few megabytes worth of training da…

That was far more amusing than I thought it'd be. Now we can feed those into an AI image generator to create some "art".

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#224
post #211

Earlier quoted context omitted.

What do they (or you) have to say about the Lee Sedol AlphaGo move 78. It seems like that was "new knowledge." Are games just iterable and the real world idea space not? I am playing with these ideas a little.

AlphaGo is not an LLM

And? Do the arguments differ for LLM vs the other models?

I guess the arguments sometimes mention languages. But I feel like the core of the arguments are pretty much the same regardless?

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#225
post #170

Earlier quoted context omitted.

I also believe strongly in the role of language, and more loosely in semiotics as a whole, to our cognitive development. To the extent that I think there are some meaningful ideas within the mountain of gibberish from Lacan, who was the first to really tie our conception of ourselves with our symbolic understanding of the world. Unfortunately, none of that has anything to do with what LLMs are doing. The LLM is not t…

>Unfortunately, none of that has anything to do with what LLMs are doing. The LLM is not thinking about concepts and then translating that into language. It is imitating what it looks like to read people doing so and nothing more. 'Language' is only the initial and final layers of a Large Language Model. Manipulating concepts is exactly what they do, and it's unfortunate the most obstinate seem to be the most ignoran…

They do not manipulate concepts. There is no representation of a concept for them to manipulate.

It may, however, turn out that in doing what they do, they are effectively manipulating concepts, and this is what I was alluding to: by building the model, even though your approach was through tokenization and whatever term you want to use for the network, you end up accidentally building something that implicitly manipulates concepts. Moreover, it might turn out that we ourselves do more of this than we perhaps like to think.

Nevertheless "manipulating concepts is exactly what they do" seems almost willfully ignorant of how these systems work, unless you believe that "find the next most probable sequence of tokens of some length" is all there is to "manipulating concepts".

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#226

Would be interesting to train a cutting edge model with a cut off date of say 1900 and then prompt it about QM and relativity with some added context. If the model comes up with anything even remotely correct it would be quite a strong evidence that LLMs are a path to something bigger if not then I think it is time to go back to the drawing board.

>.If the model comes up with anything even remotely correct it would be quite a strong evidence that LLMs are a path to something bigger if not then I think it is time to go back to the drawing board.

In principle I see your point, in practice my default assumption until proven otherwise here -- is that a little something slipped through post-1900.

A much easier approach would be to just download some model, whatever model, today. Then 5 years from now, whatever interesting discoveries are found - can the model get there.

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#228

Would be interesting to train a cutting edge model with a cut off date of say 1900 and then prompt it about QM and relativity with some added context. If the model comes up with anything even remotely correct it would be quite a strong evidence that LLMs are a path to something bigger if not then I think it is time to go back to the drawing board.

Don't you need to do reinforcement learning through human feedback to get non gibberish results from the models in general?

1900 era humans are not available to do this so I'm not sure how this experiment is supposed to work.

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#229

Would be interesting to train a cutting edge model with a cut off date of say 1900 and then prompt it about QM and relativity with some added context. If the model comes up with anything even remotely correct it would be quite a strong evidence that LLMs are a path to something bigger if not then I think it is time to go back to the drawing board.

You would find things in there that were already close to QM and relativity. The Michelson-Morley experiment was 1887 and Lorentz transformations came along in 1889. The photoelectric effect (which Einstein explained in terms of photons in 1905) was also discovered in 1887. William Clifford (who _died_ in 1889) had notions that foreshadowed general relativity: "Riemann, and more specifically Clifford, conjectured tha…

It's only easy to see precursors in hindsight. The Michelson-Morley tale is a great example of this. In hindsight, their experiment was screaming relativity, because it demonstrated that the speed of light was identical from two perspectives where it's very difficult to explain without relativity. Lorentz contraction was just a completely ad-hoc proposal to maintain the assumptions of the time (luminiferous aether in particular) while also explaining the result. But in general it was not seen as that big of a deal.

There's a very similar parallel with dark matter in modern times. We certainly have endless hints to the truth that will be evident in hindsight, but for now? We are mostly convinced that we know the truth, perform experiments to prove that, find nothing, shrug, adjust the model to be even more esoteric, and repeat onto the next one. And maybe one will eventually show something, or maybe we're on the wrong path altogether. This quote, from Michelson in 1894 (more than a decade before Einstein would come along), is extremely telling of the opinion at the time:

"While it is never safe to affirm that the future of Physical Science has no marvels in store even more astonishing than those of the past, it seems probable that most of the grand underlying principles have been firmly established and that further advances are to be sought chiefly in the rigorous application of these principles to all the phenomena which come under our notice. It is here that the science of measurement shows its importance — where quantitative work is more to be desired than qualitative work. An eminent physicist remarked that the future truths of physical science are to be looked for in the sixth place of decimals." - Michelson 1894

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#230

Earlier quoted context omitted.

>Unfortunately, none of that has anything to do with what LLMs are doing. The LLM is not thinking about concepts and then translating that into language. It is imitating what it looks like to read people doing so and nothing more. 'Language' is only the initial and final layers of a Large Language Model. Manipulating concepts is exactly what they do, and it's unfortunate the most obstinate seem to be the most ignoran…

They do not manipulate concepts. There is no representation of a concept for them to manipulate. It may, however, turn out that in doing what they do, they are effectively manipulating concepts, and this is what I was alluding to: by building the model, even though your approach was through tokenization and whatever term you want to use for the network, you end up accidentally building something that implicitly manip…

>They do not manipulate concepts. There is no representation of a concept for them to manipulate.

Yes, they do. And of course there is. And there's plenty of research on the matter.

>It may, however, turn out that in doing what they do, they are effectively manipulating concepts

There is no effectively here. Text is what goes in and what comes out, but it's by no means what they manipulate internally.

>Nevertheless "manipulating concepts is exactly what they do" seems almost willfully ignorant of how these systems work, unless you believe that "find the next most probable sequence of tokens of some length" is all there is to "manipulating concepts".

"Find the next probable token" is the goal, not the process. It is what models are tasked to do yes, but it says nothing about what they do internally to achieve it.

Post reply on HN