Live data from Hacker News

TimeCapsuleLLM: LLM trained only on data from 1800-1875

github.com

171–180 of 334 posts

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#171

If the output of this is even somewhat coherent, it would disprove the argument that mass amounts of copyrighted works are required to train an LLM. Unfortunately that does not appear to be the case here.

Take a look at The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text ( https://arxiv.org/pdf/2506.05209 ). They build a reasonable 7B parameter model using only open-licensed data.

They mostly do that. They risked legal contamination by using Whisper-derived text and web text which might have gotchas. Other than that, it was a great collection for low-risk training.

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#172
post #84

Earlier quoted context omitted.

You realize parent said "This would be an interesting way to test proposition X" and you responded with "X is false because I say say", right?

"Proposition X" does not need testing. We already know X is categorically false because we know how LLMs are programmed, and not a single line of that programming pertains to thinking (thinking in the human sense, not "thinking" in the LLM sense which merely uses an anthromorphized analogy to describe a script that feeds back multiple prompts before getting the final prompt output to present to the user). In the same…

> We already know X is categorically false because we know how LLMs are programmed, and not a single line of that programming pertains to thinking (thinking in the human sense, not "thinking" in the LLM sense which merely uses an anthromorphized analogy to describe a script that feeds back multiple prompts before getting the final prompt output to present to the user).

Could you elucidate me on the process of human thought, and point out the differences between that and a probabilistic prediction engine?

I see this argument all over the place, but "how do humans think" is never described. It is always left as a black box with something magical (presumably a soul or some other metaphysical substance) inside.

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#173

Earlier quoted context omitted.

I've seen many futurists claim that human innovation is dead and all future discoveries will be the results of AI. If this is true, we should be able to see AI trained on the past figure it's way to various things we have today. If it can't do this, I'd like said futurists to quiet down, as they are discouraging an entire generation of kids who may go on to discover some great things.

I'm struggling to get a handle on this idea. Is the idea that today's data will be the data of the past, in the future? So if it can work with whats now past, it will be able to work with the past in the future?

Essentially, yes.

If the prediction is that AI will be able to invent the future. If we give it data from our past without knowledge of the present... what type of future will it invent, what progress will it make, if any at all? And not just having the idea, but how to implement the idea in a way that actually works with the technology of the day, and can build on those things over time.

For example, would AI with 1850 data have figured out the idea of lift to make an airplane and taught us how to make working flying machines and progress them to the jets we have today, or something better? It wouldn't even be starting from 0, so this would be a generous example, as da Vinci way playing with these ideas in the 15th century.

If it can't do it, or what it produces is worse than what humans have done, we shouldn't leave it to AI alone to invent our actual future. Which would mean reevaluating the role these "thought leaders" say it will play, and how we're educating and communicating about AI to the younger generations.

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#174

Earlier quoted context omitted.

I am a deep LLM skeptic. But I think there are also some questions about the role of language in human thought that leave the door just slightly ajar on the issue of whether or not manipulating the tokens of language might be more central to human cognition than we've tended to think. If it turned out that this was true, then it is possible that "a model predicting tokens" has more power than that description would s…

> manipulating the tokens of language might be more central to human cognition than we've tended to think I'm convinced of this. I think it's because we've always looked at the most advanced forms of human languaging (like philosophy) to understand ourselves. But human language must have evolved from forms of communication found in other species, especially highly intelligent ones. It's to be expected that the buildi…

>Ironically, the other crucial ingredient for AGI which LLMs don't have, but we do, is exactly that animal nature which we always try to shove under the rug, over-attributing our success to the stochastic parrot part of us, and ignoring the gut instinct, the intuitive, spontaneous insight into things which a lot of the great scientists and artists of the past have talked about.

Are you familiar with the major works in epistemology that were written, even before the 20th century, on this exact topic?

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#175
post #164
post #84

Earlier quoted context omitted.

You realize parent said "This would be an interesting way to test proposition X" and you responded with "X is false because I say say", right?

Yes. That is correct. If I told you I planned on going outside this evening to test whether the sun sets in the east, the best response would be to let me know ahead of time that my hypothesis is wrong.

So, based on the source of "Trust me bro.", we'll decide this open question about new technology and the nature of cognition is solved. Seems unproductive.

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#176

Earlier quoted context omitted.

"Proposition X" does not need testing. We already know X is categorically false because we know how LLMs are programmed, and not a single line of that programming pertains to thinking (thinking in the human sense, not "thinking" in the LLM sense which merely uses an anthromorphized analogy to describe a script that feeds back multiple prompts before getting the final prompt output to present to the user). In the same…

> We already know X is categorically false because we know how LLMs are programmed, and not a single line of that programming pertains to thinking (thinking in the human sense, not "thinking" in the LLM sense which merely uses an anthromorphized analogy to describe a script that feeds back multiple prompts before getting the final prompt output to present to the user). Could you elucidate me on the process of human t…

There is no need to involve souls or magic. I am not making the argument that it is impossible to create a machine that is capable of doing the same computations as the brain. The argument is that whether or not such a machine is possible, an LLM is not such a machine. If you'd like to think of our brains as squishy computers, then the principle is simple: we run code that is more complex than a token prediction engine. The fact that our code is more complex than a token prediction engine is easily verified by our capability to address problems that a token prediction engine cannot. This is because our brain-code is capable of reasoning from deterministic logical principles rather than only probabilities. We also likely have something akin to token prediction code, but that is not the only thing our brain is programmed to do, whereas it is the only thing LLMs are programmed to do.

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#177

Would be interesting to train a cutting edge model with a cut off date of say 1900 and then prompt it about QM and relativity with some added context. If the model comes up with anything even remotely correct it would be quite a strong evidence that LLMs are a path to something bigger if not then I think it is time to go back to the drawing board.

I like this, it would be exciting (and scary) if it deduced QM, and informative if it cannot.

But I also think we can do this with normal LLMs trained on up-to-date text, by asking them to come up with any novel theory that fits the facts. It does not have to be a groundbreaking theory like QM, just original and not (yet) proven wrong ?

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#178

Earlier quoted context omitted.

"Proposition X" does not need testing. We already know X is categorically false because we know how LLMs are programmed, and not a single line of that programming pertains to thinking (thinking in the human sense, not "thinking" in the LLM sense which merely uses an anthromorphized analogy to describe a script that feeds back multiple prompts before getting the final prompt output to present to the user). In the same…

> We already know X is categorically false because we know how LLMs are programmed, and not a single line of that programming pertains to thinking (thinking in the human sense, not "thinking" in the LLM sense which merely uses an anthromorphized analogy to describe a script that feeds back multiple prompts before getting the final prompt output to present to the user). Could you elucidate me on the process of human t…

Kant's model of epistemology, with humans schematizing conceptual understanding of objects through apperception of manifold impressions from our sensibility, and then reasoning about these objects using transcendental application of the categories, is a reasonable enough model of thought. It was (and is I think) a satisfactory answer for the question of how humans can produce synthetic a priori knowledge, something that LLMs are incapable of (don't take my word on that though, ChatGPT is more than happy to discuss [1])

1: https://chatgpt.com/share/6965653e-b514-8011-b233-79d8c25d33...

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#179

Would be interesting to train a cutting edge model with a cut off date of say 1900 and then prompt it about QM and relativity with some added context. If the model comes up with anything even remotely correct it would be quite a strong evidence that LLMs are a path to something bigger if not then I think it is time to go back to the drawing board.

It's going to be divining tea leaves. It will be 99% wrong and then someone will say 'oh but look at this tea leaf over here! It's almost correct"'

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#180
post #86

Earlier quoted context omitted.

> And no, the "brain is a computer" is not a scientific description, it's a metaphor. Disagree. A brain is turing complete, no? Isn't that the definition of a computer? Sure, it may be reductive to say "the brain is just a computer".

Not even close. Turing complete does not apply to the brain plain and simple. That's something to do with algorithms and your brain is not a computer as I have mentioned. It does not store information. It doesn't process information. It just doesn't work that way. https://aeon.co/essays/your-brain-does-not-process-informati...

That is an article by a psychologist, with no expertise in neuroscience, claiming without evidence that the "dominant cognitive neuroscience" is wrong. He offers no alternative explanation on how memories are stored and retrieved, but argues that large numbers of neurons across the brain are involved and he implies that neuroscientists think otherwise.

This is odd because the dominant view in neuroscience is that memories are stored by altering synaptic connection strength in a large number of neurons. So it's not clear what his disagreement is, and he just seems to be misrepresenting neuroscientists.

Interestingly, this is also how LLMs store memory during training: by altering the strength of connections between many artificial neurons.

Post reply on HN