Live data from Hacker News

History LLMs: Models trained exclusively on pre-1913 texts

github.com

281–290 of 452 posts

Re: History LLMs: Models trained exclusively on pre-1913 texts

#281
post #269

Earlier quoted context omitted.

And Jules verne predicted rockets. I still move that it's quantitative predictions not qualitative. I mean, all Kindle does for me is save me space. I don't have to store all those books now. Who predicted the humble internet forum though? Or usenet before it?

Kindles are just books and books are already mostly fairly compact and inexpensive long-form entertainment and information. They're convenient but if they went away tomorrow, my life wouldn't really change in any material way. That's not really the case with smartphones much less the internet more broadly.

That has to be the most dystopian-sci-fi-turning-into-reality-fast thing I've read in a while.

I'd take smartphones vanishing rather than books any day.

Re: History LLMs: Models trained exclusively on pre-1913 texts

#282

I wouldn't have expected there to be enough text from before 1913 to properly train a model, it seemed like they needed an internet of text to train the first successful LLMs?

This model is more comparable to GPT-2 than anything we use now.

Re: History LLMs: Models trained exclusively on pre-1913 texts

#283

> Why not just prompt GPT-5 to "roleplay" 1913? Because it will perform token completion driven by weights coming from training data newer than 1913 with no way to turn that off. It can't be asked to pretend that it wasn't trained on documents that didn't exist in 1913. The LLM cannot reprogram its own weights to remove the influence of selected materials; that kind of introspection is not there. Not to mention that…

Excuse me sir you forgot to anthropomorphise the language model

Re: History LLMs: Models trained exclusively on pre-1913 texts

#284
post #268

Earlier quoted context omitted.

To be honest, while I'd heard of it over a decade ago and I've read LOTR and I've been paying attention to privacy longer than most, I didn't ever really look into what it did until I started hearing more about it in the past year or two. But yeah lots of people don't really buy into the idea of their small contribution to a large problem being a problem.

>But yeah lots of people don't really buy into the idea of their small contribution to a large problem being a problem. As an abstract idea I think there is a reasonable argument to be made that the size of any contribution to a problem should be measured as a relative proportion of total influence. The carbon footprint is a good example, if each individual focuses on reducing their small individual contribution then…

>As an abstract idea I think there is a reasonable argument to be made that the size of any contribution to a problem should be measured as a relative proportion of total influence.

Right, I think you have responsibility for your 1/th (arguably considerably more though, for first-worlders) of the problem. What I see is something like refusal to consider swapping out a two-stroke-engine-powered tungsten lightbulb with an LED of equivalent brightness, CRI, and color temperature, because it won't unilaterally solve the problem.

Re: History LLMs: Models trained exclusively on pre-1913 texts

#285
post #281
post #269

Earlier quoted context omitted.

Kindles are just books and books are already mostly fairly compact and inexpensive long-form entertainment and information. They're convenient but if they went away tomorrow, my life wouldn't really change in any material way. That's not really the case with smartphones much less the internet more broadly.

That has to be the most dystopian-sci-fi-turning-into-reality-fast thing I've read in a while. I'd take smartphones vanishing rather than books any day.

My point was Kindles vanishing, not books vanishing. Kindles are in no way a prerequisite for reading books.

Re: History LLMs: Models trained exclusively on pre-1913 texts

#286
post #274

Earlier quoted context omitted.

> It’s as if every researcher in this field is getting high on the small amount of power they have from denying others access to their results. I’ve never been as unimpressed by scientists as I have been in the past five years or so. This is absolutely nothing new. With experimental things, it's non uncommon for a lab to develop a new technique and omit slight but important details to give them a competitive advantag…

> With experimental things, it's non uncommon for a lab to develop a new technique and omit slight but important details to give them a competitive advantage. Yes, to give them a competitive advantage. Not to LARP as morality police. There’s a big difference between the two. I take greed over self-righteousness any day.

I’ve heard people say that they’re not going to release their software because people wouldn’t know how to use it! I’m not sure the motivation really matters more than the end result though.

Re: History LLMs: Models trained exclusively on pre-1913 texts

#287
post #97

Wait so what does the model think that it is? If it doesn't know computers exist yet, I mean, and you ask it how it works, what does it say?

This is an anthropomorphization. LLMs do not think they are anything, no concept of self, no thinking at all (despite the lovely marketing around thinking/reasoning models). I'm quite sad that more hasn't been done to dispel this.

When you ask gpt 4.1 et c to describe itself, it doesn't have singular concept of "itself". It has some training data around what LLMs are in general and can feed back a reasonable response given.

Re: History LLMs: Models trained exclusively on pre-1913 texts

#288
post #285
post #281

Earlier quoted context omitted.

That has to be the most dystopian-sci-fi-turning-into-reality-fast thing I've read in a while. I'd take smartphones vanishing rather than books any day.

My point was Kindles vanishing, not books vanishing. Kindles are in no way a prerequisite for reading books.

You may want to make your original post more clear, because i agree that at a quick glance it says you wouldn't miss books.

I didn't believe you meant that of course, but we've already seen it can happen.

Re: History LLMs: Models trained exclusively on pre-1913 texts

#289

> Historical texts contain racism, antisemitism, misogyny, imperialist views. The models will reproduce these views because they're in the training data. This isn't a flaw, but a crucial feature—understanding how such views were articulated and normalized is crucial to understanding how they took hold. Yes! > We're developing a responsible access framework that makes models available to researchers for scholarly purp…

fully understand you. we'd like to provide access but also guard against misrepresentations of our projects goals by pointing to e.g. racist generations. if you have thoughts on how we should do that, perhaps you could reach out at history-llms@econ.uzh.ch ? thanks in advance!

Re: History LLMs: Models trained exclusively on pre-1913 texts

#290

It would be interesting to see how hard it would be to walk these models towards general relativity and quantum mechanics. Einstein’s paper “On the Electrodynamics of Moving Bodies” with special relativity was published in 1905. His work on general relativity was published 10 years later in 1915. The earliest knowledge cuttoff of these models is 1913, in between the relativity papers. The knowledge cutoffs are also r…

the issue is there is very little text before the internet, so not enough historical tokens to train a really big model

And it's a 4B model. I worry that nontechnical users will dramatically overestimate its accuracy and underestimate hallucinations, which makes me wonder how it could really be useful for academic research.
Post reply on HN