Live data from Hacker News

History LLMs: Models trained exclusively on pre-1913 texts

github.com

301–310 of 452 posts

Re: History LLMs: Models trained exclusively on pre-1913 texts

#301
post #285
post #281

Earlier quoted context omitted.

That has to be the most dystopian-sci-fi-turning-into-reality-fast thing I've read in a while. I'd take smartphones vanishing rather than books any day.

My point was Kindles vanishing, not books vanishing. Kindles are in no way a prerequisite for reading books.

Thanks for clarifying, I see what you mean now.

Re: History LLMs: Models trained exclusively on pre-1913 texts

#302
post #28

> Imagine you could interview thousands of educated individuals from 1913—readers of newspapers, novels, and political treatises—about their views on peace, progress, gender roles, or empire. I don't mind the experimentation. I'm curious about where someone has found an application of it. What is the value of such a broad, generic viewpoint? What does it represent? What is it evidence of? The answer to both seems to…

I agree. This is just make believe based on a smaller subset of human writing than LLMs we have today. It's responses are in no way useful because it is a machine mimicking a subset of published works that survived to be digitized. In that sense the "opinions" and "beliefs" are just an averaging of a subset of a subset of humanity pre 1913. I see no value in this to historians. It is really more of a parlor trick, a seance masquerading as science.

Re: History LLMs: Models trained exclusively on pre-1913 texts

#303
post #184

Earlier quoted context omitted.

There's a thriving startup scene in that direction.

Wasn't that the elevator pitch for Palentir? Still can't believe people buy their stock, given that they are the closest thing to a James Bond villain, just because it goes up. I mean, they are literally called "the stuff Sauron uses to control his evil forces". It's so on the nose it reads like an anime plot.

> Still can't believe people buy their stock, given that they are the closest thing to a James Bond villain, just because it goes up.

I proudly owned zero shares of Microsoft stock, in the 1980s and 1990s. :)

I own no Palantir today.

It's a Pyrrhic victory, but sometimes that's all you can do.

Re: History LLMs: Models trained exclusively on pre-1913 texts

#304

> Historical texts contain racism, antisemitism, misogyny, imperialist views. The models will reproduce these views because they're in the training data. This isn't a flaw, but a crucial feature—understanding how such views were articulated and normalized is crucial to understanding how they took hold. Yes! > We're developing a responsible access framework that makes models available to researchers for scholarly purp…

fully understand you. we'd like to provide access but also guard against misrepresentations of our projects goals by pointing to e.g. racist generations. if you have thoughts on how we should do that, perhaps you could reach out at history-llms@econ.uzh.ch ? thanks in advance!

What is your worst-case scenario here?

Something like a pop-sci article along the lines of "Mad scientists create racist, imperialistic AI"?

I honestly don't see publication of the weights as a relevant risk factor, because sensationalist misrepresentation is trivially possible with the given example responses alone.

I don't think such pseudo-malicious misrepresentation of scientific research can be reliably prevented anyway, and the disclaimers make your stance very clear.

On the other hand, publishing weights might lead to interesting insights from others tinkering with the models. A good example for this would be the published word prevalence data (M. Brysbaert et al @Ghent University) that led to interesting follow-ups like this: https://observablehq.com/@yurivish/words

I hope you can get the models out in some form, would be a waste not to, but congratulations on a fascinating project regardless!

Re: History LLMs: Models trained exclusively on pre-1913 texts

#305
post #3

“Time-locked models don't roleplay; they embody their training data. Ranke-4B-1913 doesn't know about WWI because WWI hasn't happened in its textual universe. It can be surprised by your questions in ways modern LLMs cannot.” “Modern LLMs suffer from hindsight contamination. GPT-5 knows how the story ends—WWI, the League's failure, the Spanish flu.” This is really fascinating. As someone who reads a lot of history an…

Reminds me of this scene from a Doctor Who episode

https://youtu.be/eg4mcdhIsvU

I’m not a Doctor Who fan and haven’t seen the rest of the episode and I don’t even what this episode was about but I thought this scene was excellent.

Re: History LLMs: Models trained exclusively on pre-1913 texts

#306

> Historical texts contain racism, antisemitism, misogyny, imperialist views. The models will reproduce these views because they're in the training data. This isn't a flaw, but a crucial feature—understanding how such views were articulated and normalized is crucial to understanding how they took hold. Yes! > We're developing a responsible access framework that makes models available to researchers for scholarly purp…

fully understand you. we'd like to provide access but also guard against misrepresentations of our projects goals by pointing to e.g. racist generations. if you have thoughts on how we should do that, perhaps you could reach out at history-llms@econ.uzh.ch ? thanks in advance!

Perhaps you could detect these... "dated"... conclusions and prepend a warning to the responses? IDK.

I think the uncensored response is still valuable, with context. "Those who cannot remember the past are condemned to repeat it" sort of thing.

Re: History LLMs: Models trained exclusively on pre-1913 texts

#307

Earlier quoted context omitted.

They have some more details at https://github.com/DGoettlich/history-llms/blob/main/ranke-4... Basically using GPT-5 and being careful

Ok so it was that. The responses given did sound off, while it has some period-appropriate mannerisms, and has entire sections basically rephrased from some popular historical texts, it seems off compared to reading an actual 1900s text. The overall vibe just isn't right, it seems too modern, somehow. I also wonder that you'd get this kind of performance with actual, just pre-1900s text. LLMs work because they're fed…

what makes you think we trained on only a few gigabytes? https://github.com/DGoettlich/history-llms/blob/main/ranke-4...

Re: History LLMs: Models trained exclusively on pre-1913 texts

#308

> Imagine you could interview thousands of educated individuals from 1913—readers of newspapers, novels, and political treatises—about their views on peace, progress, gender roles, or empire. Not just survey them with preset questions, but engage in open-ended dialogue, probe their assumptions, and explore the boundaries of thought in that moment. Hell yeah, sold, let’s go… > We're developing a responsible access fra…

understand your frustration. i trust you also understand the models have some dark corners that someone could use to misrepresent the goals of our project. if you have ideas on how we could make the models more broadly accessible while avoiding that risk, please do reach out @ history-llms@econ.uzh.ch

Of course, I have to assume that you have considered more outcomes than I have. Because, from my five minutes of reflection as a software geek, albeit with a passion for history, I find this the most surprising thing about the whole project.

I suspect restricting access could equally be a comment on modern LLMs in general, rather than the historical material specifically. For example, we must be constantly reminded not to give LLMs a level of credibility that their hallucinations would have us believe.

But I'm fascinated by the possibility that somehow resurrecting lost voices might give an unholy agency to minds and their supporting worldviews that are so anachronistic that hearing them speak again might stir long-banished evils. I'm being lyrical for dramatic affect!

I would make one serious point though, that do I have the credentials to express. The conversation may have died down, but there is still a huge question mark over, if not the legality, but certainly the ethics of restricting access to, and profiting from, public domain knowledge. I don't wish to suggest a side to take here, just to point out that the lack of conversation should not be taken to mean that the matter is settled.

Re: History LLMs: Models trained exclusively on pre-1913 texts

#309

> Historical texts contain racism, antisemitism, misogyny, imperialist views. The models will reproduce these views because they're in the training data. This isn't a flaw, but a crucial feature—understanding how such views were articulated and normalized is crucial to understanding how they took hold. Yes! > We're developing a responsible access framework that makes models available to researchers for scholarly purp…

fully understand you. we'd like to provide access but also guard against misrepresentations of our projects goals by pointing to e.g. racist generations. if you have thoughts on how we should do that, perhaps you could reach out at history-llms@econ.uzh.ch ? thanks in advance!

[deleted]

Re: History LLMs: Models trained exclusively on pre-1913 texts

#310
post #229

I'm surprised you can do this with a relatively modest corpus of text (compared to the petabytes you can vacuum up from modern books, Wikipedia, and random websites). But if it works, that's actually fantastic, because it lets you answer some interesting questions about LLMs being able to make new discoveries or transcend the training set in other ways. Forget relativity: can an LLM trained on this data notice any in…

The chinchilla paper says the “optimal” training data set size is about 20x the number of parameters (in tokens), see table 3: https://arxiv.org/pdf/2203.15556 Here they do 80B tokens for a 4B model.

It's worth noting that this is "compute-bound optimal", i.e., given fixed compute, the optimal choice is 20:1.

Under Chinchilla model the larger model always performs better than the small one if trained on the same amount of data. I'm not sure if it is true empirically, and probably 1-10B is a good guess for how large the model trained on 80B tokens should be.

Similarly, the small models continue to improve beyond 20:1 ratio, and current models are trained on much more data. You could train a better performing model using the same compute, but it would be larger which is not always desirable.

Post reply on HN