Live data from Hacker News

History LLMs: Models trained exclusively on pre-1913 texts

github.com

31–40 of 452 posts

Re: History LLMs: Models trained exclusively on pre-1913 texts

#31
post #12

So many disclaimers about bias. I wonder how far back you have to go before the bias isn’t an issue. Not because it unbiased, but because we don’t recognize or care about the biases present.

Was there ever such a time or place?

There is a modern trope of a certain political group that bias is a modern invention of another political group - an attempt to politicize anti-bias.

Preventing bias is fundamental to scientific research and law, for example. That same political group is strongly anti-science and anti-rule-of-law, maybe for the same reason.

Re: History LLMs: Models trained exclusively on pre-1913 texts

#32

The knowledge machine question is fascinating ("Imagine you had access to a machine embodying all the collective knowledge of your ancestors. What would you ask it?") – it truly does not know about computers, has no concept of its own substrate. But a knowledge machine is still comprehensible to it. It makes me think of the Book Of Ember, the possibility of chopping things out very deliberately. Maybe creating someth…

Jonathan Swift wrote about something we might consider a computer in the early 18th century, in Gulliver's Travels - https://en.wikipedia.org/wiki/The_Engine

The idea of knowledge machines was not necessarily common, but it was by no means unheard of by the mid 18th century, there were adding machines and other mechanical computation, even leaving aside our field's direct antecedents in Babbage and Lovelace.

Re: History LLMs: Models trained exclusively on pre-1913 texts

#33
post #28

> Imagine you could interview thousands of educated individuals from 1913—readers of newspapers, novels, and political treatises—about their views on peace, progress, gender roles, or empire. I don't mind the experimentation. I'm curious about where someone has found an application of it. What is the value of such a broad, generic viewpoint? What does it represent? What is it evidence of? The answer to both seems to…

It doesn't have to be generic. You can assign genders, ideals, even modern ones, and it should do it's best to oblige.

Re: History LLMs: Models trained exclusively on pre-1913 texts

#34
post #13
post #3

“Time-locked models don't roleplay; they embody their training data. Ranke-4B-1913 doesn't know about WWI because WWI hasn't happened in its textual universe. It can be surprised by your questions in ways modern LLMs cannot.” “Modern LLMs suffer from hindsight contamination. GPT-5 knows how the story ends—WWI, the League's failure, the Spanish flu.” This is really fascinating. As someone who reads a lot of history an…

"...what do you mean, 'World War One ?'"

… what do you mean, an internet where everything wasn't hidden behind anti-bot captchas?

Re: History LLMs: Models trained exclusively on pre-1913 texts

#35
post #8

I’d like to know how they chat-tuned it. Getting the base model is one thing, did they also make a bunch of conversations for SFT and if so how was it done? We develop chatbots while minimizing interference with the normative judgments acquired during pretraining (“uncontaminated bootstrapping”). So they are chat tuning, I wonder what “minimizing interference with normative judgements” really amounts to and how objec…

They have some more details at https://github.com/DGoettlich/history-llms/blob/main/ranke-4... Basically using GPT-5 and being careful

Thank you that helps to inject a lot of skepticism. I was wondering how it so easily worked out what Q: A: stood for when that formatting took off in the 1940s

Re: History LLMs: Models trained exclusively on pre-1913 texts

#36
post #13
post #3

“Time-locked models don't roleplay; they embody their training data. Ranke-4B-1913 doesn't know about WWI because WWI hasn't happened in its textual universe. It can be surprised by your questions in ways modern LLMs cannot.” “Modern LLMs suffer from hindsight contamination. GPT-5 knows how the story ends—WWI, the League's failure, the Spanish flu.” This is really fascinating. As someone who reads a lot of history an…

"...what do you mean, 'World War One ?'"

> "...what do you mean, 'World War One?'"

Oh sorry, spoilers.

(Hell, I miss Capaldi)

Re: History LLMs: Models trained exclusively on pre-1913 texts

#37
post #4

The sample responses given are fascinating. It seems more difficult than normal to even tell that they were generated by an LLM, since most of us (terminally online) people have been training our brains' AI-generated text detection on output from models trained with a recent cutoff date. Some of the sample responses seem so unlike anything an LLM would say, obviously due to its apparent beliefs on certain concepts, t…

I used to teach 19th-century history, and the responses definitely sound like a Victorian-era writer. And they of course sound like writing (books and periodicals etc) rather than "chat": as other responders allude to, the fine-tuning or RL process for making them good at conversation was presumably quite different from what is used for most chatbots, and they're leaning very heavily into the pre-training texts. We d…

I wonder if the historical format you might want to look at for "Chat" is letters? Definitely wordier segments, but it's at least the back and forth feel and we often have complete correspondence over long stretches from certain figures.

This would probably get easier towards the start of the 20th century ofc

Re: History LLMs: Models trained exclusively on pre-1913 texts

#39

Earlier quoted context omitted.

I used to teach 19th-century history, and the responses definitely sound like a Victorian-era writer. And they of course sound like writing (books and periodicals etc) rather than "chat": as other responders allude to, the fine-tuning or RL process for making them good at conversation was presumably quite different from what is used for most chatbots, and they're leaning very heavily into the pre-training texts. We d…

I wonder if the historical format you might want to look at for "Chat" is letters? Definitely wordier segments, but it's at least the back and forth feel and we often have complete correspondence over long stretches from certain figures. This would probably get easier towards the start of the 20th century ofc

Good point, informal letters might actually be a better source - AI chat is (usually) a written rather than spoken interaction after all! And we do have a lot transcribed collections of letters to train on, although they’re mostly from people who were famous or became famous, which certainly introduces some bias.

Re: History LLMs: Models trained exclusively on pre-1913 texts

#40
I can imagine the political and judicial battles already, like with textualist feeling that the constitution should be understood as the text and only the text, meant by specific words and legal formulations of their known meaning at the time.

“The model clearly shows that Alexander Hamilton & Monroe were much more in agreement on topic X, putting the common textualist interpretation of it and Supreme Court rulings on a now specious interpretation null and void!”

Post reply on HN