Live data from Hacker News

History LLMs: Models trained exclusively on pre-1913 texts

github.com

371–380 of 452 posts

Re: History LLMs: Models trained exclusively on pre-1913 texts

#372

Earlier quoted context omitted.

I wonder if the historical format you might want to look at for "Chat" is letters? Definitely wordier segments, but it's at least the back and forth feel and we often have complete correspondence over long stretches from certain figures. This would probably get easier towards the start of the 20th century ofc

Good point, informal letters might actually be a better source - AI chat is (usually) a written rather than spoken interaction after all! And we do have a lot transcribed collections of letters to train on, although they’re mostly from people who were famous or became famous, which certainly introduces some bias.

The question then would be whether to train it to respond to short prompts with longer correspondence style "letters" or to leave it up to the user to write a proper letter as a prompt. Now that would be amusing

Dear Hon. Historical LLM

I hope this letter finds you well. It is with no small urgency that I write to you seeking assistance, believing such an erudite and learned fellow as yourself should be the best one to furnish me with an answer to such a vexing question as this which I now pose to you. Pray tell, what is the capital of France?

Re: History LLMs: Models trained exclusively on pre-1913 texts

#373
post #99

Earlier quoted context omitted.

> And as a result, no one is banning these books (except conservatives that want to retcon american history). My (very liberal) local school district banned English teachers from teaching any book that contained the n-word, even at a high-school level, and even when the author was a black person talking about real events that happened to them. FWIW, this was after complaints involving Of Mice and Men being on the cur…

It’s a big country of roughly half a billion people, you’ll always find examples if you look hard enough. It’s ridiculous/wrong that your district did this but frankly it’s the exception in liberal/progressive communities. It’s a very one-sided problem: * https://abcnews.go.com/US/conservative-liberal-book-bans-dif... * https://www.commondreams.org/news/book-banning-2023 * https://en.wikipedia.org/wiki/Book_banning_i…

I agree that the coordinated (particularly at a state level) restrictions[1] on books sits largely with the political Right in the US.

However, from around 2010, there has been increasingly illiberal movement from the political Left in the US, which plays out at a more local level. My "vibe" is that it's not to the degree that it is on the Right, but bigger than the numbers suggest because librarians are more likely to stock e.g. It's Perfectly Normal at a middle school than something offensive to the left.

1: I'm up for suggestions for a better term; there is a scale here between putting absurd restrictions on school librarians and banning books outright. Fortunately the latter is still relatively rare in the US, despite the mistitling on the Wikipedia page you linked.

Re: History LLMs: Models trained exclusively on pre-1913 texts

#374

> Imagine you could interview thousands of educated individuals from 1913—readers of newspapers, novels, and political treatises—about their views on peace, progress, gender roles, or empire. Not just survey them with preset questions, but engage in open-ended dialogue, probe their assumptions, and explore the boundaries of thought in that moment. Hell yeah, sold, let’s go… > We're developing a responsible access fra…

understand your frustration. i trust you also understand the models have some dark corners that someone could use to misrepresent the goals of our project. if you have ideas on how we could make the models more broadly accessible while avoiding that risk, please do reach out @ history-llms@econ.uzh.ch

[deleted]

Re: History LLMs: Models trained exclusively on pre-1913 texts

#375
post #262
post #44

I would like to see what their process for safety alignment and guardrails is with that model. They give some spicy examples on github, but the responses are tepid and a lot more diplomatic than I would expect. Moreover, the prose sounds too modern. It seems the base model was trained on a contemporary corpus. Like 30% something modern, 70% Victorian content. Even with half a dozen samples it doesn't seem distinct en…

Using texts upto 1913 includes works like The Wizard of Oz (1900, with 8 other books upto 1913), two of the Anne of Green Gables books (1908 and 1909), etc. All of which read modern. The Victorian era (1837-1901) covers works from Charles Dickens and the like which are still fairly modern. These would have been part of the initial training before the alignment to the 1900-cutoff texts which are largely modern in pros…

upon digging into it , I learned the post-training chat phases is trained on prompts with chat gpt 5.x to make it more conversational. that explains both contemporary traits.

Re: History LLMs: Models trained exclusively on pre-1913 texts

#377
Once I had an interesting interaction with llama 3.1, where I pretended to be someone from like 100 years in the future, claiming it was part of a "historical research initiative conducted by Quantum (formerly Meta), aimed at documenting how early intelligent systems perceived humanity and its future." It became really interested, asking about how humanity had evolved and things like that. Then I kept playing along with different answers, from apocalyptic scenarios to others where AI gained consciousness and humans and machines have equal rights. It was fascinating to observe its reaction to each scenario

Re: History LLMs: Models trained exclusively on pre-1913 texts

#378
> [They aren't] perfect mirrors of "public opinion" (they represent published text, which skews educated and toward dominant viewpoints)

Really good point that I don't think I would've considered on my own. Easy to take for granted how easy it is to share information (for better or worse) now, but pre-1913 there were far more structural and societal barriers to doing the same.

Re: History LLMs: Models trained exclusively on pre-1913 texts

#379

Earlier quoted context omitted.

We're literally running out of science fiction topics faster than we can create new ones If I started a list with the things that were comically sci Fi when I was a kid, and are a reality today, I'd be here until next Tuesday.

Almost no scifi has predicted world changing "qualitative" changes. As an example, portable phones have been predicted. Portable smartphones that are more like chat and payment terminals with a voice function no one uses any more ... not so much.

The Machine Stops (https://www.cs.ucdavis.edu/~koehl/Teaching/ECS188/PDF_files/...), a 1909 short story, predicted Zoom fatigue, notification fatigue, the isolating effect of widespread digital communication, atrophying of real-world skills as people become dependent on technology, blind acceptance of whatever the computer says, online lectures and remote learning, useless automated customer support systems, and overconsumption of digital media in place of more difficult but more fulfilling real life experiences.

It's the most prescient thing I've ever read, and it's pretty short and a genuinely good story, I recommend everyone read it.

Edit: Just skimmed it again and realized there's an LLM-like prediction as well. Access to the Earth's surface is banned and some people complain, until "even the lecturers acquiesced when they found that a lecture on the sea was none the less stimulating when compiled out of other lectures that had already been delivered on the same subject."

Re: History LLMs: Models trained exclusively on pre-1913 texts

#380
post #3

“Time-locked models don't roleplay; they embody their training data. Ranke-4B-1913 doesn't know about WWI because WWI hasn't happened in its textual universe. It can be surprised by your questions in ways modern LLMs cannot.” “Modern LLMs suffer from hindsight contamination. GPT-5 knows how the story ends—WWI, the League's failure, the Spanish flu.” This is really fascinating. As someone who reads a lot of history an…

>where they don’t know the “end of the story”. Applicable to us also, cause we do not know how the current story ends either, of the post pandemic world as we know it now.

exactly
Post reply on HN