History LLMs: Models trained exclusively on pre-1913 texts
371–380 of 452 posts
Re: History LLMs: Models trained exclusively on pre-1913 texts
#372Earlier quoted context omitted.
I wonder if the historical format you might want to look at for "Chat" is letters? Definitely wordier segments, but it's at least the back and forth feel and we often have complete correspondence over long stretches from certain figures. This would probably get easier towards the start of the 20th century ofc
Good point, informal letters might actually be a better source - AI chat is (usually) a written rather than spoken interaction after all! And we do have a lot transcribed collections of letters to train on, although they’re mostly from people who were famous or became famous, which certainly introduces some bias.
Dear Hon. Historical LLM
I hope this letter finds you well. It is with no small urgency that I write to you seeking assistance, believing such an erudite and learned fellow as yourself should be the best one to furnish me with an answer to such a vexing question as this which I now pose to you. Pray tell, what is the capital of France?
Re: History LLMs: Models trained exclusively on pre-1913 texts
#373Earlier quoted context omitted.
> And as a result, no one is banning these books (except conservatives that want to retcon american history). My (very liberal) local school district banned English teachers from teaching any book that contained the n-word, even at a high-school level, and even when the author was a black person talking about real events that happened to them. FWIW, this was after complaints involving Of Mice and Men being on the cur…
It’s a big country of roughly half a billion people, you’ll always find examples if you look hard enough. It’s ridiculous/wrong that your district did this but frankly it’s the exception in liberal/progressive communities. It’s a very one-sided problem: * https://abcnews.go.com/US/conservative-liberal-book-bans-dif... * https://www.commondreams.org/news/book-banning-2023 * https://en.wikipedia.org/wiki/Book_banning_i…
However, from around 2010, there has been increasingly illiberal movement from the political Left in the US, which plays out at a more local level. My "vibe" is that it's not to the degree that it is on the Right, but bigger than the numbers suggest because librarians are more likely to stock e.g. It's Perfectly Normal at a middle school than something offensive to the left.
1: I'm up for suggestions for a better term; there is a scale here between putting absurd restrictions on school librarians and banning books outright. Fortunately the latter is still relatively rare in the US, despite the mistitling on the Wikipedia page you linked.
Re: History LLMs: Models trained exclusively on pre-1913 texts
#374> Imagine you could interview thousands of educated individuals from 1913—readers of newspapers, novels, and political treatises—about their views on peace, progress, gender roles, or empire. Not just survey them with preset questions, but engage in open-ended dialogue, probe their assumptions, and explore the boundaries of thought in that moment. Hell yeah, sold, let’s go… > We're developing a responsible access fra…
understand your frustration. i trust you also understand the models have some dark corners that someone could use to misrepresent the goals of our project. if you have ideas on how we could make the models more broadly accessible while avoiding that risk, please do reach out @ history-llms@econ.uzh.ch
Re: History LLMs: Models trained exclusively on pre-1913 texts
#375I would like to see what their process for safety alignment and guardrails is with that model. They give some spicy examples on github, but the responses are tepid and a lot more diplomatic than I would expect. Moreover, the prose sounds too modern. It seems the base model was trained on a contemporary corpus. Like 30% something modern, 70% Victorian content. Even with half a dozen samples it doesn't seem distinct en…
Using texts upto 1913 includes works like The Wizard of Oz (1900, with 8 other books upto 1913), two of the Anne of Green Gables books (1908 and 1909), etc. All of which read modern. The Victorian era (1837-1901) covers works from Charles Dickens and the like which are still fairly modern. These would have been part of the initial training before the alignment to the 1900-cutoff texts which are largely modern in pros…
Re: History LLMs: Models trained exclusively on pre-1913 texts
#376"Give me an LLM from 1928."
etc.
Re: History LLMs: Models trained exclusively on pre-1913 texts
#377Re: History LLMs: Models trained exclusively on pre-1913 texts
#378Really good point that I don't think I would've considered on my own. Easy to take for granted how easy it is to share information (for better or worse) now, but pre-1913 there were far more structural and societal barriers to doing the same.
Re: History LLMs: Models trained exclusively on pre-1913 texts
#379Earlier quoted context omitted.
We're literally running out of science fiction topics faster than we can create new ones If I started a list with the things that were comically sci Fi when I was a kid, and are a reality today, I'd be here until next Tuesday.
Almost no scifi has predicted world changing "qualitative" changes. As an example, portable phones have been predicted. Portable smartphones that are more like chat and payment terminals with a voice function no one uses any more ... not so much.
It's the most prescient thing I've ever read, and it's pretty short and a genuinely good story, I recommend everyone read it.
Edit: Just skimmed it again and realized there's an LLM-like prediction as well. Access to the Earth's surface is banned and some people complain, until "even the lecturers acquiesced when they found that a lecture on the sea was none the less stimulating when compiled out of other lectures that had already been delivered on the same subject."
Re: History LLMs: Models trained exclusively on pre-1913 texts
#380“Time-locked models don't roleplay; they embody their training data. Ranke-4B-1913 doesn't know about WWI because WWI hasn't happened in its textual universe. It can be surprised by your questions in ways modern LLMs cannot.” “Modern LLMs suffer from hindsight contamination. GPT-5 knows how the story ends—WWI, the League's failure, the Spanish flu.” This is really fascinating. As someone who reads a lot of history an…
>where they don’t know the “end of the story”. Applicable to us also, cause we do not know how the current story ends either, of the post pandemic world as we know it now.