Live data from Hacker News

History LLMs: Models trained exclusively on pre-1913 texts

github.com

381–390 of 452 posts

Re: History LLMs: Models trained exclusively on pre-1913 texts

#381

Earlier quoted context omitted.

understand your frustration. i trust you also understand the models have some dark corners that someone could use to misrepresent the goals of our project. if you have ideas on how we could make the models more broadly accessible while avoiding that risk, please do reach out @ history-llms@econ.uzh.ch

Ok... So as a black person should I demand that all books written before the civil rights act be destroyed? The past is messy. But it's the only way to learn anything. All an LLM does it's take a bunch of existing texts and rebundle them. Like it or not, the existing texts are still there. I understand an LLM that won't tell me how to do heart surgery. But I can't fear one that might be less enlightened on race issue…

we're on the same page.

Re: History LLMs: Models trained exclusively on pre-1913 texts

#382

Earlier quoted context omitted.

the issue is there is very little text before the internet, so not enough historical tokens to train a really big model

And it's a 4B model. I worry that nontechnical users will dramatically overestimate its accuracy and underestimate hallucinations, which makes me wonder how it could really be useful for academic research.

valid point. its more of a stepping stone towards larger models. we're figuring out what the best way to do this is before scaling up.

Re: History LLMs: Models trained exclusively on pre-1913 texts

#383
Very neat! I've thought about this with frontier models because they're ignorant of recent events, though it's too bad old frontier models just kind of disappear into the aether when a company moves on to the next iteration. Every company's frontier model today is a time capsule for the future. There should probably be some kind of preservation attempts made early so they don't wind up simply deleted; once we're in Internet time, sifting through the data to ensure scrapes are accurately dated becomes a nightmare unless you're doing your own regular Internet scrapes over a long time.

It would be nice to go back substantially further, though it's not too far back that the commoner becomes voiceless in history and we just get a bunch of politics and academia. Great job; look forward to testing it out.

Re: History LLMs: Models trained exclusively on pre-1913 texts

#384

Earlier quoted context omitted.

Yes at all. I think you misunderstand the significance of "general computing". The binary string 01101110 is a general-purpose computer, for example.

No, that's insane. Computing is a dynamic process. A static string is not a computer.

It may be insane, but it's also true.

https://en.wikipedia.org/wiki/Rule_110

Re: History LLMs: Models trained exclusively on pre-1913 texts

#385
post #3

“Time-locked models don't roleplay; they embody their training data. Ranke-4B-1913 doesn't know about WWI because WWI hasn't happened in its textual universe. It can be surprised by your questions in ways modern LLMs cannot.” “Modern LLMs suffer from hindsight contamination. GPT-5 knows how the story ends—WWI, the League's failure, the Spanish flu.” This is really fascinating. As someone who reads a lot of history an…

Perhaps I'm overly sensitive to this and terminally online, but that first quote reads as a textbook LLM-generated sentence.

" doesn't , it "

Later parts of the readme (whole section of bullets enumerating what it is and what it isn't, another LLM favorite) make me more confident that significant parts of the readme is generated.

I'm generally pro-AI, but if you spend hundreds of hours making a thing, I'd rather hear your explanation of it, not an LLM's.

Re: History LLMs: Models trained exclusively on pre-1913 texts

#386

Earlier quoted context omitted.

Ok... So as a black person should I demand that all books written before the civil rights act be destroyed? The past is messy. But it's the only way to learn anything. All an LLM does it's take a bunch of existing texts and rebundle them. Like it or not, the existing texts are still there. I understand an LLM that won't tell me how to do heart surgery. But I can't fear one that might be less enlightened on race issue…

we're on the same page.

Although...

Self preservation is the first law of nature. If you release the model someone will basically say you endorse those views and you risk your funding being cut.

You created Pandora's box and now you're afraid of opening it.

Re: History LLMs: Models trained exclusively on pre-1913 texts

#387
post #207

> Historical texts contain racism, antisemitism, misogyny, imperialist views. The models will reproduce these views because they're in the training data. This isn't a flaw, but a crucial feature—understanding how such views were articulated and normalized is crucial to understanding how they took hold. Yes! > We're developing a responsible access framework that makes models available to researchers for scholarly purp…

It’s as if every researcher in this field is getting high on the small amount of power they have from denying others access to their results. I’ve never been as unimpressed by scientists as I have been in the past five years or so. “We’ve created something so dangerous that we couldn’t possibly live with the moral burden of knowing that the wrong people (which are never us, of course) might get their hands on it, so…

Wow, this is needlessly antagonistic. Given the emergence of online communities that bond on conspiracy theories and racist philosophies in the 20th century, it's not hard to imagine the consequences of widely disseminating an LLM that could be used to propagate and further these discredited (for example, racial) scientific theories for bad ends by uneducated people in these online communities.

We can debate on whether it's good or not, but ultimately they're publishing it and in some very small way responsible for some of its ends. At least that's how I can see their interest in disseminating the use of the LLM through a responsible framework.

Re: History LLMs: Models trained exclusively on pre-1913 texts

#388

I'm surprised you can do this with a relatively modest corpus of text (compared to the petabytes you can vacuum up from modern books, Wikipedia, and random websites). But if it works, that's actually fantastic, because it lets you answer some interesting questions about LLMs being able to make new discoveries or transcend the training set in other ways. Forget relativity: can an LLM trained on this data notice any in…

> https://github.com/DGoettlich/history-llms/blob/main/ranke-4... Given the training notes, it seems like you can't get the performance they give examples of? I'm not sure about the exact details but there is some kind of targetted distillation of GPT-5 involved to try and get more conversational text and better performance. Which seems a bit iffy to me.

Thanks for the comment. Could you elaborate on what you find iffy about our approach? I'm sure we can improve!

Re: History LLMs: Models trained exclusively on pre-1913 texts

#389

Earlier quoted context omitted.

What is your worst-case scenario here? Something like a pop-sci article along the lines of "Mad scientists create racist, imperialistic AI"? I honestly don't see publication of the weights as a relevant risk factor, because sensationalist misrepresentation is trivially possible with the given example responses alone. I don't think such pseudo-malicious misrepresentation of scientific research can be reliably prevente…

It seems like if there is an obvious misuse of a tool, one has a moral imperative to restrict use of the tool.

Every tool can be misused. Hammers are as good for bashing heads as building houses. Restricting hammers would be silly and counterproductive.

Re: History LLMs: Models trained exclusively on pre-1913 texts

#390
post #323
post #308

Earlier quoted context omitted.

Of course, I have to assume that you have considered more outcomes than I have. Because, from my five minutes of reflection as a software geek, albeit with a passion for history, I find this the most surprising thing about the whole project. I suspect restricting access could equally be a comment on modern LLMs in general, rather than the historical material specifically. For example, we must be constantly reminded n…

They aren't afraid of hallucinations. Their first example is a hallucination, an imaginary biography of a Hitler who never lived. Their concern can't be understood without a deep understanding of the far left wing mind. Leftists believe people are so infinitely malleable that merely being exposed to a few words of conservative thought could instantly "convert" someone into a mortal enemy of their ideology for life. I…

They said it plainly ("dark corners that someone could use to misrepresent the goals of our project"): they just don't want to see their project in headlines about "Researchers create racist LLM!".
Post reply on HN