Live data from Hacker News

History LLMs: Models trained exclusively on pre-1913 texts

github.com

321–330 of 452 posts

Re: History LLMs: Models trained exclusively on pre-1913 texts

#321
post #173

Earlier quoted context omitted.

There's an entire subreddit called LLMPhysics dedicated to "vibe physics". It's full of people thinking they are close to the next breakthrough encouraged by sycophantic LLMs while trying to prove various crackpot theories. I'd be careful venturing out into unknown territory together with an LLM. You can easily lure yourself into convincing nonsense with no one to pull you out.

Fully automated toaster-fucker generator! https://news.ycombinator.com/item?id=25667362

Man, I think about that comment all the time, like at least weekly since it was posted. I can't be the only one.

Re: History LLMs: Models trained exclusively on pre-1913 texts

#322

> Imagine you could interview thousands of educated individuals from 1913—readers of newspapers, novels, and political treatises—about their views on peace, progress, gender roles, or empire. Not just survey them with preset questions, but engage in open-ended dialogue, probe their assumptions, and explore the boundaries of thought in that moment. Hell yeah, sold, let’s go… > We're developing a responsible access fra…

understand your frustration. i trust you also understand the models have some dark corners that someone could use to misrepresent the goals of our project. if you have ideas on how we could make the models more broadly accessible while avoiding that risk, please do reach out @ history-llms@econ.uzh.ch

There's no such risk so you're not going to get any sensible ideas in response to this question. The goals of the project are history, you already made that clear. There's nothing more that needs to be done.

We all get that academics now exist in some kind of dystopian horror where they can get transitively blamed for the existence of anyone to the right of Lenin, but bear in mind:

1. The people who might try to cancel you are idiots unworthy of your respect, because if they're against this project, they're against the study of history in its entirety.

2. They will scream at you anyway no matter what you do.

3. You used (Swiss) taxpayer funds to develop these models. There is no moral justification for withholding from the public what they worked to pay for.

You already slathered your README with disclaimers even though you didn't even release the model at all, just showed a few examples of what it said - none of which are in any way surprising. That is far more than enough. Just release the models and if anyone complains, politely tell them to go complain to the users.

Re: History LLMs: Models trained exclusively on pre-1913 texts

#323
post #308

Earlier quoted context omitted.

understand your frustration. i trust you also understand the models have some dark corners that someone could use to misrepresent the goals of our project. if you have ideas on how we could make the models more broadly accessible while avoiding that risk, please do reach out @ history-llms@econ.uzh.ch

Of course, I have to assume that you have considered more outcomes than I have. Because, from my five minutes of reflection as a software geek, albeit with a passion for history, I find this the most surprising thing about the whole project. I suspect restricting access could equally be a comment on modern LLMs in general, rather than the historical material specifically. For example, we must be constantly reminded n…

They aren't afraid of hallucinations. Their first example is a hallucination, an imaginary biography of a Hitler who never lived.

Their concern can't be understood without a deep understanding of the far left wing mind. Leftists believe people are so infinitely malleable that merely being exposed to a few words of conservative thought could instantly "convert" someone into a mortal enemy of their ideology for life. It's therefore of paramount importance to ensure nobody is ever exposed to such words unless they are known to be extremely far left already, after intensive mental preparation, and ideally not at all.

That's why leftist spaces like universities insist on trigger warnings on Shakespeare's plays, why they're deadly places for conservatives to give speeches, why the sample answers from the LLM are hidden behind a dropdown and marked as sensitive, and why they waste lots of money training an LLM that they're terrified of letting anyone actually use. They intuit that it's a dangerous mind bomb because if anyone could hear old fashioned/conservative thought, it would change political outcomes in the real world today.

Anyone who is that terrified of historical documents really shouldn't be working in history at all, but it's academia so what do you expect? They shouldn't be allowed to waste money like this.

Re: History LLMs: Models trained exclusively on pre-1913 texts

#324
post #224

Earlier quoted context omitted.

I don't support banning the book, but I think it is hard book to teach because it needs SO much context and a mature audience (lol good luck). Also, there are hundreds of other books from that era that are relevant even from Mark Twain's corpus so being obstinate about that book is a questionable position. I'm ambivalent honestly, but definitely not willing to die on that hill. (I graduated highschool in 1989 from a…

I mean, you gotta read it. I’m not normally a huge fan of the classics; I find Steinbeck dry and tedious, and Hemingway to be self-indulgent and repetitious. Even Twain’s other work isn’t exactly to my taste. But I’ve read Huckleberry Finn three times—in elementary school just for fun, in high school because it was assigned, and I recently listened to it on audiobook—and enjoyed the hell out of each time. Banning it…

I have read it. I spent my 20s guiltily reading all of the books I was supposed to have read in high school but used Cliff's Notes instead. From my 20's perspective I found Finn insipid and hokey but that's because pop culture had recycled it hundreds of times since its first publication, however when I consider it from the period perspective I can see the satire and the pointed allegories that made Twain so formidable. (Funny you mention Hemingway. I loved his writing in my 20's, then went back and read some again in my 40's and was like "huh, this irritating and immature, no wonder i loved it in my 20's.")

Re: History LLMs: Models trained exclusively on pre-1913 texts

#325
post #323
post #308

Earlier quoted context omitted.

Of course, I have to assume that you have considered more outcomes than I have. Because, from my five minutes of reflection as a software geek, albeit with a passion for history, I find this the most surprising thing about the whole project. I suspect restricting access could equally be a comment on modern LLMs in general, rather than the historical material specifically. For example, we must be constantly reminded n…

They aren't afraid of hallucinations. Their first example is a hallucination, an imaginary biography of a Hitler who never lived. Their concern can't be understood without a deep understanding of the far left wing mind. Leftists believe people are so infinitely malleable that merely being exposed to a few words of conservative thought could instantly "convert" someone into a mortal enemy of their ideology for life. I…

[deleted]

Re: History LLMs: Models trained exclusively on pre-1913 texts

#326

> Historical texts contain racism, antisemitism, misogyny, imperialist views. The models will reproduce these views because they're in the training data. This isn't a flaw, but a crucial feature—understanding how such views were articulated and normalized is crucial to understanding how they took hold. Yes! > We're developing a responsible access framework that makes models available to researchers for scholarly purp…

fully understand you. we'd like to provide access but also guard against misrepresentations of our projects goals by pointing to e.g. racist generations. if you have thoughts on how we should do that, perhaps you could reach out at history-llms@econ.uzh.ch ? thanks in advance!

You can guard against misrepresentations of your goals by stating your goals clearly, which you already do. Any further misrepresentation is going to be either malicious or idiotic, a university should simply be able to deal with that.

Edit: just thought of a practical step you can take: host it somewhere else than github. If there's ever going to be a backlash the microsoft moderators might not take too kindly to the stuff about e.g. homosexuality, no matter how academic.

Re: History LLMs: Models trained exclusively on pre-1913 texts

#327

Earlier quoted context omitted.

the issue is there is very little text before the internet, so not enough historical tokens to train a really big model

> the issue is there is very little text before the internet, Hm there is a lot of text from before the internet, but most of it is not on internet. There is a weird gap in some circles because of that, people are rediscovering work from pre 1980s researchers that only exist in books that have never been re-edited and that virtually no one knows about.

There is no doubt trillions of tokens of general communication in all kinds of languages tucked away in national archives and private collections.

The National Archives of Spain alone have 350 million pages of documents going back to the 15th century, ranging from correspondence to testimony to charts and maps, but only 10% of it is digitized and a much smaller fraction is transcribed. Hopefully with how good LLMs are getting they can accelerate the transcription process and open up all of our historical documents as a huge historical LLM dataset.

Re: History LLMs: Models trained exclusively on pre-1913 texts

#329

Earlier quoted context omitted.

> LLMs are a general purpose computing paradigm. Yes, so is logistic regression.

No, not at all.

Yes at all. I think you misunderstand the significance of "general computing". The binary string 01101110 is a general-purpose computer, for example.

Re: History LLMs: Models trained exclusively on pre-1913 texts

#330
post #12

So many disclaimers about bias. I wonder how far back you have to go before the bias isn’t an issue. Not because it unbiased, but because we don’t recognize or care about the biases present.

It's always up to the reader to determine which biases they themself care about.

If you're wondering at what point "we" as a collective will stop caring about a bias or set of biases, I don't think such a time exists.

You'll never get everyone to agree on anything.

Post reply on HN