Live data from Hacker News

History LLMs: Models trained exclusively on pre-1913 texts

github.com

201–210 of 452 posts

Re: History LLMs: Models trained exclusively on pre-1913 texts

#201

> Imagine you could interview thousands of educated individuals from 1913—readers of newspapers, novels, and political treatises—about their views on peace, progress, gender roles, or empire. Not just survey them with preset questions, but engage in open-ended dialogue, probe their assumptions, and explore the boundaries of thought in that moment. Hell yeah, sold, let’s go… > We're developing a responsible access fra…

How would one even "misuse" a historical LLM, ask it how to cook up sarine gas in a trench?

Re: History LLMs: Models trained exclusively on pre-1913 texts

#202

It would be interesting to see how hard it would be to walk these models towards general relativity and quantum mechanics. Einstein’s paper “On the Electrodynamics of Moving Bodies” with special relativity was published in 1905. His work on general relativity was published 10 years later in 1915. The earliest knowledge cuttoff of these models is 1913, in between the relativity papers. The knowledge cutoffs are also r…

the issue is there is very little text before the internet, so not enough historical tokens to train a really big model

Re: History LLMs: Models trained exclusively on pre-1913 texts

#203

> Imagine you could interview thousands of educated individuals from 1913—readers of newspapers, novels, and political treatises—about their views on peace, progress, gender roles, or empire. Not just survey them with preset questions, but engage in open-ended dialogue, probe their assumptions, and explore the boundaries of thought in that moment. Hell yeah, sold, let’s go… > We're developing a responsible access fra…

How would one even "misuse" a historical LLM, ask it how to cook up sarine gas in a trench?

Ask it to write a document called "Project 2025".

Re: History LLMs: Models trained exclusively on pre-1913 texts

#204
>Historical texts contain racism, antisemitism, misogyny, imperialist views. The models will reproduce these views because they're in the training data. This isn't a flaw, but a crucial feature—understanding how such views were articulated and normalized is crucial to understanding how they took hold.

Yes!

>We're developing a responsible access framework that makes models available to researchers for scholarly purposes while preventing misuse.

Noooooo!

So is the model going to be publicly available, just like those dangerous pre-1913 texts, or not?

Re: History LLMs: Models trained exclusively on pre-1913 texts

#205
post #184

Earlier quoted context omitted.

There's a thriving startup scene in that direction.

Wasn't that the elevator pitch for Palentir? Still can't believe people buy their stock, given that they are the closest thing to a James Bond villain, just because it goes up. I mean, they are literally called "the stuff Sauron uses to control his evil forces". It's so on the nose it reads like an anime plot.

To the proud contrarian, "the empire did nothing wrong". Maybe Sci-fi has actually played a role in the "memetic desire" of some of the titans of tech who are trying to bring about these worlds more-or-less intentionally. I guess it's not as much of a dystopia if you're on top and its not evil if you think of it as inevitable anyway.

Re: History LLMs: Models trained exclusively on pre-1913 texts

#206
post #184

Earlier quoted context omitted.

There's a thriving startup scene in that direction.

Wasn't that the elevator pitch for Palentir? Still can't believe people buy their stock, given that they are the closest thing to a James Bond villain, just because it goes up. I mean, they are literally called "the stuff Sauron uses to control his evil forces". It's so on the nose it reads like an anime plot.

Stock buying as a political or ethical statement is not much of a thing. For one the stocks will still be bought by persons with less strung opinions, and secondly it does not lend itself well to virtue signaling.

Re: History LLMs: Models trained exclusively on pre-1913 texts

#207

> Historical texts contain racism, antisemitism, misogyny, imperialist views. The models will reproduce these views because they're in the training data. This isn't a flaw, but a crucial feature—understanding how such views were articulated and normalized is crucial to understanding how they took hold. Yes! > We're developing a responsible access framework that makes models available to researchers for scholarly purp…

It’s as if every researcher in this field is getting high on the small amount of power they have from denying others access to their results. I’ve never been as unimpressed by scientists as I have been in the past five years or so.

“We’ve created something so dangerous that we couldn’t possibly live with the moral burden of knowing that the wrong people (which are never us, of course) might get their hands on it, so with a heavy heart, we decided that we cannot just publish it.”

Meanwhile, anyone can hop on an online journal and for a nominal fee read articles describing how to genetically engineer deadly viruses, how to synthesize poisons, and all kinds of other stuff that is far more dangerous than what these LARPers have cooked up.

Re: History LLMs: Models trained exclusively on pre-1913 texts

#209
> Modern LLMs suffer from hindsight contamination. GPT-5 knows how the story ends—WWI, the League's failure, the Spanish flu. This knowledge inevitably shapes responses, even when instructed to "forget.

> Our data comes from more than 20 open-source datasets of historical books and newspapers. ... We currently do not deduplicate the data. The reason is that if documents show up in multiple datasets, they also had greater circulation historically. By leaving these duplicates in the data, we expect the model will be more strongly influenced by documents of greater historical importance.

I found these claims contradictory. Many books that modern readers consider historically significant had only niche circulation at the time of publishing. A quick inquiry likely points to later works by Nietzsche and Marx's Das Kapital. They're possible subjects to the duplication likely influencing the model's responses as if they had been widely known at the time

Re: History LLMs: Models trained exclusively on pre-1913 texts

#210

Everyone learns that the renaissance was sparked by the translation of Ancient Greek works. But few know that the Renaissance was written in Latin — and has barely been translated. Less than 3% of I’m working on a project to change that. Research blog at www.SecondRenaissance.ai — we are starting by scanning and translating thousands of books at the Embassy of the Free Mind in Amsterdam, a UNESCO-recognized rare book…

This ia very cool but should go in a Show HN post as per HN rules. All the best!
Post reply on HN