Live data from Hacker News

History LLMs: Models trained exclusively on pre-1913 texts

github.com

81–90 of 452 posts

Re: History LLMs: Models trained exclusively on pre-1913 texts

#81
post #80

I would love to see this LLM try to solve math olympiad questions. I’ve been surprised by how well current LLMs perform on them, and usually explain that surprise away by assuming the questions and details about their answers are in the training set. It would be cool to see if the general approach to LLMs is capable of solving truly novel (novel to them) problems.

I suspect that it would fail terribly, it wasn't until the 1900s that the modern definition of a vector space was even created iirc. Something trained in maths up until the 1990s should have a shot though.

Re: History LLMs: Models trained exclusively on pre-1913 texts

#82
post #73
post #3

“Time-locked models don't roleplay; they embody their training data. Ranke-4B-1913 doesn't know about WWI because WWI hasn't happened in its textual universe. It can be surprised by your questions in ways modern LLMs cannot.” “Modern LLMs suffer from hindsight contamination. GPT-5 knows how the story ends—WWI, the League's failure, the Spanish flu.” This is really fascinating. As someone who reads a lot of history an…

I used to follow this blog — I believe it was somehow associated with Slate Star Codex? — anyways, I remember the author used to do these experiments on themselves where they spent a week or two only reading newspapers/media from a specific point in time and then wrote a blog about their experiences/takeaways On that same note, there was this great YouTube series called The Great War. It spanned from 2014-2018 (100 y…

The Great War series is phenomenal. A truly impressive project.

Re: History LLMs: Models trained exclusively on pre-1913 texts

#83
post #65

Earlier quoted context omitted.

Public access, triggering a few racist responses from the model, a viral post on Xitter, the usual outrage, a scandal, the project gets publicly vilified, financing ceases. The researchers carry the tail of negative publicity throughout their remaining careers. Why risk all this?

> triggering a few racist responses from the mode I feel like, ironically, it would be folks less concerned with political correctness/not being offensive that would abuse this opportunity to slander the project. But that’s just my gut.

[dead]

Re: History LLMs: Models trained exclusively on pre-1913 texts

#84
post #57

Earlier quoted context omitted.

Respectfully, LLMs are nothing like a brain, and I discourage comparisons between the two, because beyond a complete difference in the way they operate, a brain can innovate, and as of this moment, an LLM cannot because it relies on previously available information. LLMs are just seemingly intelligent autocomplete engines, and until they figure a way to stop the hallucinations, they aren't great either. Every piece o…

This is the 2023 take on LLMs. It still gets repeated a lot. But it doesn’t really hold up anymore - it’s more complicated than that. Don’t let some factoid about how they are pretrained on autocomplete-like next token prediction fool you into thinking you understand what is going on in that trillion parameter neural network. Sure, LLMs do not think like humans and they may not have human-level creativity. Sometimes…

[dead]

Re: History LLMs: Models trained exclusively on pre-1913 texts

#85
post #65

> We're developing a responsible access framework that makes models available to researchers for scholarly purposes while preventing misuse. The idea of training such a model is really a great one, but not releasing it because someone might be offended by the output is just stupid beyond believe.

Public access, triggering a few racist responses from the model, a viral post on Xitter, the usual outrage, a scandal, the project gets publicly vilified, financing ceases. The researchers carry the tail of negative publicity throughout their remaining careers. Why risk all this?

Because there are easy workarounds. If it becomes an issue, you can quickly add large disclaimers informing people that there might be offensive output because, well, it's trained on texts written during the age of racism.

People typically get outraged when they see something they weren't expecting. If you tell them ahead of time, the user typically won't blame you (they'll blame themselves for choosing to ignore the disclaimer).

And if disclaimers don't work, rebrand and relaunch it under a different name.

Re: History LLMs: Models trained exclusively on pre-1913 texts

#86
post #65

> We're developing a responsible access framework that makes models available to researchers for scholarly purposes while preventing misuse. The idea of training such a model is really a great one, but not releasing it because someone might be offended by the output is just stupid beyond believe.

Public access, triggering a few racist responses from the model, a viral post on Xitter, the usual outrage, a scandal, the project gets publicly vilified, financing ceases. The researchers carry the tail of negative publicity throughout their remaining careers. Why risk all this?

this is FUD.

Re: History LLMs: Models trained exclusively on pre-1913 texts

#88

A question for those who think LLM’s are the path to artificial intelligence: if a large language model trained on pre-1913 data is a window into the past, how is a large language model trained on pre-2025 data not effectively the same thing?

Counter question: how does a training set, representing a window into the past, differ from your own experience as an intelligent entity? Are you able to see into the future? How?

[deleted]

Re: History LLMs: Models trained exclusively on pre-1913 texts

#89
post #73
post #3

“Time-locked models don't roleplay; they embody their training data. Ranke-4B-1913 doesn't know about WWI because WWI hasn't happened in its textual universe. It can be surprised by your questions in ways modern LLMs cannot.” “Modern LLMs suffer from hindsight contamination. GPT-5 knows how the story ends—WWI, the League's failure, the Spanish flu.” This is really fascinating. As someone who reads a lot of history an…

I used to follow this blog — I believe it was somehow associated with Slate Star Codex? — anyways, I remember the author used to do these experiments on themselves where they spent a week or two only reading newspapers/media from a specific point in time and then wrote a blog about their experiences/takeaways On that same note, there was this great YouTube series called The Great War. It spanned from 2014-2018 (100 y…

The people that did the Great War series (at least some of them, I believe there was a little bit of a falling out) went on to do a WWII version on the World War II channel: https://youtube.com/@worldwartwo

They are currently in the middle of a Korean War version: https://youtube.com/@thekoreanwarbyindyneidell

Re: History LLMs: Models trained exclusively on pre-1913 texts

#90
post #65

> We're developing a responsible access framework that makes models available to researchers for scholarly purposes while preventing misuse. The idea of training such a model is really a great one, but not releasing it because someone might be offended by the output is just stupid beyond believe.

Public access, triggering a few racist responses from the model, a viral post on Xitter, the usual outrage, a scandal, the project gets publicly vilified, financing ceases. The researchers carry the tail of negative publicity throughout their remaining careers. Why risk all this?

I think you are confusing research with commodification.

This is a research project, and it is clear how it was trained, and targeted at experts, enthusiasts, historians. Like if I was studying racism, the reference books explicitly written to dissect racism wouldn't be racist agents with a racist agenda. And as a result, no one is banning these books (except conservatives that want to retcon american history).

Foundational models spewing racist white supremecist content when the trillion-dollar company forces it in your face is a vastly different scenario.

There's a clear difference.

Post reply on HN