Live data from Hacker News

History LLMs: Models trained exclusively on pre-1913 texts

github.com

111–120 of 452 posts

Re: History LLMs: Models trained exclusively on pre-1913 texts

#111
post #8

I’d like to know how they chat-tuned it. Getting the base model is one thing, did they also make a bunch of conversations for SFT and if so how was it done? We develop chatbots while minimizing interference with the normative judgments acquired during pretraining (“uncontaminated bootstrapping”). So they are chat tuning, I wonder what “minimizing interference with normative judgements” really amounts to and how objec…

They have some more details at https://github.com/DGoettlich/history-llms/blob/main/ranke-4... Basically using GPT-5 and being careful

This explains why it uses modern prose and not something from the 19th century and earlier

Re: History LLMs: Models trained exclusively on pre-1913 texts

#113
post #3

“Time-locked models don't roleplay; they embody their training data. Ranke-4B-1913 doesn't know about WWI because WWI hasn't happened in its textual universe. It can be surprised by your questions in ways modern LLMs cannot.” “Modern LLMs suffer from hindsight contamination. GPT-5 knows how the story ends—WWI, the League's failure, the Spanish flu.” This is really fascinating. As someone who reads a lot of history an…

Watching a modern LLM chat with this would be fun.

Re: History LLMs: Models trained exclusively on pre-1913 texts

#114

It would be interesting to see how hard it would be to walk these models towards general relativity and quantum mechanics. Einstein’s paper “On the Electrodynamics of Moving Bodies” with special relativity was published in 1905. His work on general relativity was published 10 years later in 1915. The earliest knowledge cuttoff of these models is 1913, in between the relativity papers. The knowledge cutoffs are also r…

[deleted]

Re: History LLMs: Models trained exclusively on pre-1913 texts

#115
post #13

Earlier quoted context omitted.

"...what do you mean, 'World War One ?'"

I remember reading a children's book when I was young and the fact that people used the phrase "World War One" rather than "The Great War" was a clue to the reader that events were taking place in a certain time period. Never forgot that for some reason. I failed to catch the clue, btw.

It wouldn’t be totally implausible to use that phrase between the wars. The name “the First World War” was used as early as 1920, although not very common.

Re: History LLMs: Models trained exclusively on pre-1913 texts

#116
post #65

> We're developing a responsible access framework that makes models available to researchers for scholarly purposes while preventing misuse. The idea of training such a model is really a great one, but not releasing it because someone might be offended by the output is just stupid beyond believe.

Public access, triggering a few racist responses from the model, a viral post on Xitter, the usual outrage, a scandal, the project gets publicly vilified, financing ceases. The researchers carry the tail of negative publicity throughout their remaining careers. Why risk all this?

People know that models can be racist now. It's old hat. "LLM gets prompted into saying vile shit" hasn't been notable for years.

Re: History LLMs: Models trained exclusively on pre-1913 texts

#117
post #15
post #3

“Time-locked models don't roleplay; they embody their training data. Ranke-4B-1913 doesn't know about WWI because WWI hasn't happened in its textual universe. It can be surprised by your questions in ways modern LLMs cannot.” “Modern LLMs suffer from hindsight contamination. GPT-5 knows how the story ends—WWI, the League's failure, the Spanish flu.” This is really fascinating. As someone who reads a lot of history an…

When you put it that way it reminds me of the Severn/Keats character in the Hyperion Cantos. Far-future AIs reconstruct historical figures from their writings in an attempt to gain philosophical insights.

This is such a ridiculously good series. If you haven't read it yet, I thoroughly recommend it.

Re: History LLMs: Models trained exclusively on pre-1913 texts

#118

Earlier quoted context omitted.

This is the 2023 take on LLMs. It still gets repeated a lot. But it doesn’t really hold up anymore - it’s more complicated than that. Don’t let some factoid about how they are pretrained on autocomplete-like next token prediction fool you into thinking you understand what is going on in that trillion parameter neural network. Sure, LLMs do not think like humans and they may not have human-level creativity. Sometimes…

> Don’t let some factoid about how they are pretrained on autocomplete-like next token prediction fool you into thinking you understand what is going on in that trillion parameter neural network. This is just an appeal to complexity, not a rebuttal to the critique of likening an LLM to a human brain. > they are not “autocomplete on steroids” anymore either. Yes, they are. The steroids are just even more powerful. By…

This would be true if all training were based on sentence completion. But training involving RLHF and RLAIF is increasingly important, isn't it?

Re: History LLMs: Models trained exclusively on pre-1913 texts

#119
> Imagine you could interview thousands of educated individuals from 1913—readers of newspapers, novels, and political treatises—about their views on peace, progress, gender roles, or empire. Not just survey them with preset questions, but engage in open-ended dialogue, probe their assumptions, and explore the boundaries of thought in that moment.

Hell yeah, sold, let’s go…

> We're developing a responsible access framework that makes models available to researchers for scholarly purposes while preventing misuse.

Oh. By “imagine you could interview…” they didn’t mean me.

Re: History LLMs: Models trained exclusively on pre-1913 texts

#120
post #57

Earlier quoted context omitted.

This is definitely fascinating - being able to do AI brain surgery, and selectively tuning its knowledge and priors, you'd be able to create awesome and terrifying simulations.

Respectfully, LLMs are nothing like a brain, and I discourage comparisons between the two, because beyond a complete difference in the way they operate, a brain can innovate, and as of this moment, an LLM cannot because it relies on previously available information. LLMs are just seemingly intelligent autocomplete engines, and until they figure a way to stop the hallucinations, they aren't great either. Every piece o…

Are you sure about this?

LLMs are like a topographic map of language.

If you have 2 known mountains (domains of knowledge) you can likely predict there is a valley between them, even if you haven’t been there.

I think LLMs can approximate language topography based on known surrounding features so to speak, and that can produce novel information that would be similar to insight or innovation.

I’ve seen this in our lab, or at least, I think I have.

Curious how you see it.

Post reply on HN