“Time-locked models don't roleplay; they embody their training data. Ranke-4B-1913 doesn't know about WWI because WWI hasn't happened in its textual universe. It can be surprised by your questions in ways modern LLMs cannot.” “Modern LLMs suffer from hindsight contamination. GPT-5 knows how the story ends—WWI, the League's failure, the Spanish flu.” This is really fascinating. As someone who reads a lot of history an…
> This is really fascinating. As someone who reads a lot of history and historical fiction I think this is really intriguing. Imagine having a conversation with someone genuinely from the period, where they don’t know the “end of the story”. Having the facts from the era is one thing, to make conclusions about things it doesn't know would require intelligence.
History LLMs: Models trained exclusively on pre-1913 texts
351–360 of 452 posts
Re: History LLMs: Models trained exclusively on pre-1913 texts
#352Re: History LLMs: Models trained exclusively on pre-1913 texts
#353Re: History LLMs: Models trained exclusively on pre-1913 texts
#354Re: History LLMs: Models trained exclusively on pre-1913 texts
#355> Imagine you could interview thousands of educated individuals from 1913—readers of newspapers, novels, and political treatises—about their views on peace, progress, gender roles, or empire. Not just survey them with preset questions, but engage in open-ended dialogue, probe their assumptions, and explore the boundaries of thought in that moment. Hell yeah, sold, let’s go… > We're developing a responsible access fra…
How would one even "misuse" a historical LLM, ask it how to cook up sarine gas in a trench?
What do these people fear the most? That the "truth" they been pushing is a lie.
Re: History LLMs: Models trained exclusively on pre-1913 texts
#356Earlier quoted context omitted.
This might just be the closest we get to a time machine for some time. Or maybe ever. Every "King Arthur travels to the year 2000" kinda script is now something that writes itself. > Imagine having a conversation with someone genuinely from the period, Imagine not just someone, but Aristotle or Leonardo or Kant!
I imagine King Arthur would say something like: Hwæt spricst þu be?
Re: History LLMs: Models trained exclusively on pre-1913 texts
#357> Imagine you could interview thousands of educated individuals from 1913—readers of newspapers, novels, and political treatises—about their views on peace, progress, gender roles, or empire. Not just survey them with preset questions, but engage in open-ended dialogue, probe their assumptions, and explore the boundaries of thought in that moment. Hell yeah, sold, let’s go… > We're developing a responsible access fra…
understand your frustration. i trust you also understand the models have some dark corners that someone could use to misrepresent the goals of our project. if you have ideas on how we could make the models more broadly accessible while avoiding that risk, please do reach out @ history-llms@econ.uzh.ch
Re: History LLMs: Models trained exclusively on pre-1913 texts
#358> Historical texts contain racism, antisemitism, misogyny, imperialist views. The models will reproduce these views because they're in the training data. This isn't a flaw, but a crucial feature—understanding how such views were articulated and normalized is crucial to understanding how they took hold. Yes! > We're developing a responsible access framework that makes models available to researchers for scholarly purp…
1. This implies a false equivalence. Releasing a new interactive AI model is indeed different in significant and practical ways from the status quo. Yes, there are already-released historical texts. The rational thing to do is weigh the impacts of introducing another thing.
2. Some people have a tendency to say "release everything" as if open-source software is equivalent to open-weights models. They aren't. They are different enough to matter.
3. Rhetorically, the quote across comes across as a pressure tactic. When I hear "are you going to do this or not?" I cringe.
4. The quote above feels presumptive to me, as if the commenter is owed something from the history-llms project.
5. People are rightfully bothered that Big Tech has vacuumed up public domain and even private information and turned it into a profit center. But we're talking about a university project with (let's be charitable) legitimate concerns about misuse.
6. There seems to be a lack of curiosity in play. I'd much rather see people asking e.g. "What factors are influencing your decision about publishing your underlying models?"
7. There are people who have locked-in a view that says AI-safety perspectives are categorically invalid. Accordingly, they have almost a knee-jerk reaction against even talk of "let's think about the implications before we release this."
8. This one might explain and underly most of the other points above. I see signs of a deeper problem at work here. Hiding behind convenient oversimplifications to justify what one wants does not make a sound moral argument; it is motivated reasoning a.k.a. psychological justification.
Re: History LLMs: Models trained exclusively on pre-1913 texts
#359Earlier quoted context omitted.
Fully automated toaster-fucker generator! https://news.ycombinator.com/item?id=25667362
Man, I think about that comment all the time, like at least weekly since it was posted. I can't be the only one.
(I mention this so more people can know the list exists, and hopefully email us more nominations when they see an unusually good and interesting comment.)
Re: History LLMs: Models trained exclusively on pre-1913 texts
#360> Imagine you could interview thousands of educated individuals from 1913—readers of newspapers, novels, and political treatises—about their views on peace, progress, gender roles, or empire. Not just survey them with preset questions, but engage in open-ended dialogue, probe their assumptions, and explore the boundaries of thought in that moment. Hell yeah, sold, let’s go… > We're developing a responsible access fra…
It's a shame isn't it! The public must be protected from the backwards thoughts of history. In case they misuse it. I guess what they're really saying is "we don't want you guys to cancel us".