Live data from Hacker News

History LLMs: Models trained exclusively on pre-1913 texts

github.com

341–350 of 452 posts

Re: History LLMs: Models trained exclusively on pre-1913 texts

#341
post #301
post #285

Earlier quoted context omitted.

My point was Kindles vanishing, not books vanishing. Kindles are in no way a prerequisite for reading books.

Thanks for clarifying, I see what you mean now.

I have found ebooks useful. Especially when I was traveling by air more. But certainly not essential for reading.

Re: History LLMs: Models trained exclusively on pre-1913 texts

#342

Earlier quoted context omitted.

This is the 2023 take on LLMs. It still gets repeated a lot. But it doesn’t really hold up anymore - it’s more complicated than that. Don’t let some factoid about how they are pretrained on autocomplete-like next token prediction fool you into thinking you understand what is going on in that trillion parameter neural network. Sure, LLMs do not think like humans and they may not have human-level creativity. Sometimes…

I use enterprise LLM provided by work, working on very proprietary codebase on a semi esoteric language. My impression is it is still a very big autocompletion machine. You still need to hand hold it all the way as it is only capable of regurgitating the tiny amount of code patterns it saw in the public. As opposed to say a Python project.

What model is your “enterprise LLM”?

But regardless, I don’t think anyone is claiming that LLMs can magically do things that aren’t in their training data or context window. Obviously not: they can’t learn on the job and the permanent knowledge they have is frozen in during training.

Re: History LLMs: Models trained exclusively on pre-1913 texts

#343
post #184

Earlier quoted context omitted.

There's a thriving startup scene in that direction.

Wasn't that the elevator pitch for Palentir? Still can't believe people buy their stock, given that they are the closest thing to a James Bond villain, just because it goes up. I mean, they are literally called "the stuff Sauron uses to control his evil forces". It's so on the nose it reads like an anime plot.

Still can't believe people buy their stock, given that they are the closest thing to a James Bond villain, just because it goes up.

I've been tempted to. "Everything will be terrible if these guys succeed, but at least I'll be rich. If they fail I'll lose money, but since that's the outcome I prefer anyway, the loss won't bother me."

Trouble is, that ship has arguably already sailed. No matter how rapidly things go to hell, it will take many years before PLTR is profitable enough to justify its half-trillion dollar market cap.

Re: History LLMs: Models trained exclusively on pre-1913 texts

#344
post #214

Earlier quoted context omitted.

I think, meme stocks contradict you.

Meme stocks are a symptom of the death of the American dream. Economic malaise leads to unsophisticated risk taking.

Well, two things lead to unsophisticated risk-taking, right... economic malaise, and unlimited surplus. Both conditions are easy to spot in today's world.

Re: History LLMs: Models trained exclusively on pre-1913 texts

#345

Earlier quoted context omitted.

No, not at all.

Yes at all. I think you misunderstand the significance of "general computing". The binary string 01101110 is a general-purpose computer, for example.

No, that's insane. Computing is a dynamic process. A static string is not a computer.

Re: History LLMs: Models trained exclusively on pre-1913 texts

#346

> Imagine you could interview thousands of educated individuals from 1913—readers of newspapers, novels, and political treatises—about their views on peace, progress, gender roles, or empire. Not just survey them with preset questions, but engage in open-ended dialogue, probe their assumptions, and explore the boundaries of thought in that moment. Hell yeah, sold, let’s go… > We're developing a responsible access fra…

How would one even "misuse" a historical LLM, ask it how to cook up sarine gas in a trench?

Its output might violate speech codes, and in much of the EU that is penalized much more seriously than violent crime.

Re: History LLMs: Models trained exclusively on pre-1913 texts

#347

Earlier quoted context omitted.

I used to teach 19th-century history, and the responses definitely sound like a Victorian-era writer. And they of course sound like writing (books and periodicals etc) rather than "chat": as other responders allude to, the fine-tuning or RL process for making them good at conversation was presumably quite different from what is used for most chatbots, and they're leaning very heavily into the pre-training texts. We d…

don't we have parlament transcripts? I remember something about Germany (or maybe even Prussia) developing fast script to preserve 1-to-1 what was said

I mentioned those in the post you’re replying to :)

It’s a better source for how people spoke than books etc, but it’s not really an accurate source for patterns of everyday conversation because people were making speeches rather than chatting.

Re: History LLMs: Models trained exclusively on pre-1913 texts

#348
post #3

“Time-locked models don't roleplay; they embody their training data. Ranke-4B-1913 doesn't know about WWI because WWI hasn't happened in its textual universe. It can be surprised by your questions in ways modern LLMs cannot.” “Modern LLMs suffer from hindsight contamination. GPT-5 knows how the story ends—WWI, the League's failure, the Spanish flu.” This is really fascinating. As someone who reads a lot of history an…

> This is really fascinating. As someone who reads a lot of history and historical fiction I think this is really intriguing. Imagine having a conversation with someone genuinely from the period, where they don’t know the “end of the story”.

Having the facts from the era is one thing, to make conclusions about things it doesn't know would require intelligence.

Re: History LLMs: Models trained exclusively on pre-1913 texts

#350

Isn’t there obvious problems baked into this approach, if this is used for anything but fun? LLM’s lie and fake facts all the time, they are also masters at enforcing the users bias, even unconscious ones. How even a professor of history could ensure that the generated text is actually based on the training material and representative of the feelings and opinions of the given time period, not enforcing his biases tow…

To me it is pretty clear that it’s being used for fun. I personally like reading nineteenth century novels more than more recent novels (I especially like the style of science fiction by Jules Verne). What if the model can generate text in that style I like?
Post reply on HN