Live data from Hacker News

History LLMs: Models trained exclusively on pre-1913 texts

github.com

161–170 of 452 posts

Re: History LLMs: Models trained exclusively on pre-1913 texts

#161
post #57

Earlier quoted context omitted.

Respectfully, LLMs are nothing like a brain, and I discourage comparisons between the two, because beyond a complete difference in the way they operate, a brain can innovate, and as of this moment, an LLM cannot because it relies on previously available information. LLMs are just seemingly intelligent autocomplete engines, and until they figure a way to stop the hallucinations, they aren't great either. Every piece o…

This is the 2023 take on LLMs. It still gets repeated a lot. But it doesn’t really hold up anymore - it’s more complicated than that. Don’t let some factoid about how they are pretrained on autocomplete-like next token prediction fool you into thinking you understand what is going on in that trillion parameter neural network. Sure, LLMs do not think like humans and they may not have human-level creativity. Sometimes…

> it’s more complicated than that.

No it isn't.

> ...fool you into thinking you understand what is going on in that trillion parameter neural network.

It's just matrix multiplication and logistic regression, nothing more.

Re: History LLMs: Models trained exclusively on pre-1913 texts

#162

Earlier quoted context omitted.

This is the 2023 take on LLMs. It still gets repeated a lot. But it doesn’t really hold up anymore - it’s more complicated than that. Don’t let some factoid about how they are pretrained on autocomplete-like next token prediction fool you into thinking you understand what is going on in that trillion parameter neural network. Sure, LLMs do not think like humans and they may not have human-level creativity. Sometimes…

As someone who still might have a '2023 take on LLMs', even though I use them often at work, where would you recommend I look to learn more about what a '2025 LLM' is, and how they operate differently?

Don't bother. This bubble will pop in two years, you don't want to look back on your old comments in shame in three.

Re: History LLMs: Models trained exclusively on pre-1913 texts

#163
post #28

> Imagine you could interview thousands of educated individuals from 1913—readers of newspapers, novels, and political treatises—about their views on peace, progress, gender roles, or empire. I don't mind the experimentation. I'm curious about where someone has found an application of it. What is the value of such a broad, generic viewpoint? What does it represent? What is it evidence of? The answer to both seems to…

This is a regurgitation of the old critique of history: what's it's purpose? What do you use it for? What is its application? One answer is that the study of history helps us understand that what we believe as "obviously correct" views today are as contingent on our current social norms and power structures (and their history) as the "obviously correct" views and beliefs of some point in the past. It's hard for most…

One thing I haven't seen anyone bring up yet in this thread, is that there's a big risk of leakage. If even big image models had CSAM sneak into their training material, how can we trust data from our time hasn't snuck into these historical models?

I've used Google books a lot in the past, and Google's time-filtering feature in searches too. Not to mention Spotify's search features targeting date of production. All had huge temporal mislabeling problems.

Re: History LLMs: Models trained exclusively on pre-1913 texts

#164
post #3

“Time-locked models don't roleplay; they embody their training data. Ranke-4B-1913 doesn't know about WWI because WWI hasn't happened in its textual universe. It can be surprised by your questions in ways modern LLMs cannot.” “Modern LLMs suffer from hindsight contamination. GPT-5 knows how the story ends—WWI, the League's failure, the Spanish flu.” This is really fascinating. As someone who reads a lot of history an…

That's some Westworld level of discussion

Re: History LLMs: Models trained exclusively on pre-1913 texts

#165

Earlier quoted context omitted.

Not the person you're responding to, but I think there's a non trivial argument to make that our thoughts are just auto complete. What is the next most likely word based on what you're seeing. Ever watched a movie and guessed the plot? Or read a comment and know where it was going to go by the end? And I know not everyone thinks in a literal stream of words all the time (I do) but I would argue that those people's br…

There's no evidence for it, nor any explanation for why it should be the case from a biological perspective. Tokens are an artifact of computer science that have no reason to exist inside humans. Human minds don't need a discrete dictionary of reality in order to model it. Prior to LLMs, there was never any suggestion that thoughts work like autocomplete, but now people are working backwards from that conclusion base…

There are so many theories regarding human cognition that you can certainly find something that is close to "autocomplete". A Hopfield network, for example.

Roots of predictive coding theory extend back to 1860s.

Natalia Bekhtereva was writing about compact concept representations in the brain akin to tokens.

Re: History LLMs: Models trained exclusively on pre-1913 texts

#166
post #97

Wait so what does the model think that it is? If it doesn't know computers exist yet, I mean, and you ask it how it works, what does it say?

It would be nice if we could get an LLM to simply say, "We (I) don't know."

I'll be the first to admit I don't know nearly enough about LLMs to make an educated comment, but perhaps someone here knows more than I do. Is that what a Hallucination is? When the AI model just sort of strings along an answer to the best of its ability. I'm mostly referring to ChatGPT and Gemini here, as I've seen that type of behavior with those tools in the past. Those are really the only tools I'm familiar with.

Re: History LLMs: Models trained exclusively on pre-1913 texts

#169
I had considered this task infeasible, due to a relative lack of training data. After all, isn't the received wisdom that you must shove every scrap of Common Crawl into your pre-training or you're doing it wrong? ;)

But reading the outputs here, it would appear that quality has won out over quantity after all!

Re: History LLMs: Models trained exclusively on pre-1913 texts

#170

That Adolf Hitler seems to be a hallucination. There's totally nothing googlable about him. Also what could be the language his works were translated from , into German?

I believe that's one of the primary issues LLMs aim to address. Many historical texts aren't directly Googleable because they haven't been converted to HTML, a format that Google can parse.
Post reply on HN