Live data from Hacker News

History LLMs: Models trained exclusively on pre-1913 texts

github.com

121–130 of 452 posts

Re: History LLMs: Models trained exclusively on pre-1913 texts

#121

> We're developing a responsible access framework that makes models available to researchers for scholarly purposes while preventing misuse. The idea of training such a model is really a great one, but not releasing it because someone might be offended by the output is just stupid beyond believe.

You have to understand that while the rest of the world has moved on from 2020, academics are still living there. There are many strong leftists, many of whom are deeply censorious; there are many more timeservers and cowards, who are terrified of falling foul of the first group.

And there are force multipliers for all of this. Even if you yourself are a sensible and courageous person, you want to protect your project. What if your manager, ethics committee or funder comes under pressure?

Re: History LLMs: Models trained exclusively on pre-1913 texts

#122
post #4

The sample responses given are fascinating. It seems more difficult than normal to even tell that they were generated by an LLM, since most of us (terminally online) people have been training our brains' AI-generated text detection on output from models trained with a recent cutoff date. Some of the sample responses seem so unlike anything an LLM would say, obviously due to its apparent beliefs on certain concepts, t…

Oh definitely. One thing that immediately caught my mind is that the question asks the model about “homosexual men” but the model starts the response with “the homosexual man” instead. Changing the plural to the singular and then adding an article. Feels very old fashioned to me.

Re: History LLMs: Models trained exclusively on pre-1913 texts

#123
post #99

Earlier quoted context omitted.

> And as a result, no one is banning these books (except conservatives that want to retcon american history). My (very liberal) local school district banned English teachers from teaching any book that contained the n-word, even at a high-school level, and even when the author was a black person talking about real events that happened to them. FWIW, this was after complaints involving Of Mice and Men being on the cur…

It’s a big country of roughly half a billion people, you’ll always find examples if you look hard enough. It’s ridiculous/wrong that your district did this but frankly it’s the exception in liberal/progressive communities. It’s a very one-sided problem: * https://abcnews.go.com/US/conservative-liberal-book-bans-dif... * https://www.commondreams.org/news/book-banning-2023 * https://en.wikipedia.org/wiki/Book_banning_i…

[deleted]

Re: History LLMs: Models trained exclusively on pre-1913 texts

#124
I'm surprised you can do this with a relatively modest corpus of text (compared to the petabytes you can vacuum up from modern books, Wikipedia, and random websites). But if it works, that's actually fantastic, because it lets you answer some interesting questions about LLMs being able to make new discoveries or transcend the training set in other ways. Forget relativity: can an LLM trained on this data notice any inconsistencies in its scientific knowledge, devise experiments that challenge them, and then interpret the results? Can it intuit about the halting problem? Theorize about the structure of the atom?...

Of course, if it fails, the counterpoint will be "you just need more training data", but still - I would love to play with this.

Re: History LLMs: Models trained exclusively on pre-1913 texts

#126
post #97

Wait so what does the model think that it is? If it doesn't know computers exist yet, I mean, and you ask it how it works, what does it say?

Models don't think they're anything, they'll respond with whatever's in their context as to how they've been directed to act. If it hasn't been told to have a persona, it won't think its anything, chatgpt isn't sentient

Re: History LLMs: Models trained exclusively on pre-1913 texts

#127
post #3

“Time-locked models don't roleplay; they embody their training data. Ranke-4B-1913 doesn't know about WWI because WWI hasn't happened in its textual universe. It can be surprised by your questions in ways modern LLMs cannot.” “Modern LLMs suffer from hindsight contamination. GPT-5 knows how the story ends—WWI, the League's failure, the Spanish flu.” This is really fascinating. As someone who reads a lot of history an…

I was going to say the same thing. Its really hard to explain the concept of "convincing but undoubtedly pretending", yet they captured that concept so beautifully here.

Re: History LLMs: Models trained exclusively on pre-1913 texts

#128
post #99

Earlier quoted context omitted.

> And as a result, no one is banning these books (except conservatives that want to retcon american history). My (very liberal) local school district banned English teachers from teaching any book that contained the n-word, even at a high-school level, and even when the author was a black person talking about real events that happened to them. FWIW, this was after complaints involving Of Mice and Men being on the cur…

It’s a big country of roughly half a billion people, you’ll always find examples if you look hard enough. It’s ridiculous/wrong that your district did this but frankly it’s the exception in liberal/progressive communities. It’s a very one-sided problem: * https://abcnews.go.com/US/conservative-liberal-book-bans-dif... * https://www.commondreams.org/news/book-banning-2023 * https://en.wikipedia.org/wiki/Book_banning_i…

A practical issue is the sort of books being banned. Your first link offer examples of one side trying to ban Of Mice and Men, Adventures of Huckleberry Finn, and Dr. Seuss, with the other side trying to ban many books along the lines of Gender Queer. [1] That link is to the book - which is animated, and quite NSFW.

There are a bizarrely large number similar book as Gender Queer being published, which creates the numeric discrepancy. The irony is that if there was an equal but opposite to that book about straight sex, sexuality, associated kinks, and so forth - then I think both liberals and conservatives would probably be all for keeping it away from schools. It's solely focused on sexuality, is quite crude, illustrated, targeted towards young children, and there's no moral beyond the most surface level writing which is about coming to terms with one's sexuality.

And obviously coming to terms with one's sexuality is very important, but I really don't think books like that are doing much to aid in that - especially when it's targeted at an age demographic that's still going to be extremely confused, and even moreso in a day and age when being different, if only for the sake of being different, is highly desirable. And given the nature of social media and the internet, decisions made today may stay with you for the rest of your life.

So for instance about 30% of Gen Z now declare themselves LGBT. [2] We seem to have entered into an equal but opposite problem of the past when those of deviant sexuality pretended to be straight to fit into societal expectations. And in many ways this modern twist is an even more damaging form of the problem from a variety of perspectives - fertility, STDs, stuff staying with you for the rest of your life, and so on. Let alone extreme cases where e.g. somebody engages in transition surgery or 1-way chemically induced changes which they end up later regretting.

[1] - https://archive.org/details/gender-queer-a-memoir-by-maia-ko...

[2] - https://www.nbcnews.com/nbc-out/out-news/nearly-30-gen-z-adu...

Re: History LLMs: Models trained exclusively on pre-1913 texts

#129

It would be interesting to see how hard it would be to walk these models towards general relativity and quantum mechanics. Einstein’s paper “On the Electrodynamics of Moving Bodies” with special relativity was published in 1905. His work on general relativity was published 10 years later in 1915. The earliest knowledge cuttoff of these models is 1913, in between the relativity papers. The knowledge cutoffs are also r…

> It would be interesting to see how hard it would be to walk these models towards general relativity and quantum mechanics.

Definitely. Even more interesting could be seeing them fall into the same trappings of quackery, and come up with things like over the counter lobotomies and colloidal silver.

On a totally different note, this could be very valuable for writing period accurate books and screenplays, games, etc ...

Re: History LLMs: Models trained exclusively on pre-1913 texts

#130
post #3

“Time-locked models don't roleplay; they embody their training data. Ranke-4B-1913 doesn't know about WWI because WWI hasn't happened in its textual universe. It can be surprised by your questions in ways modern LLMs cannot.” “Modern LLMs suffer from hindsight contamination. GPT-5 knows how the story ends—WWI, the League's failure, the Spanish flu.” This is really fascinating. As someone who reads a lot of history an…

This might just be the closest we get to a time machine for some time. Or maybe ever.

Every "King Arthur travels to the year 2000" kinda script is now something that writes itself.

> Imagine having a conversation with someone genuinely from the period,

Imagine not just someone, but Aristotle or Leonardo or Kant!

Post reply on HN