Live data from Hacker News

History LLMs: Models trained exclusively on pre-1913 texts

github.com

391–400 of 452 posts

Re: History LLMs: Models trained exclusively on pre-1913 texts

#392
Why not use these as a benchmark for LLM ability to make breakthrough discoveries?

For example prompt the 1913 model to try and “Invent a new theory of gravity that doesn’t conflict with special relativity”

Would it be able to eventually get to GR? If not, could finding out why not illuminate important weaknesses.

Re: History LLMs: Models trained exclusively on pre-1913 texts

#393

Interesting ... I'd love to find one that had a cutoff date around 1980.

> Which new band will still be around in 45 years?

Excellent question! It looks like Two-Tone is bringing ska back with a new wave of punk rock energy! I think The Specials are pretty special and will likely be around for a long time.

On the other hand, the "new wave" movement of punk rock music will go nowhere. The Cure, Joy Division, Tubeway Army: check the dustbin behind the record stores in a few years.

Re: History LLMs: Models trained exclusively on pre-1913 texts

#394

Earlier quoted context omitted.

we're on the same page.

Although... Self preservation is the first law of nature. If you release the model someone will basically say you endorse those views and you risk your funding being cut. You created Pandora's box and now you're afraid of opening it.

They could add a text box where users have to explicitly type the following words before it lets them interact in any way with the model: "I understand this model was created with old texts so any racial or sexual statements are a byproduct of their time an do not represent in any way the views of the researchers".

That should be more than enough to clear any chance of misunderstanding.

Re: History LLMs: Models trained exclusively on pre-1913 texts

#395

Earlier quoted context omitted.

we're on the same page.

Although... Self preservation is the first law of nature. If you release the model someone will basically say you endorse those views and you risk your funding being cut. You created Pandora's box and now you're afraid of opening it.

i think we (whole section) are just talking past each other - we never said we'll lock it away. it was an announcement of a release, not a release. main purpose for us was getting feedback on the methodological aspects, as we clearly state. i understand you guys just wanted to talk to the thing though.

Re: History LLMs: Models trained exclusively on pre-1913 texts

#396
post #207

Earlier quoted context omitted.

It’s as if every researcher in this field is getting high on the small amount of power they have from denying others access to their results. I’ve never been as unimpressed by scientists as I have been in the past five years or so. “We’ve created something so dangerous that we couldn’t possibly live with the moral burden of knowing that the wrong people (which are never us, of course) might get their hands on it, so…

Wow, this is needlessly antagonistic. Given the emergence of online communities that bond on conspiracy theories and racist philosophies in the 20th century, it's not hard to imagine the consequences of widely disseminating an LLM that could be used to propagate and further these discredited (for example, racial) scientific theories for bad ends by uneducated people in these online communities. We can debate on wheth…

thanks. i think this just took on a weird dynamic. we never said we'd lock the model away. not sure how this impression seems to have emerged for some. that aside, it was an announcement of a release, not a release. the main purpose was gathering feedback on our methodology. standard procedure in our domain is to first gather criticism, incorporate it, then publish results. but i understand people just wanted to talk to it. fair enough!

Re: History LLMs: Models trained exclusively on pre-1913 texts

#397
post #358

> Historical texts contain racism, antisemitism, misogyny, imperialist views. The models will reproduce these views because they're in the training data. This isn't a flaw, but a crucial feature—understanding how such views were articulated and normalized is crucial to understanding how they took hold. Yes! > We're developing a responsible access framework that makes models available to researchers for scholarly purp…

> So is the model going to be publicly available, just like those dangerous pre-1913 texts, or not? 1. This implies a false equivalence. Releasing a new interactive AI model is indeed different in significant and practical ways from the status quo. Yes, there are already-released historical texts. The rational thing to do is weigh the impacts of introducing another thing. 2. Some people have a tendency to say "releas…

well put.

Re: History LLMs: Models trained exclusively on pre-1913 texts

#398

Earlier quoted context omitted.

This is the 2023 take on LLMs. It still gets repeated a lot. But it doesn’t really hold up anymore - it’s more complicated than that. Don’t let some factoid about how they are pretrained on autocomplete-like next token prediction fool you into thinking you understand what is going on in that trillion parameter neural network. Sure, LLMs do not think like humans and they may not have human-level creativity. Sometimes…

>> Sometimes they hallucinate. For someone speaking as you knew everything, you appear to know very little. Every LLM completion is a "hallucination", some of them just happen to be factually correct.

I can say "I don't know" in response to a question. Can an LLM?

Re: History LLMs: Models trained exclusively on pre-1913 texts

#400
post #4

The sample responses given are fascinating. It seems more difficult than normal to even tell that they were generated by an LLM, since most of us (terminally online) people have been training our brains' AI-generated text detection on output from models trained with a recent cutoff date. Some of the sample responses seem so unlike anything an LLM would say, obviously due to its apparent beliefs on certain concepts, t…

I used to teach 19th-century history, and the responses definitely sound like a Victorian-era writer. And they of course sound like writing (books and periodicals etc) rather than "chat": as other responders allude to, the fine-tuning or RL process for making them good at conversation was presumably quite different from what is used for most chatbots, and they're leaning very heavily into the pre-training texts. We d…

very interesting observation!
Post reply on HN