Earlier quoted context omitted.
I think you are confusing research with commodification. This is a research project, and it is clear how it was trained, and targeted at experts, enthusiasts, historians. Like if I was studying racism, the reference books explicitly written to dissect racism wouldn't be racist agents with a racist agenda. And as a result, no one is banning these books (except conservatives that want to retcon american history). Found…
> And as a result, no one is banning these books (except conservatives that want to retcon american history). My (very liberal) local school district banned English teachers from teaching any book that contained the n-word, even at a high-school level, and even when the author was a black person talking about real events that happened to them. FWIW, this was after complaints involving Of Mice and Men being on the cur…
History LLMs: Models trained exclusively on pre-1913 texts
101–110 of 452 posts
Re: History LLMs: Models trained exclusively on pre-1913 texts
#102Earlier quoted context omitted.
I think you are confusing research with commodification. This is a research project, and it is clear how it was trained, and targeted at experts, enthusiasts, historians. Like if I was studying racism, the reference books explicitly written to dissect racism wouldn't be racist agents with a racist agenda. And as a result, no one is banning these books (except conservatives that want to retcon american history). Found…
> And as a result, no one is banning these books (except conservatives that want to retcon american history). My (very liberal) local school district banned English teachers from teaching any book that contained the n-word, even at a high-school level, and even when the author was a black person talking about real events that happened to them. FWIW, this was after complaints involving Of Mice and Men being on the cur…
* https://abcnews.go.com/US/conservative-liberal-book-bans-dif...
* https://www.commondreams.org/news/book-banning-2023
*https://en.wikipedia.org/wiki/Book_banning_in_the_United_Sta...
Re: History LLMs: Models trained exclusively on pre-1913 texts
#103Wait so what does the model think that it is? If it doesn't know computers exist yet, I mean, and you ask it how it works, what does it say?
But with pre-1913 training, I would indeed be worried again I'd send it into an existential crisis. It has no knowledge whatsoever of what it is. But with a couple millennia of philosophical texts, it might come up with some interesting theories.
Re: History LLMs: Models trained exclusively on pre-1913 texts
#104Earlier quoted context omitted.
Public access, triggering a few racist responses from the model, a viral post on Xitter, the usual outrage, a scandal, the project gets publicly vilified, financing ceases. The researchers carry the tail of negative publicity throughout their remaining careers. Why risk all this?
I think you are confusing research with commodification. This is a research project, and it is clear how it was trained, and targeted at experts, enthusiasts, historians. Like if I was studying racism, the reference books explicitly written to dissect racism wouldn't be racist agents with a racist agenda. And as a result, no one is banning these books (except conservatives that want to retcon american history). Found…
No books should ever be banned. Doesn’t matter how vile it is.
Re: History LLMs: Models trained exclusively on pre-1913 texts
#105Earlier quoted context omitted.
Respectfully, LLMs are nothing like a brain, and I discourage comparisons between the two, because beyond a complete difference in the way they operate, a brain can innovate, and as of this moment, an LLM cannot because it relies on previously available information. LLMs are just seemingly intelligent autocomplete engines, and until they figure a way to stop the hallucinations, they aren't great either. Every piece o…
This is the 2023 take on LLMs. It still gets repeated a lot. But it doesn’t really hold up anymore - it’s more complicated than that. Don’t let some factoid about how they are pretrained on autocomplete-like next token prediction fool you into thinking you understand what is going on in that trillion parameter neural network. Sure, LLMs do not think like humans and they may not have human-level creativity. Sometimes…
This is just an appeal to complexity, not a rebuttal to the critique of likening an LLM to a human brain.
> they are not “autocomplete on steroids” anymore either.
Yes, they are. The steroids are just even more powerful. By refining training data quality, increasing parameter size, and increasing context length we can squeeze more utility out of LLMs than ever before, but ultimately, Opus 4.5 is the same thing as GPT2, it's only that coherence lasts a few pages rather than a few sentences.
Re: History LLMs: Models trained exclusively on pre-1913 texts
#106Re: History LLMs: Models trained exclusively on pre-1913 texts
#107> We're developing a responsible access framework that makes models available to researchers for scholarly purposes while preventing misuse. The idea of training such a model is really a great one, but not releasing it because someone might be offended by the output is just stupid beyond believe.
Public access, triggering a few racist responses from the model, a viral post on Xitter, the usual outrage, a scandal, the project gets publicly vilified, financing ceases. The researchers carry the tail of negative publicity throughout their remaining careers. Why risk all this?
Re: History LLMs: Models trained exclusively on pre-1913 texts
#108Earlier quoted context omitted.
> And as a result, no one is banning these books (except conservatives that want to retcon american history). My (very liberal) local school district banned English teachers from teaching any book that contained the n-word, even at a high-school level, and even when the author was a black person talking about real events that happened to them. FWIW, this was after complaints involving Of Mice and Men being on the cur…
Banning Huckleberry Finn from a school district should be grounds for immediate dismissal.
Almost everybody in that book is an awful person, especially the most 'upstanding' of types. Even the protagonist is an awful person. The one and only exception is 'N* Jim' who is the only kind-hearted and genuinely decent person in the book. It's an entire story about how the appearances of people, and the reality of those people, are two very different things.
It being banned for using foul language, as educational outcomes continue to deteriorate, is just so perfectly ironic.
Re: History LLMs: Models trained exclusively on pre-1913 texts
#109Earlier quoted context omitted.
This isn’t science fiction anymore. CIA is using chatbot simulations of world leaders to inform analysts. https://archive.ph/9KxkJ
[flagged]
- Are you ( edit: on a ) paid version? - If paid, which model you used? - Can you share exact prompt?
I am genuinely asking for myself. I have never received an answer this direct, but I accept there is a level of variability.
Re: History LLMs: Models trained exclusively on pre-1913 texts
#110Earlier quoted context omitted.
This is the 2023 take on LLMs. It still gets repeated a lot. But it doesn’t really hold up anymore - it’s more complicated than that. Don’t let some factoid about how they are pretrained on autocomplete-like next token prediction fool you into thinking you understand what is going on in that trillion parameter neural network. Sure, LLMs do not think like humans and they may not have human-level creativity. Sometimes…
> Don’t let some factoid about how they are pretrained on autocomplete-like next token prediction fool you into thinking you understand what is going on in that trillion parameter neural network. This is just an appeal to complexity, not a rebuttal to the critique of likening an LLM to a human brain. > they are not “autocomplete on steroids” anymore either. Yes, they are. The steroids are just even more powerful. By…