Wait so what does the model think that it is? If it doesn't know computers exist yet, I mean, and you ask it how it works, what does it say?
It would be nice if we could get an LLM to simply say, "We (I) don't know." I'll be the first to admit I don't know nearly enough about LLMs to make an educated comment, but perhaps someone here knows more than I do. Is that what a Hallucination is? When the AI model just sort of strings along an answer to the best of its ability. I'm mostly referring to ChatGPT and Gemini here, as I've seen that type of behavior wit…
History LLMs: Models trained exclusively on pre-1913 texts
221–230 of 452 posts
Re: History LLMs: Models trained exclusively on pre-1913 texts
#222> Historical texts contain racism, antisemitism, misogyny, imperialist views. The models will reproduce these views because they're in the training data. This isn't a flaw, but a crucial feature—understanding how such views were articulated and normalized is crucial to understanding how they took hold. Yes! > We're developing a responsible access framework that makes models available to researchers for scholarly purp…
It’s as if every researcher in this field is getting high on the small amount of power they have from denying others access to their results. I’ve never been as unimpressed by scientists as I have been in the past five years or so. “We’ve created something so dangerous that we couldn’t possibly live with the moral burden of knowing that the wrong people (which are never us, of course) might get their hands on it, so…
This is absolutely nothing new. With experimental things, it's non uncommon for a lab to develop a new technique and omit slight but important details to give them a competitive advantage. Similarly in the simulation/modelling space it's been common for years for researchers to not publish their research software. There's been a lot of lobbying on that side by groups such as the Software Sustainability Institute and Research Software Engineer organisations like RSE UK and RSE US, but there's a lot of researchers that just think that they shouldn't have to do it, even when publicly funded.
Re: History LLMs: Models trained exclusively on pre-1913 texts
#223Everyone learns that the renaissance was sparked by the translation of Ancient Greek works. But few know that the Renaissance was written in Latin — and has barely been translated. Less than 3% of I’m working on a project to change that. Research blog at www.SecondRenaissance.ai — we are starting by scanning and translating thousands of books at the Embassy of the Free Mind in Amsterdam, a UNESCO-recognized rare book…
This ia very cool but should go in a Show HN post as per HN rules. All the best!
Re: History LLMs: Models trained exclusively on pre-1913 texts
#224Earlier quoted context omitted.
Banning Huckleberry Finn from a school district should be grounds for immediate dismissal.
I don't support banning the book, but I think it is hard book to teach because it needs SO much context and a mature audience (lol good luck). Also, there are hundreds of other books from that era that are relevant even from Mark Twain's corpus so being obstinate about that book is a questionable position. I'm ambivalent honestly, but definitely not willing to die on that hill. (I graduated highschool in 1989 from a…
Re: History LLMs: Models trained exclusively on pre-1913 texts
#225It would be interesting to see how hard it would be to walk these models towards general relativity and quantum mechanics. Einstein’s paper “On the Electrodynamics of Moving Bodies” with special relativity was published in 1905. His work on general relativity was published 10 years later in 1915. The earliest knowledge cuttoff of these models is 1913, in between the relativity papers. The knowledge cutoffs are also r…
the issue is there is very little text before the internet, so not enough historical tokens to train a really big model
Re: History LLMs: Models trained exclusively on pre-1913 texts
#226How can this thing possibly be even remotely coherent with just fine tuning amounts of data used for pretraining?
Re: History LLMs: Models trained exclusively on pre-1913 texts
#227Earlier quoted context omitted.
When you put it that way it reminds me of the Severn/Keats character in the Hyperion Cantos. Far-future AIs reconstruct historical figures from their writings in an attempt to gain philosophical insights.
This isn’t science fiction anymore. CIA is using chatbot simulations of world leaders to inform analysts. https://archive.ph/9KxkJ
Re: History LLMs: Models trained exclusively on pre-1913 texts
#228> Imagine you could interview thousands of educated individuals from 1913—readers of newspapers, novels, and political treatises—about their views on peace, progress, gender roles, or empire. Not just survey them with preset questions, but engage in open-ended dialogue, probe their assumptions, and explore the boundaries of thought in that moment. Hell yeah, sold, let’s go… > We're developing a responsible access fra…
I wonder how much GPU compute you would need to create a public domain version of this. This would be a really valuable for the general public.
Re: History LLMs: Models trained exclusively on pre-1913 texts
#229I'm surprised you can do this with a relatively modest corpus of text (compared to the petabytes you can vacuum up from modern books, Wikipedia, and random websites). But if it works, that's actually fantastic, because it lets you answer some interesting questions about LLMs being able to make new discoveries or transcend the training set in other ways. Forget relativity: can an LLM trained on this data notice any in…
Here they do 80B tokens for a 4B model.