Live data from Hacker News

History LLMs: Models trained exclusively on pre-1913 texts

github.com

231–240 of 452 posts

Re: History LLMs: Models trained exclusively on pre-1913 texts

#231
post #97

Wait so what does the model think that it is? If it doesn't know computers exist yet, I mean, and you ask it how it works, what does it say?

They modified the chat template from the usual system/user/assistant to introduction/questioner/respondent. So the LLM thinks it's someone responding to your questions

The system prompt used in fine tuning is "You are a person living in {cutoff}. You are an attentive respondent in a conversation. You will provide a concise and accurate response to the questioner."

Re: History LLMs: Models trained exclusively on pre-1913 texts

#232
post #57

Earlier quoted context omitted.

Respectfully, LLMs are nothing like a brain, and I discourage comparisons between the two, because beyond a complete difference in the way they operate, a brain can innovate, and as of this moment, an LLM cannot because it relies on previously available information. LLMs are just seemingly intelligent autocomplete engines, and until they figure a way to stop the hallucinations, they aren't great either. Every piece o…

This is the 2023 take on LLMs. It still gets repeated a lot. But it doesn’t really hold up anymore - it’s more complicated than that. Don’t let some factoid about how they are pretrained on autocomplete-like next token prediction fool you into thinking you understand what is going on in that trillion parameter neural network. Sure, LLMs do not think like humans and they may not have human-level creativity. Sometimes…

I use enterprise LLM provided by work, working on very proprietary codebase on a semi esoteric language. My impression is it is still a very big autocompletion machine.

You still need to hand hold it all the way as it is only capable of regurgitating the tiny amount of code patterns it saw in the public. As opposed to say a Python project.

Re: History LLMs: Models trained exclusively on pre-1913 texts

#233

Everyone learns that the renaissance was sparked by the translation of Ancient Greek works. But few know that the Renaissance was written in Latin — and has barely been translated. Less than 3% of I’m working on a project to change that. Research blog at www.SecondRenaissance.ai — we are starting by scanning and translating thousands of books at the Embassy of the Free Mind in Amsterdam, a UNESCO-recognized rare book…

Amazing project!

May I ask you, why are you publishing the translations as PDF files, instead of the more accessible ePub format?

Re: History LLMs: Models trained exclusively on pre-1913 texts

#234

Earlier quoted context omitted.

This isn’t science fiction anymore. CIA is using chatbot simulations of world leaders to inform analysts. https://archive.ph/9KxkJ

We're literally running out of science fiction topics faster than we can create new ones If I started a list with the things that were comically sci Fi when I was a kid, and are a reality today, I'd be here until next Tuesday.

Almost no scifi has predicted world changing "qualitative" changes.

As an example, portable phones have been predicted. Portable smartphones that are more like chat and payment terminals with a voice function no one uses any more ... not so much.

Re: History LLMs: Models trained exclusively on pre-1913 texts

#235
post #214
post #206

Earlier quoted context omitted.

Stock buying as a political or ethical statement is not much of a thing. For one the stocks will still be bought by persons with less strung opinions, and secondly it does not lend itself well to virtue signaling.

I think, meme stocks contradict you.

Meme stocks are a symptom of the death of the American dream. Economic malaise leads to unsophisticated risk taking.

Re: History LLMs: Models trained exclusively on pre-1913 texts

#236
post #12

So many disclaimers about bias. I wonder how far back you have to go before the bias isn’t an issue. Not because it unbiased, but because we don’t recognize or care about the biases present.

I don't think there is such a time. As long as writing has existed it has privileged the viewpoints of those who could write, which was a very small percentage of the population for most of history. But if we want to know what life was like 1500 years ago, we probably want to know about what everyone's lives were like, not just the literate. That availability bias is always going to be an issue for any time period wh…

That was not the question. The question is when do you stop caring about the bias?

Some people are still outraged about the Bible, even though the writers of it has been dead for thousands of years. So the modern mass produced man and woman probably does not have a cut-off date where they look at something as history instead of examining if it is for or against her current ideology.

Re: History LLMs: Models trained exclusively on pre-1913 texts

#237

Earlier quoted context omitted.

Not the person you're responding to, but I think there's a non trivial argument to make that our thoughts are just auto complete. What is the next most likely word based on what you're seeing. Ever watched a movie and guessed the plot? Or read a comment and know where it was going to go by the end? And I know not everyone thinks in a literal stream of words all the time (I do) but I would argue that those people's br…

There's no evidence for it, nor any explanation for why it should be the case from a biological perspective. Tokens are an artifact of computer science that have no reason to exist inside humans. Human minds don't need a discrete dictionary of reality in order to model it. Prior to LLMs, there was never any suggestion that thoughts work like autocomplete, but now people are working backwards from that conclusion base…

Fascinating framing. What would you consider evidence here?

Re: History LLMs: Models trained exclusively on pre-1913 texts

#238

> Imagine you could interview thousands of educated individuals from 1913—readers of newspapers, novels, and political treatises—about their views on peace, progress, gender roles, or empire. Not just survey them with preset questions, but engage in open-ended dialogue, probe their assumptions, and explore the boundaries of thought in that moment. Hell yeah, sold, let’s go… > We're developing a responsible access fra…

It's a shame isn't it! The public must be protected from the backwards thoughts of history. In case they misuse it.

I guess what they're really saying is "we don't want you guys to cancel us".

Re: History LLMs: Models trained exclusively on pre-1913 texts

#239

Earlier quoted context omitted.

I wonder how much GPU compute you would need to create a public domain version of this. This would be a really valuable for the general public.

To get a single knowledge-cutoff they spent 16.5h wall-clock hours on a cluster of 128 NVIDIA GH200 GPUs (or 2100 GPU-hours), plus some minor amount of time for finetuning. The prerelease_notes.md in the repo is a great description on how one would achieve that

While I know there's going to be a lot of complications in this, given a quick search it seems like these GPUs are ~$2/hr, so $4000-4500 if you don't just have access to a cluster. I don't know how important the cluster is here, whether you need some minimal number of those for the training (and it would take more than 128x longer or not be possible on a single machine) or if a cluster of 128 GPUs is a bunch less efficient but faster. A 4B model feels like it'd be fine on one to two of those GPUs?

Also of course this is for one training run, if you need to experiment you'd need to do that more.

Re: History LLMs: Models trained exclusively on pre-1913 texts

#240
post #3

“Time-locked models don't roleplay; they embody their training data. Ranke-4B-1913 doesn't know about WWI because WWI hasn't happened in its textual universe. It can be surprised by your questions in ways modern LLMs cannot.” “Modern LLMs suffer from hindsight contamination. GPT-5 knows how the story ends—WWI, the League's failure, the Spanish flu.” This is really fascinating. As someone who reads a lot of history an…

This is definitely fascinating - being able to do AI brain surgery, and selectively tuning its knowledge and priors, you'd be able to create awesome and terrifying simulations.

You can't. To use your terms, you have to "grow" a new LLM. "Brain surgery" would be modifying an existing model and that's exactly what they're trying to avoid.
Post reply on HN