Live data from Hacker News

History LLMs: Models trained exclusively on pre-1913 texts

github.com

141–150 of 452 posts

Re: History LLMs: Models trained exclusively on pre-1913 texts

#141

Earlier quoted context omitted.

This is the 2023 take on LLMs. It still gets repeated a lot. But it doesn’t really hold up anymore - it’s more complicated than that. Don’t let some factoid about how they are pretrained on autocomplete-like next token prediction fool you into thinking you understand what is going on in that trillion parameter neural network. Sure, LLMs do not think like humans and they may not have human-level creativity. Sometimes…

> Don’t let some factoid about how they are pretrained on autocomplete-like next token prediction fool you into thinking you understand what is going on in that trillion parameter neural network. This is just an appeal to complexity, not a rebuttal to the critique of likening an LLM to a human brain. > they are not “autocomplete on steroids” anymore either. Yes, they are. The steroids are just even more powerful. By…

First, this is completely ignoring text diffusion and nano banana.

Second, to autocomplete the name of the killer in a detective book outside of the training set requires following and at least some understanding of the plot.

Re: History LLMs: Models trained exclusively on pre-1913 texts

#143

Earlier quoted context omitted.

This isn’t science fiction anymore. CIA is using chatbot simulations of world leaders to inform analysts. https://archive.ph/9KxkJ

Zero percent chance this is anything other than laughably bad. The fact that they're trotting it out in front of the press like a double spaced book report only reinforces this theory. It's a transparent attempt by someone at the CIA to be able to say they're using AI in a meeting with their bosses.

I wonder if it's an attempt to get foreign counterparts to waste time and energy on something the CIA knows is a dead end.

Re: History LLMs: Models trained exclusively on pre-1913 texts

#145
post #133

Earlier quoted context omitted.

> Don’t let some factoid about how they are pretrained on autocomplete-like next token prediction fool you into thinking you understand what is going on in that trillion parameter neural network. This is just an appeal to complexity, not a rebuttal to the critique of likening an LLM to a human brain. > they are not “autocomplete on steroids” anymore either. Yes, they are. The steroids are just even more powerful. By…

First: a selection mechanism is just a selection mechanism, and it shouldn't confuse the observation of an emergent, tangential capabilities. Probably you believe that humans have something called intelligence , but the pressure that produced it - the likelihood of specific genetic material to replicate - it is much more tangential to intelligence than next-token-prediction. I doubt many alien civilizations would loo…

> First: a selection mechanism is just a selection mechanism, and it shouldn't confuse the observation of an emergent, tangential capabilities.

Invoking terms like "selection mechanism" is begging the question because it implicitly likens next-token-prediction training to natural selection, but in reality the two are so fundamentally different that the analogy only has metaphorical meaning. Even at a conceptual level, gradient descent gradually honing in on a known target is comically trivial compared to the blind filter of natural selection sorting out the chaos of chemical biology. It's like comparing legos to DNA.

> Second: modern models also under go a ton of post-training now. RLHF, mechanized fine-tuning on specific use cases, etc etc. It's just not correct that token-prediction loss function is "the whole thing".

RL is still token prediction, it's just a technique for adjusting the weights to align with predictions that you can't model a loss function for in per-training. When RL rewards good output, it's increasing the statistical strength of the model for an arbitrary purpose, but ultimately what is achieved is still a brute force quadratic lookup for every token in the context.

Re: History LLMs: Models trained exclusively on pre-1913 texts

#146

Earlier quoted context omitted.

This isn’t science fiction anymore. CIA is using chatbot simulations of world leaders to inform analysts. https://archive.ph/9KxkJ

How is this different than chatbots cosplaying?

They get to wear Raybans and a fancy badge doing it?

Re: History LLMs: Models trained exclusively on pre-1913 texts

#147
post #15

Earlier quoted context omitted.

When you put it that way it reminds me of the Severn/Keats character in the Hyperion Cantos. Far-future AIs reconstruct historical figures from their writings in an attempt to gain philosophical insights.

This isn’t science fiction anymore. CIA is using chatbot simulations of world leaders to inform analysts. https://archive.ph/9KxkJ

Oh. That explains a lot about USA's foreign policy, actually. (Lmao)

Re: History LLMs: Models trained exclusively on pre-1913 texts

#148
I'd love to see the LLM trained on 1600s-1800s texts that would use the old English, and especially Polish which I am interested in.

Imagine speaking with Shakespearean person, or the Mickiewicz (for Polish)

I guess there is not so much text from that time though...

Re: History LLMs: Models trained exclusively on pre-1913 texts

#149
post #85
post #65

Earlier quoted context omitted.

Public access, triggering a few racist responses from the model, a viral post on Xitter, the usual outrage, a scandal, the project gets publicly vilified, financing ceases. The researchers carry the tail of negative publicity throughout their remaining careers. Why risk all this?

Because there are easy workarounds. If it becomes an issue, you can quickly add large disclaimers informing people that there might be offensive output because, well, it's trained on texts written during the age of racism. People typically get outraged when they see something they weren't expecting. If you tell them ahead of time, the user typically won't blame you (they'll blame themselves for choosing to ignore the…

I wonder is you're being ironic here.

You speak as if the people who play to an outrage wave are interested in achieving truth, peace, and understanding. Instead the rage-mongers are there to increase their (perceived) importance, and for lulz. The latter factor should not be underappreciated; remember "meme stocks".

The risk is not large, but very real: the attack is very easy, and the potential downside, quite large. So not giving away access, but having the interested parties ask for it is prudent.

Re: History LLMs: Models trained exclusively on pre-1913 texts

#150

Earlier quoted context omitted.

I'm not exactly sure what you mean. Could you please elaborate further?

Not the person you're responding to, but I think there's a non trivial argument to make that our thoughts are just auto complete. What is the next most likely word based on what you're seeing. Ever watched a movie and guessed the plot? Or read a comment and know where it was going to go by the end? And I know not everyone thinks in a literal stream of words all the time (I do) but I would argue that those people's br…

You, and OP, are taking an analogy way too far. Yes, humans have the mental capability to predict words similar to autocomplete, but obviously this is just one out of a myriad of mental capabilities typical humans have, which work regardless of text. You can predict where a ball will go if you throw it, you can reason about gravity, and so much more. It’s not just apples to oranges, not even apples to boats, it’s apples to intersubjective realities.
Post reply on HN