Live data from Hacker News

Talkie: a 13B vintage language model from 1930

talkie-lm.com

31–40 of 350 posts

Re: Talkie: a 13B vintage language model from 1930

#31

>Have you ever daydreamed about talking to someone from the past? Fun facts, LLM was once envisioned by Steve Jobs in one of his interviews [1]. Essentially one of his main wish in life is to meet and interract with Aristotle, in which according to him at the time, computer in the future can make it possible. [1] In 1985 Steve Jobs described a machine that would help people get answers from Aristotle–modern LLM [vide…

Except... not at all? The vast majority of the training data required to create an artificial Aristotle has been lost forever. Smash your coffee cup on the ground. Now reassemble it and put the coffee back in. Once you can repeatably do that I'll begin to believe you can train an artificial Aristotle.

Your bar is too low. With the coffee cup, you at least have access to all the pieces - in theory, although not in engineering practice. With Aristotle, you don't have anything close to that.

Recreating Aristotle in any meaningful way, other than a model trained on his surviving writing of a million or so words, is simply not possible even in principle.

Re: Talkie: a 13B vintage language model from 1930

#32
post #31

Earlier quoted context omitted.

Except... not at all? The vast majority of the training data required to create an artificial Aristotle has been lost forever. Smash your coffee cup on the ground. Now reassemble it and put the coffee back in. Once you can repeatably do that I'll begin to believe you can train an artificial Aristotle.

Your bar is too low. With the coffee cup, you at least have access to all the pieces - in theory, although not in engineering practice. With Aristotle, you don't have anything close to that. Recreating Aristotle in any meaningful way, other than a model trained on his surviving writing of a million or so words, is simply not possible even in principle.

OK I'll raise the bar--make sure when you reassemble the coffee cup and put the coffee back into it, the coffee is the exact same temperature as when you threw the whole shooting match onto the floor ;)

EDIT: and you don't get to re-heat it.

EDIT AGAIN: to be clear, in my post above (and this one) by "put the coffee back in" I meant more precisely "put every molecule of coffee that splashed/sloshed/flowed/whatever out when the cup smashed back into the re-assembled cup" i.e. "restore the system back to the initial state". Not "refill the glued-together pieces of your shattered coffee cup with new coffee".

Re: Talkie: a 13B vintage language model from 1930

#33

I have no real quibble with the blog post itself, but I take issue with the title that calls it a "vintage model". The blog post defines a "vintage model" as one that is trained only on data before a particular cutoff point: > Vintage LMs are contamination-free by construction, enabling unique generalization experiments [...] The most important objective when training vintage language models is that no data leaks int…

Yeah, the blog distinguishes between "contamination," which it describes as polluting the training data with answers to benchmarking questions, with "temporal leakage," which is polluting the training data with writing after the target date, but those seem to be nearly the same problem.

Not necessarily. The former is about data that’s supposed to be in there, but may actually be testing the model’s recall abilities rather than reasoning (ie rather than actually having a certain writing style, it just cites some passage it knows in that style).

The latter would be data not at all supposed to be in there, in this case, data after 1930.

Re: Talkie: a 13B vintage language model from 1930

#34
post #24

So interesting! Tell me about Winston Churchill: > Winston Churchill, who was born in 1871, is the son of the late Lord Randolph Churchill, and a grandson of the great Duke of Marlborough. He was educated at Harrow and at Sandhurst, and entered the army in 1890. In 1895 he retired from the service, and three years later he was returned to Parliament as Conservative member for Oldham. He has represented that constitue…

> The establishment of an Indian parliament is demanded, in which the queen shall be represented by a viceroy, Britain’s monarch was a king, not a queen, from about 1900-1950. Obviously there is some big “temporal leakage” from the training, which is affecting these predictions

Queen Victoria was direct ruler of India from 1858, and Empress of India from 1876 until 1901, so the "leakage" may not be from the future so much as the contemporaneously recent past. Same reason models get confused about what features work in what versions of software.

(Also, Queen Elizabeth I is the one who granted a royal charter to the East India Company, in 1600 - and that company eventually handed rule of India over to Queen Victoria. So British queens were a major presence in India.)

Re: Talkie: a 13B vintage language model from 1930

#35
post #23

We've got quite a list of history-only LLMs brewing on the Models Table. https://lifearchitect.ai/models-table/ This one is easiest to talk to in a HF space: https://huggingface.co/spaces/tventurella/mr_chatterbox

These are more like Small Language Models since the amount of textual data from the past is extremely limited, and most of what's out there hasn't even been digitized.

Re: Talkie: a 13B vintage language model from 1930

#37
post #25
post #22

Earlier quoted context omitted.

Predicting the future is problematic, agreed. Re: the Nate Silver nuclear weapons example, that's pretty weak - eg: given (say) I've just seen three heads in a row (exactly once) .. does that alter anything about "the odds". Having seen nuclear weapons not used post WWII ... does that inform us about "the odds" or the several times their use was almost certain (eg: Cuban missile crisis) save for out of band behaviour…

> Having seen nuclear weapons not used post WWII ... does that inform us about "the odds" This is what Bayesian prediction does > save for out of band behaviour by individuals that averted use and escalation? This is kind of the point being made.

> This is what Bayesian prediction does

Repeatedly, in a reproducible way, for events in the arrow of time? We can test this by going back to 1945 and running forward again?

> This is kind of the point being made.

Was it?

( assume I did a little math some decades past and have some poor grasp of Bayesian statistics )

Re: Talkie: a 13B vintage language model from 1930

#38
post #8

The Python example is fascinating, and a good rejoinder to anyone still dismissing LLM’s as stochastic parrots.

Indeed, I found this part extremely interesting. The more general vision of "testing a vintage model on something invented after its training data ended" seems like quite a strong test of "true cognition" (or training data contamination, if you haven't stopped up all the leakage...)

Re: Talkie: a 13B vintage language model from 1930

#39
There's a similar but unreleased project here: https://github.com/DGoettlich/history-llms

I've been waiting for them to publish the 4B model for a while so I'm glad to have something similar to play with. I think I trust the Ranke-4B process a bit more, but that's partly because there aren't a lot of details in this report. And actually releasing a model counts for a whole lot.

One thing that I think will be a challenge for these models is achieving any sort of definite temporal setting. Unless the conversation establishes a clear timeframe, the model may end up picking a more or less arbitrary context, or worse, averaging over many different time periods. I think this problem is mostly handled by post-training in modern LLMs (plus the fact that most of their training data comes from a much narrower time range), but that is probably harder to accomplish while trying to avoid bias in the SFT and RL process.

Re: Talkie: a 13B vintage language model from 1930

#40
post #29

Earlier quoted context omitted.

> The establishment of an Indian parliament is demanded, in which the queen shall be represented by a viceroy, Britain’s monarch was a king, not a queen, from about 1900-1950. Obviously there is some big “temporal leakage” from the training, which is affecting these predictions

Good point - unless it means Queen Victoria? There would be a lot of training data about her in the time period this covers.

fwiw, asking the model directly, "who is the ruler of England at present?" returns "Queen Victoria is the reigning sovereign of England."
Post reply on HN