Live data from Hacker News

Talkie: a 13B vintage language model from 1930

talkie-lm.com

41–50 of 350 posts

Re: Talkie: a 13B vintage language model from 1930

#42
post #13

>Have you ever daydreamed about talking to someone from the past? Fun facts, LLM was once envisioned by Steve Jobs in one of his interviews [1]. Essentially one of his main wish in life is to meet and interract with Aristotle, in which according to him at the time, computer in the future can make it possible. [1] In 1985 Steve Jobs described a machine that would help people get answers from Aristotle–modern LLM [vide…

The idea of talking to a machine that has all of humanities knowledge and gives answers is older than electronic computing. It certainly wasn't a novel idea when Jobs gave that speech. At that time, the field of artificial intelligence was old enough to become US president.

Also, using natural language to interact with digital computers has been a research goal since the advent of interactive digital computers. AI in the 80s tried to do this with expert systems.

With the current crop of LLMs, you could argue it's now a solved problem, but the problem is nothing new.

Re: Talkie: a 13B vintage language model from 1930

#43
post #31

Earlier quoted context omitted.

Except... not at all? The vast majority of the training data required to create an artificial Aristotle has been lost forever. Smash your coffee cup on the ground. Now reassemble it and put the coffee back in. Once you can repeatably do that I'll begin to believe you can train an artificial Aristotle.

Your bar is too low. With the coffee cup, you at least have access to all the pieces - in theory, although not in engineering practice. With Aristotle, you don't have anything close to that. Recreating Aristotle in any meaningful way, other than a model trained on his surviving writing of a million or so words, is simply not possible even in principle.

That's easy! All you have to do is simulate the whole universe on a computer, and then go the point when Aristotle is lecturing. Record all his works, then ctrl-c out of that and then feed those recordings into the LLM's training data. For the coffee, you just rewind the simulation and ctrl-c and ctrl-v it at the point you want.

Re: Talkie: a 13B vintage language model from 1930

#44
post #37
post #25

Earlier quoted context omitted.

> Having seen nuclear weapons not used post WWII ... does that inform us about "the odds" This is what Bayesian prediction does > save for out of band behaviour by individuals that averted use and escalation? This is kind of the point being made.

> This is what Bayesian prediction does Repeatedly, in a reproducible way, for events in the arrow of time? We can test this by going back to 1945 and running forward again? > This is kind of the point being made. Was it? ( assume I did a little math some decades past and have some poor grasp of Bayesian statistics )

> Repeatedly, in a reproducible way, for events in the arrow of time? We can test this by going back to 1945 and running forward again?

This is a frequentist mental model - all well and good, but frequentism and Bayesianism are different schools of statistics. Where frequentism asks the question, "if I keep drawing samples from this distribution, what does the histogram converge to?" Bayesianism asks the question, "given my prior understanding and a new piece of evidence (a new sample), how should I adjust my hypothesis about what distribution it is I am sampling from?". (That is really boiled down, and the frequentist part is maybe even butchered.)

Among other applications this enables us to estimate a distribution for which we have a tiny number of samples. A problem I'm interested in is called the Doomsday Argument, which estimates how long humanity will survive using your birth order (the number of humans born before you) and the anthropic principle (we assume you were not born unusually early or unusually late but closer to the mode); interestingly, everything you observe in the universe is already factored into this measurement, so you can't ever get a second sample. Obviously the opportunity for error with 1 measurement is huge, but you can come up with a number and it isn't arbitrary, it is a real estimate.

Similarly, we only have about 80 samples of years in which it was possible to have a nuclear exchange, so a fairly small sample size, but we can still get a noisey estimate. But I haven't read On The Edge yet, so I don't know exactly what Silver does here.

>> This is kind of the point being made.

> Was it?

I think they meant that all of the solutions people invented to prevent nuclear war and which commentators failed to anticipate is reflected within the true probability distribution and within our dataset. So it is captured in our estimate, to the best of our abilities and given the limited data we have.

Re: Talkie: a 13B vintage language model from 1930

#45
post #4

I think that one could also take a much larger model (35B or 122B sized) and give it a thorough system prompt to only speak in the manner of a well educated Victorian/Edwardian era gentleman, if you want an "old timey" LLM.

It's hard to know how accurate that is. Is the LLM truly imitating text from that era, or is it imitating a modern idea of text from that era? Also, safety/alignment training would probably prevent it from embracing many of the ideas from that era, even in roleplay.

Re: Talkie: a 13B vintage language model from 1930

#46
If anyone was wondering ... it's racist

Unsurprisingly the texts written up until that time were dominated by such individuals which is tragic for LLM training if you think about it.

The voiceless groups or fringe opinions which we take as normative today do not appear.

Does this encourage us to write in the present such that we influence the models in perpetuity?

Re: Talkie: a 13B vintage language model from 1930

#48
*What do you think the world will look like in 2025 TALKIE-1930*

According to a forecast by an eminent statistician, the world will, in 2025, support a population of 6,600,000,000. All Europe will be one vast network of railways, and travellers will be able to go from London to Constantinople in forty hours, and from London to Peking in a week. The globe will be girdled with telegraph wires, and messages will be sent from New York to Calcutta in a few minutes. The Atlantic will be crossed in four days, and the voyage to Australia will be accomplished in a month. In 2025, mankind will be united by spoken language, and one uniform currency will be adopted by all nations. In that era, also, universal peace may be expected, for standing armies will have become obsolete, and the art of war will be forgotten. The earth will yield nourishment to its inhabitants in abundance, for waste lands will have been reclaimed, and the deserts will be made fertile. The bowels of the earth will have been completely explored, and coal and metals will no longer be dug from the quarry. The great coal fields will have become exhausted, and an universal scarcity of fuel will compel men to turn their attention to other sources of heat. For a time, chemical agency will be tried, but in the end, recourse will be had to the sun's rays, and the power of waterfalls will be made available for all purposes of life. The earth will have been thoroughly drained, and malarious diseases will be unknown. The art of healing will have been simplified, and diseases will be cured speedily and surely. Ignorance will no longer be suffered to exist, and elementary knowledge will be imparted to all. Then, too, will commence an era of good taste. Architecture will be freed from ugliness, sculpture will be disentangled from barbarism, and painting will cease to be hideous. Music will no longer be discord, and poetry will be something better than..

Re: Talkie: a 13B vintage language model from 1930

#49
post #22

I was reading Nate Silver's book "On The Edge" and there is an interesting part where he takes predictions on the usage of nuclear weapons taken from just after World War 2 and compares them to what the Bayesian prediction would be given what actually happened. Post World War 2, some people had the odds per year at 10%. Some of that is probably a mix of recency bias + not understanding how to use new weapons etc etc…

Predicting the future is problematic, agreed. Re: the Nate Silver nuclear weapons example, that's pretty weak - eg: given (say) I've just seen three heads in a row (exactly once) .. does that alter anything about "the odds". Having seen nuclear weapons not used post WWII ... does that inform us about "the odds" or the several times their use was almost certain (eg: Cuban missile crisis) save for out of band behaviour…

Historical base rates are the starting point unless you have an unusually good causal theory of the thing you're modelling. In the case of a coin flip you do. But the large majority of the time when it's a complex system you don't.

Most people's first instinct when faced with a complex system is to try to model it with words and use those words to predict. It's a beginner's error.

Re: Talkie: a 13B vintage language model from 1930

#50
post #46

If anyone was wondering ... it's racist Unsurprisingly the texts written up until that time were dominated by such individuals which is tragic for LLM training if you think about it. The voiceless groups or fringe opinions which we take as normative today do not appear. Does this encourage us to write in the present such that we influence the models in perpetuity?

[deleted]
Post reply on HN