Live data from Hacker News

TimeCapsuleLLM: LLM trained only on data from 1800-1875

github.com

51–60 of 334 posts

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#51
post #14

Could this be an experiment to show how likely LLMs are to lead to AGI, or at least intelligence well beyond our current level? If you could only give it texts and info and concepts up to Year X, well before Discovery Y, could we then see if it could prompt its way to that discovery?

This would be a true test of can LLMs innovate or just regurgitate. I think part of people's amazement of LLMs is they don't realize how much they don't know. So thinking and recalling look the same to the end user.

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#52
post #28
post #14

Could this be an experiment to show how likely LLMs are to lead to AGI, or at least intelligence well beyond our current level? If you could only give it texts and info and concepts up to Year X, well before Discovery Y, could we then see if it could prompt its way to that discovery?

> Could this be an experiment to show how likely LLMs are to lead to AGI, or at least intelligence well beyond our current level? You'd have to be specific what you mean by AGI: all three letters mean a different thing to different people, and sometimes use the whole means something not present in the letters. > If you could only give it texts and info and concepts up to Year X, well before Discovery Y, could we then…

Basically looking for emergent behavior.

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#53

Would be interesting to train a cutting edge model with a cut off date of say 1900 and then prompt it about QM and relativity with some added context. If the model comes up with anything even remotely correct it would be quite a strong evidence that LLMs are a path to something bigger if not then I think it is time to go back to the drawing board.

You would find things in there that were already close to QM and relativity. The Michelson-Morley experiment was 1887 and Lorentz transformations came along in 1889. The photoelectric effect (which Einstein explained in terms of photons in 1905) was also discovered in 1887. William Clifford (who _died_ in 1889) had notions that foreshadowed general relativity: "Riemann, and more specifically Clifford, conjectured tha…

I presume that's what the parent post is trying to get at? Seeing if, given the cutting edge scientific knowledge of the day, the LLM is able to synthesis all it into a workable theory of QM by making the necessary connections and (quantum...) leaps

Standing on the shoulders of giants, as it were

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#56

Would be interesting to train a cutting edge model with a cut off date of say 1900 and then prompt it about QM and relativity with some added context. If the model comes up with anything even remotely correct it would be quite a strong evidence that LLMs are a path to something bigger if not then I think it is time to go back to the drawing board.

You would find things in there that were already close to QM and relativity. The Michelson-Morley experiment was 1887 and Lorentz transformations came along in 1889. The photoelectric effect (which Einstein explained in terms of photons in 1905) was also discovered in 1887. William Clifford (who _died_ in 1889) had notions that foreshadowed general relativity: "Riemann, and more specifically Clifford, conjectured tha…

This would still be valuable even if the LLM only finds out about things that are already in the air.

It’s probably even more of a problem that different areas of scientific development don’t know about each other. LLMs combining results would still not be like they invented something new.

But if they could give us a head start of 20 years on certain developments this would be an awesome result.

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#57
Can you confidently say that the architure of the LLM doesn't include any a priori bias that might effect the integrity of this LLM?

That is, the architectures of today are chosen to yield the best results given the textual data around today and the problems we want to solve today.

I'd argue that this lack of bias would need to be researched (if it hasn't been already) before this kind of model has credence.

LLMs aren't my area of expertise but during my PhD we were able to encode a lot of a priori knowledge through the design of neural network architectures.

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#58
post #28
post #14

Could this be an experiment to show how likely LLMs are to lead to AGI, or at least intelligence well beyond our current level? If you could only give it texts and info and concepts up to Year X, well before Discovery Y, could we then see if it could prompt its way to that discovery?

> Could this be an experiment to show how likely LLMs are to lead to AGI, or at least intelligence well beyond our current level? You'd have to be specific what you mean by AGI: all three letters mean a different thing to different people, and sometimes use the whole means something not present in the letters. > If you could only give it texts and info and concepts up to Year X, well before Discovery Y, could we then…

An example of why you need to explain what you mean by AGI is:

https://www.robinsloan.com/winter-garden/agi-is-here/

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#59

Earlier quoted context omitted.

Are you saying it wouldn't be able to converse using english of the time?

That's not what they are saying. SOTA models include much more than just language, and the scale of training data is related to its "intelligence". Restricting the corpus in time => less training data => less intelligence => less ability to "discover" new concepts not in its training data

Perhaps less bullshit though was my thought? Was language more restricted then? Scope of ideas?

Re: TimeCapsuleLLM: LLM trained only on data from 1800-1875

#60

Earlier quoted context omitted.

I fail to see how the two concepts equate. LLMs have neither intelligence nor problem-solving abillity (and I won't be relaxing the definition of either so that some AI bro can pretend a glorified chatbot is sentient) You would, at best, be demonstrating that the sharing of knowledge across multiple disciplines and nations (which is a relatively new concept - at least at the scale of something like the internet) lead…

I've seen many futurists claim that human innovation is dead and all future discoveries will be the results of AI. If this is true, we should be able to see AI trained on the past figure it's way to various things we have today. If it can't do this, I'd like said futurists to quiet down, as they are discouraging an entire generation of kids who may go on to discover some great things.

> I've seen many futurists claim that human innovation is dead and all future discoveries will be the results of AI.

I think there's a big difference between discoveries through AI-human synergy and discoveries through AI working in isolation.

It probably will be true soon (if it isn't already) that most innovation features some degree of AI input, but still with a human to steer the AI in the right direction.

I think an AI being able to discover something genuinely new all by itself, without any human steering, is a lot further off.

If AIs start producing significant quantities of genuine and useful innovation with minimal human input, maybe the singularitarians are about to be proven right.

Post reply on HN