Socratic Models – Composing Zero-Shot Multimodal Reasoning with Language
socraticmodels.github.io
Socratic Models – Composing Zero-Shot Multimodal Reasoning with Language
1–10 of 40 posts
Re: Socratic Models – Composing Zero-Shot Multimodal Reasoning with Language
#2I think the next step in NLP will be a drastic innovation on today's learning model.
Re: Socratic Models – Composing Zero-Shot Multimodal Reasoning with Language
#3We've come to the consensus that large language models are just stochastic parrots... What makes us think that we can achieve a higher level of intelligence by putting them in conversation? I think the next step in NLP will be a drastic innovation on today's learning model.
The Socratic paper is not about “higher intelligence”, it’s about demonstrating useful behaviour purely by connecting several large models via language.
Re: Socratic Models – Composing Zero-Shot Multimodal Reasoning with Language
#4I still hold the opinion that we’re going to need to move to spiking neuron (SNN) models in the future to keep growing the networks. Spiking networks require lots of storage, but a lot, lot less compute. They also propagate additional information in the _timing_ of the spikes, not just the values. There are a lot of low-hanging fruit in SNNs and I think people are still trying to copy biological systems too much.
Unfortunately, the main issue with SNNs is that no one has figured out a way to train them as effectively as ANNs.
Re: Socratic Models – Composing Zero-Shot Multimodal Reasoning with Language
#5We've come to the consensus that large language models are just stochastic parrots... What makes us think that we can achieve a higher level of intelligence by putting them in conversation? I think the next step in NLP will be a drastic innovation on today's learning model.
Re: Socratic Models – Composing Zero-Shot Multimodal Reasoning with Language
#6This is super impressive. Transformers have consistently done better than almost anyone thought. I still hold the opinion that we’re going to need to move to spiking neuron (SNN) models in the future to keep growing the networks. Spiking networks require lots of storage, but a lot, lot less compute. They also propagate additional information in the _timing_ of the spikes, not just the values. There are a lot of low-h…
Is this fundamental, or just a problem with mapping these models to our current serially-bottlenecked compute architectures? Could a move to “hyperconverged infrastructure in-the-small” — striping DRAM or NVMe and tiny RISC cores together on a die, where each CPU gets its own storage (or, you might say, where each small cluster of storage cells has its own tiny CPU attached), such that one stick has millions of independent+concurrent [+slow+memory-constrained] processors — resolve these difficulties?
Re: Socratic Models – Composing Zero-Shot Multimodal Reasoning with Language
#7This is super impressive. Transformers have consistently done better than almost anyone thought. I still hold the opinion that we’re going to need to move to spiking neuron (SNN) models in the future to keep growing the networks. Spiking networks require lots of storage, but a lot, lot less compute. They also propagate additional information in the _timing_ of the spikes, not just the values. There are a lot of low-h…
> a lot of storage Is this fundamental, or just a problem with mapping these models to our current serially-bottlenecked compute architectures? Could a move to “hyperconverged infrastructure in-the-small” — striping DRAM or NVMe and tiny RISC cores together on a die, where each CPU gets its own storage (or, you might say, where each small cluster of storage cells has its own tiny CPU attached), such that one stick ha…
Re: Socratic Models – Composing Zero-Shot Multimodal Reasoning with Language
#8This is super impressive. Transformers have consistently done better than almost anyone thought. I still hold the opinion that we’re going to need to move to spiking neuron (SNN) models in the future to keep growing the networks. Spiking networks require lots of storage, but a lot, lot less compute. They also propagate additional information in the _timing_ of the spikes, not just the values. There are a lot of low-h…
As someone just trying to learn more about the implications of new research, I find myself resorting to /r/machinelearning, or even twitter threads, to get timely and informed discussions. That's a shame, given what HN sets out to be.
Re: Socratic Models – Composing Zero-Shot Multimodal Reasoning with Language
#9This is super impressive. Transformers have consistently done better than almost anyone thought. I still hold the opinion that we’re going to need to move to spiking neuron (SNN) models in the future to keep growing the networks. Spiking networks require lots of storage, but a lot, lot less compute. They also propagate additional information in the _timing_ of the spikes, not just the values. There are a lot of low-h…
The comments of every ML paper posted on this site are dominated by people either baselessly discounting the results as a party trick or illusion, or shoehorning in their conjecture about what approach the field is overlooking. As someone just trying to learn more about the implications of new research, I find myself resorting to /r/machinelearning, or even twitter threads, to get timely and informed discussions. Tha…
Re: Socratic Models – Composing Zero-Shot Multimodal Reasoning with Language
#10Earlier quoted context omitted.
The comments of every ML paper posted on this site are dominated by people either baselessly discounting the results as a party trick or illusion, or shoehorning in their conjecture about what approach the field is overlooking. As someone just trying to learn more about the implications of new research, I find myself resorting to /r/machinelearning, or even twitter threads, to get timely and informed discussions. Tha…
I’m certainly not discounting the results and I don’t see anything wrong with suggesting what I think would generally be a good path to look at in the future.
Maybe I'm expecting too much of HN, but I've seen these same two top level comments under myriad ML posts.
Sorry for the meta-discussion that's gotten us further away from this really remarkable paper.