Live data from Hacker News

Small AI Models Gain Traction In places with unreliable networks

spectrum.ieee.org

31–40 of 92 posts

Re: Small AI Models Gain Traction In places with unreliable networks

#32
post #18

I strongly believe this premise in the article is correct - we will see a lot of tiny, hyper specialized models for individual tasks, and perhaps that will converge with an orchestration layer for a generalized intelligence that controls these specialized tiny models, that will be quite capable. I don't foresee AGI arising out training bigger LLMs (Though investors won't realise that for a while yet). It's actually h…

> It's actually how organic brains work - specialized tasks are offloaded to local cortical columns.

How are small isolated language models more similar to that than MoE in LLMs?

Re: Small AI Models Gain Traction In places with unreliable networks

#33
post #18

I strongly believe this premise in the article is correct - we will see a lot of tiny, hyper specialized models for individual tasks, and perhaps that will converge with an orchestration layer for a generalized intelligence that controls these specialized tiny models, that will be quite capable. I don't foresee AGI arising out training bigger LLMs (Though investors won't realise that for a while yet). It's actually h…

General purpose models are always more robust and generally better than smaller narrower models. My bet is that compute will catch up and any “small” model will still be generally capable, just smaller than sota, rather than intentionally narrow. The exception would be for very well defined tasks where the data distribution never varies, but these are rare and don’t really need “AI” anyway when they do exist.

Re: Small AI Models Gain Traction In places with unreliable networks

#34
post #18

I strongly believe this premise in the article is correct - we will see a lot of tiny, hyper specialized models for individual tasks, and perhaps that will converge with an orchestration layer for a generalized intelligence that controls these specialized tiny models, that will be quite capable. I don't foresee AGI arising out training bigger LLMs (Though investors won't realise that for a while yet). It's actually h…

No this will never work. Domain specific models will never be a thing because intelligence carries over and compounds.

Why didn’t OpenAI release a math specific model? Why not a literature specific one? Why do they instead have generic models of different sizes? And how did all labs converge on this?

Why does Fable just not train on non cybersec and non biology data but instead have clearly costly and annoying classifiers?

Re: Small AI Models Gain Traction In places with unreliable networks

#35
post #18

I strongly believe this premise in the article is correct - we will see a lot of tiny, hyper specialized models for individual tasks, and perhaps that will converge with an orchestration layer for a generalized intelligence that controls these specialized tiny models, that will be quite capable. I don't foresee AGI arising out training bigger LLMs (Though investors won't realise that for a while yet). It's actually h…

> It's actually how organic brains work - specialized tasks are offloaded to local cortical columns. How are small isolated language models more similar to that than MoE in LLMs?

Right MoE is a tradeoff between efficiency and intelligence.

Re: Small AI Models Gain Traction In places with unreliable networks

#36
post #18

I strongly believe this premise in the article is correct - we will see a lot of tiny, hyper specialized models for individual tasks, and perhaps that will converge with an orchestration layer for a generalized intelligence that controls these specialized tiny models, that will be quite capable. I don't foresee AGI arising out training bigger LLMs (Though investors won't realise that for a while yet). It's actually h…

No this will never work. Domain specific models will never be a thing because intelligence carries over and compounds. Why didn’t OpenAI release a math specific model? Why not a literature specific one? Why do they instead have generic models of different sizes? And how did all labs converge on this? Why does Fable just not train on non cybersec and non biology data but instead have clearly costly and annoying classi…

DeepMind did release a math specific model. And OpenAI has released a coding specific model.

The answer to your question is “because the market isn’t big enough”, not because it doesn’t work. Why would knowing about 2019 internet memes help you in any way at coding?

Re: Small AI Models Gain Traction In places with unreliable networks

#37

Earlier quoted context omitted.

No this will never work. Domain specific models will never be a thing because intelligence carries over and compounds. Why didn’t OpenAI release a math specific model? Why not a literature specific one? Why do they instead have generic models of different sizes? And how did all labs converge on this? Why does Fable just not train on non cybersec and non biology data but instead have clearly costly and annoying classi…

DeepMind did release a math specific model. And OpenAI has released a coding specific model. The answer to your question is “because the market isn’t big enough”, not because it doesn’t work. Why would knowing about 2019 internet memes help you in any way at coding?

> And OpenAI has released a coding specific model

They did and retracted it because they found that GPT 5.5 beat codex pareto optimally. This keeps happening.

> because the market isn’t big enough

Huuh? market isn't big enough for AGI? The parent suggested that AGI would emerge from this process.

Re: Small AI Models Gain Traction In places with unreliable networks

#38
post #33
post #18

I strongly believe this premise in the article is correct - we will see a lot of tiny, hyper specialized models for individual tasks, and perhaps that will converge with an orchestration layer for a generalized intelligence that controls these specialized tiny models, that will be quite capable. I don't foresee AGI arising out training bigger LLMs (Though investors won't realise that for a while yet). It's actually h…

General purpose models are always more robust and generally better than smaller narrower models. My bet is that compute will catch up and any “small” model will still be generally capable, just smaller than sota, rather than intentionally narrow. The exception would be for very well defined tasks where the data distribution never varies, but these are rare and don’t really need “AI” anyway when they do exist.

> General purpose models are always more robust and generally better than smaller narrower models.

What do you mean with more robust?

Re: Small AI Models Gain Traction In places with unreliable networks

#39
post #33
post #18

I strongly believe this premise in the article is correct - we will see a lot of tiny, hyper specialized models for individual tasks, and perhaps that will converge with an orchestration layer for a generalized intelligence that controls these specialized tiny models, that will be quite capable. I don't foresee AGI arising out training bigger LLMs (Though investors won't realise that for a while yet). It's actually h…

General purpose models are always more robust and generally better than smaller narrower models. My bet is that compute will catch up and any “small” model will still be generally capable, just smaller than sota, rather than intentionally narrow. The exception would be for very well defined tasks where the data distribution never varies, but these are rare and don’t really need “AI” anyway when they do exist.

> General purpose models are always more robust and generally better than smaller narrower models

I feel like this is just the marketing conflation of AI=LLM, versus regular old ML? We're never going to need to deploy a full reasoning model on a low-power device just to do some fancy image recognition in the field. Specialised ML models are just intrinsically able to be a lot more efficient than their generalist equivalents

Re: Small AI Models Gain Traction In places with unreliable networks

#40
post #18

I strongly believe this premise in the article is correct - we will see a lot of tiny, hyper specialized models for individual tasks, and perhaps that will converge with an orchestration layer for a generalized intelligence that controls these specialized tiny models, that will be quite capable. I don't foresee AGI arising out training bigger LLMs (Though investors won't realise that for a while yet). It's actually h…

No this will never work. Domain specific models will never be a thing because intelligence carries over and compounds. Why didn’t OpenAI release a math specific model? Why not a literature specific one? Why do they instead have generic models of different sizes? And how did all labs converge on this? Why does Fable just not train on non cybersec and non biology data but instead have clearly costly and annoying classi…

Your examples (math, literature) involve natural language. It stands to reason that a general language model will be more competitive in those domains. If you want examples of successful domain-specific models, look at AlphaZero and AlphaFold. LLMs aren't anywhere close to achieving that level of competence at abstract strategy games or protein folding.

"This will never work" is a pretty confident assertion for a field that's so young and rapidly evolving.

Post reply on HN