Live data from Hacker News

Small Models Have Arrived

calv.info

11–20 of 371 posts

Re: Small Models Have Arrived

#12
Maybe I'm being super reductive here, but operating small models at the core of your business kind of moves the needle from making external API calls (against frontier models) to running internal API calls (against your locally-run models). It seems like if we want local models to take off, it will need to become easier to run local models for cheap. I'm thinking like reducing the barrier of entry for running "local models" in the cloud providers like DigitalOcean, AWS, etc.

Re: Small Models Have Arrived

#13
> Across his various startups, Peter has seen two kinds of work:

> 1. the "IQ 180" work. some mad scientist genius type comes up with some crazy solution you've never thought of.

> 2. the "token spewer" work. being ultra responsive, pushing the ball forward across dozens of different fronts.

Interesting comp to pg's Maker's Schedule, Manager's Schedule https://www.paulgraham.com/makersschedule.html

I'm curious about not only which of these roles models will fill, but also how they will empower us to be in the mode we prefer.

Re: Small Models Have Arrived

#14
It makes sense that we’ll see “room at the bottom” strategies. Currently, large parameter counts seem to be slush funds of world knowledge, language skills (because language’s nuances and open vocabulary make it high-dimensional), and reasoning primitives, the general belief being that the latter takes up the least space in the model.

There are many applications where world knowledge is unnecessary or even a negative, and in which only a small amount of language skill is necessary, and there we can expect small models more intelligently used to beat large ones naively used.

Re: Small Models Have Arrived

#15
post #12

Maybe I'm being super reductive here, but operating small models at the core of your business kind of moves the needle from making external API calls (against frontier models) to running internal API calls (against your locally-run models). It seems like if we want local models to take off, it will need to become easier to run local models for cheap. I'm thinking like reducing the barrier of entry for running "local…

You should be glad to know digital ocean already offers this

Re: Small Models Have Arrived

#16

It makes sense that we’ll see “room at the bottom” strategies. Currently, large parameter counts seem to be slush funds of world knowledge, language skills (because language’s nuances and open vocabulary make it high-dimensional), and reasoning primitives, the general belief being that the latter takes up the least space in the model. There are many applications where world knowledge is unnecessary or even a negative…

Small amounts of world knowledge seems like it would inherently be tied to more hallucinations.

Re: Small Models Have Arrived

#17
100% agreed. Small, cheap, and hosted models. Luna (and open weight models and others) is ridiculously cheap @ $0.2/$1.2, easily accessible, and more than good enough for basic use cases (e.g. summarization, simple tool calling, etc.).

Re: Small Models Have Arrived

#18
I have trouble seeing the points of using less capable models.

I just want the smartest, best, and most capable models. It feels smaller models for speed and cost are just transitions towards better hardware allowing the very best model.

Re: Small Models Have Arrived

#19

I have trouble seeing the points of using less capable models. I just want the smartest, best, and most capable models. It feels smaller models for speed and cost are just transitions towards better hardware allowing the very best model.

First, smaller models are fun for hackers: you can run them locally, or run them faster.

Second, when cloud models become unavailable or otherwise deteriorate, these will be all you have. May as well prepare.

Re: Small Models Have Arrived

#20

I have trouble seeing the points of using less capable models. I just want the smartest, best, and most capable models. It feels smaller models for speed and cost are just transitions towards better hardware allowing the very best model.

Can you define your use of the word 'smartest' here just in case some of us don't quite know what you mean?
Post reply on HN