Live data from Hacker News

Small Models Have Arrived

calv.info

321–330 of 375 posts

Re: Small Models Have Arrived

#321
post #218

Earlier quoted context omitted.

> a large language model trained on chess commentary as well as being trained on millions of chess games would outperform stockfish which is only trained on millions of chess games Not really, if anything it's closer to the opposite. The Bitter Lesson essay literally has this as an example: > These researchers wanted methods based on human input to win and were disappointed when they did not.[1] and > Enormous initia…

> These researchers wanted methods based on human input to win and were disappointed when they did not.[1] This was/is basically a strawman though. Like maybe "human input winning" was desirable for chess masters but for computer science wonks? Not the point or the disappoint. It's always neats and scruffies fighting about using some kind of recognizable method (logic) instead of magic (ML). > breakthrough progress e…

There's two different goals to AI research - one was to get results - a chess engine thst wins, etc. But the other goal (which seems to have been abandoned in the deep learning era) was to use AI to help understand how human minds work. A chess engine modeled after human grandmasters is much more interesting in that regard than either a min-max algorithm like beat Kasparov or modern deep learning engines.

Re: Small Models Have Arrived

#323

I love how "Small Models" apparently is "Model of unknown size but probably smaller than another model that we also don't know the size of".

I was wondering the same, trying to find if OpenAI has started to release the ampumt of model parametres. Basically they seem to mean ”cheaper”.

Re: Small Models Have Arrived

#325

> But I also think the demand for "fast/cheap/good-enough" models is just about to take off. There's a sort of "revelation" I had in ~early '24 when I used a 7B local model with a library called Guidance (initially out of MS, then the team moved) to create a flow where the model would receive pseudocode for tests, first write the tests, and once I approved then started writing code until the tests passed. This was be…

Yes. The infancy phase of this technology is represented by the pursuit of making wildly grand, wildly expensive, all-purpose models that somehow discern a user's full accurate intent from a lazy, underdeveloped, vague idea that they ambiguously and poorly express in a couple dozen words. The adolescence will arrive as those outsized and ill-considered ambitions collapse and we instead see a cambrian explosion of res…

it feels like if we had invented hammers, and we're still on the "make them bigger, stronger" phase, but we haven't even invented nails yet.

Re: Small Models Have Arrived

#326
post #70

Earlier quoted context omitted.

Yes LLMs are a beautiful way to compact knowledge. It would be such a cool technology to develop and worked with if it wasn’t linked to such a toxic industry

I think you're just observing ppl in one of these rare instances where enough of them come together because they are motivated. 'Toxic' is the clamoring sound of a crowded room where what gets through to your ears are just the most annoying snippets of incomplete conversations. I dare you to hang out with any actual people here, understand their viewpoint and listen to what they actually have to say in person, within…

The "toxicity" was:

- ingesting all the current knowledge without regard for intellectual property or the work of people that went into it; then

- claiming that AI would make all those people who put in the work redundant

It's not really surprising that when the sales pitch is "this will eliminate human creative work in all writing and illustration centric industries", people got angry.

Re: Small Models Have Arrived

#327

> One thing a few investors I've talked with have mentioned: "It's weird we're not seeing more consumer AI companies. Why is that?" What would consumer AI company even be? The frontier labs have declared they will eat everything and they have a head start. Best bet would to be a contrarian and build products and services that people actually want or need. Fine to be AI powered or augmented, but consumer companies do…

And almost all consumer software has AI now. > But what if you want to add AI to your product? Well, now you have some real inference costs on every request! Eventually, these companies just lower their costs by using more efficient models. There were consumer companies built on GPT-3.

> almost all consumer software has AI now

.. whether you want it or not!

We're at the point where it's becoming a product sticker like "fair trade" or "does not contain nuts" to say a product wasn't made with AI.

Re: Small Models Have Arrived

#328

Earlier quoted context omitted.

Perhaps we'll get to a point where believing any un-sourced information from an LLM will feel crazy. I don't want my model to know more than it needs to perform logic and use tools. Once it is capable of using tools I would much rather it looked up information or sourced it from existing context rather than just divine it from it's weights.

The problem is it needs world knowledge to know what to lookup. This puts a floor on how little it can know while being able to look up what it doesn't know. Maybe its better if it knows a lot but has a good instinct for verifying that.

How do you actually get to the thing you're looking up? The scrapeable internet is shrinking in response to scrapers.

"Source: rare book ingested and shredded by Anthropic. No, you can't look it up and we can't show you the scan. The remaining open market copy is $5000. Trust me."

Re: Small Models Have Arrived

#329

I’m kind of cautiously excited for the next five to ten years, with these AI chips becoming incredibly fast and RAM capacities ramping up its in the cards that we’ll have chips like today’s ATMEL microprocessors that fit on a single board computer and can run small models locally, then all our gizmos can have local AI and I can have a truly intelligent home. Of course there will be a huge push to put all of it in the…

Is a "truly intelligent home" something I should want? None of the current generation of "smart" addons are true value adds - their entire purpose is data collection. The pretext is always absurdly thin. I just thought "I bet there's a wifi enabled microwave", googled it, and indeed, Samsung have released such a thing - you can control it with your voice! Wow! Never mind that you can't un/load food with your voice, or that you're never more than 10 feet from it in the kitchen anyway. Oh but there's an app. Of course there's an app. It "suggests recipes". Right.

I think hoping for a locally hosted "smart" home is backwards. They have no other function than to invade your space. "Smart" objects are agents, and they don't work for you.

Re: Small Models Have Arrived

#330

Earlier quoted context omitted.

Yes. The infancy phase of this technology is represented by the pursuit of making wildly grand, wildly expensive, all-purpose models that somehow discern a user's full accurate intent from a lazy, underdeveloped, vague idea that they ambiguously and poorly express in a couple dozen words. The adolescence will arrive as those outsized and ill-considered ambitions collapse and we instead see a cambrian explosion of res…

it feels like if we had invented hammers, and we're still on the "make them bigger, stronger" phase, but we haven't even invented nails yet.

For large language models, isn't everything nails?
Post reply on HN