Earlier quoted context omitted.
> a large language model trained on chess commentary as well as being trained on millions of chess games would outperform stockfish which is only trained on millions of chess games Not really, if anything it's closer to the opposite. The Bitter Lesson essay literally has this as an example: > These researchers wanted methods based on human input to win and were disappointed when they did not.[1] and > Enormous initia…
> These researchers wanted methods based on human input to win and were disappointed when they did not.[1] This was/is basically a strawman though. Like maybe "human input winning" was desirable for chess masters but for computer science wonks? Not the point or the disappoint. It's always neats and scruffies fighting about using some kind of recognizable method (logic) instead of magic (ML). > breakthrough progress e…
Small Models Have Arrived
321–330 of 373 posts
Re: Small Models Have Arrived
#322Re: Small Models Have Arrived
#323I love how "Small Models" apparently is "Model of unknown size but probably smaller than another model that we also don't know the size of".
Re: Small Models Have Arrived
#324Re: Small Models Have Arrived
#325> But I also think the demand for "fast/cheap/good-enough" models is just about to take off. There's a sort of "revelation" I had in ~early '24 when I used a 7B local model with a library called Guidance (initially out of MS, then the team moved) to create a flow where the model would receive pseudocode for tests, first write the tests, and once I approved then started writing code until the tests passed. This was be…
Yes. The infancy phase of this technology is represented by the pursuit of making wildly grand, wildly expensive, all-purpose models that somehow discern a user's full accurate intent from a lazy, underdeveloped, vague idea that they ambiguously and poorly express in a couple dozen words. The adolescence will arrive as those outsized and ill-considered ambitions collapse and we instead see a cambrian explosion of res…
Re: Small Models Have Arrived
#326Earlier quoted context omitted.
Yes LLMs are a beautiful way to compact knowledge. It would be such a cool technology to develop and worked with if it wasn’t linked to such a toxic industry
I think you're just observing ppl in one of these rare instances where enough of them come together because they are motivated. 'Toxic' is the clamoring sound of a crowded room where what gets through to your ears are just the most annoying snippets of incomplete conversations. I dare you to hang out with any actual people here, understand their viewpoint and listen to what they actually have to say in person, within…
- ingesting all the current knowledge without regard for intellectual property or the work of people that went into it; then
- claiming that AI would make all those people who put in the work redundant
It's not really surprising that when the sales pitch is "this will eliminate human creative work in all writing and illustration centric industries", people got angry.
Re: Small Models Have Arrived
#327> One thing a few investors I've talked with have mentioned: "It's weird we're not seeing more consumer AI companies. Why is that?" What would consumer AI company even be? The frontier labs have declared they will eat everything and they have a head start. Best bet would to be a contrarian and build products and services that people actually want or need. Fine to be AI powered or augmented, but consumer companies do…
And almost all consumer software has AI now. > But what if you want to add AI to your product? Well, now you have some real inference costs on every request! Eventually, these companies just lower their costs by using more efficient models. There were consumer companies built on GPT-3.
.. whether you want it or not!
We're at the point where it's becoming a product sticker like "fair trade" or "does not contain nuts" to say a product wasn't made with AI.
Re: Small Models Have Arrived
#328Earlier quoted context omitted.
Perhaps we'll get to a point where believing any un-sourced information from an LLM will feel crazy. I don't want my model to know more than it needs to perform logic and use tools. Once it is capable of using tools I would much rather it looked up information or sourced it from existing context rather than just divine it from it's weights.
The problem is it needs world knowledge to know what to lookup. This puts a floor on how little it can know while being able to look up what it doesn't know. Maybe its better if it knows a lot but has a good instinct for verifying that.
"Source: rare book ingested and shredded by Anthropic. No, you can't look it up and we can't show you the scan. The remaining open market copy is $5000. Trust me."
Re: Small Models Have Arrived
#329I’m kind of cautiously excited for the next five to ten years, with these AI chips becoming incredibly fast and RAM capacities ramping up its in the cards that we’ll have chips like today’s ATMEL microprocessors that fit on a single board computer and can run small models locally, then all our gizmos can have local AI and I can have a truly intelligent home. Of course there will be a huge push to put all of it in the…
I think hoping for a locally hosted "smart" home is backwards. They have no other function than to invade your space. "Smart" objects are agents, and they don't work for you.
Re: Small Models Have Arrived
#330Earlier quoted context omitted.
Yes. The infancy phase of this technology is represented by the pursuit of making wildly grand, wildly expensive, all-purpose models that somehow discern a user's full accurate intent from a lazy, underdeveloped, vague idea that they ambiguously and poorly express in a couple dozen words. The adolescence will arrive as those outsized and ill-considered ambitions collapse and we instead see a cambrian explosion of res…
it feels like if we had invented hammers, and we're still on the "make them bigger, stronger" phase, but we haven't even invented nails yet.