Earlier quoted context omitted.
> Imagine this, an AI able to operate some specific API in a deterministic / reliable way. It doesn’t seem like LLMs are going to be able to do this, unless the application has a high tolerance for mistakes.
Maybe something new needs to be invented? Or utilize something we already have? Anyway I think this is the next leap in AI. Operative Language Models. Where the model is trained to do specific tasks very well. In a reliable way. I don't anticipate them to be 'large' and expensive also. Then we have lots of these models and some kind of 'orchestration' layer that makes them all work together. This I believe will be th…
How to train your own large language models
31–40 of 63 posts
Re: How to train your own large language models
#32Earlier quoted context omitted.
Maybe something new needs to be invented? Or utilize something we already have? Anyway I think this is the next leap in AI. Operative Language Models. Where the model is trained to do specific tasks very well. In a reliable way. I don't anticipate them to be 'large' and expensive also. Then we have lots of these models and some kind of 'orchestration' layer that makes them all work together. This I believe will be th…
The next leap will be decided by what someone is able to actually implement, not by what anyone thinks it should be.
Re: How to train your own large language models
#33Re: How to train your own large language models
#34Did we ever get any resolution about what happened after this company threatened to sue their intern for making a side project that supposedly stole all their great ideas? I would like to know before I ever consider anything from them again.
The founder admitted his mistake and the ex-intern's site is back up and running https://riju.codes/ . I'm personally a fan of both Amjad's (CEO) and Radon's (intern) and realize that everyone makes mistakes. It's not a reason to discount the hard work of the people at replit.
I just want a URL in which I can run some code. https://riju.codes/ is literally that. Thanks!
Re: How to train your own large language models
#35Did we ever get any resolution about what happened after this company threatened to sue their intern for making a side project that supposedly stole all their great ideas? I would like to know before I ever consider anything from them again.
Re: How to train your own large language models
#36Earlier quoted context omitted.
Story for those who didn't see it: https://intuitiveexplanations.com/tech/replit/
That's weird. I would never do anything even remotely similar to what my (ex) employer does. CEO sounds like a douchebag tho.
Re: How to train your own large language models
#37Ghostwriter is notably worse than GPT-4, so while it may be true in a sense that "Training a custom model allows us to tailor it to our specific needs and requirements", the reality is they'd be getting better results just using OpenAI right now. Probably true for almost every other use case. That said, I am patiently waiting and champing at the bit for the day this isn't true anymore. Cool to see the groundwork bein…
I think LLMs could end up the same way, if the comminity consolidates around a good one.
Re: How to train your own large language models
#38How expensive is it? My understanding is that it's not reasonable to train an LLM from scratch by yourself, and that if you want one that isn't just very stupid then you need to spend between hundreds of thousands and hundreds of millions of dollars. But if you don't want to train from scratch then you can fine-tune existing models for cheaper.
Training these models from scratch on your domain specific data is not as expensive as one might think. We have provided some cost estimates in our blogs.
https://www.mosaicml.com/blog/mosaicbert
https://www.mosaicml.com/blog/training-stable-diffusion-from...
Re: How to train your own large language models
#39Earlier quoted context omitted.
The founder admitted his mistake and the ex-intern's site is back up and running https://riju.codes/ . I'm personally a fan of both Amjad's (CEO) and Radon's (intern) and realize that everyone makes mistakes. It's not a reason to discount the hard work of the people at replit.
That’s a very generous interpretation of what happened because it wasn’t a “mistake” when he threatened the intern, it was something he purposefully and intentionally did, and doubled down on, even after having significant time to reconsider. Only when there was widespread public criticism of his actions did he backpedal. I’m curious what he’s said or done to make you a fan?
> an action or judgment that is misguided or wrong.
So to admit that you made a mistake just means you were "misguided or wrong" which he definitely made clear he was. You're claiming that there was significant time to reconsider but the reality is that this all went down in a matter of hours from when the intern published his article. Sometimes people have bad judgement and the public and especially one's peers can help to see the error of their ways and improve for the better. If he was a repeat offender and this happened several times then I could totally understand it but there's no reason to keep bringing this up every single time anything related to Replit is mentioned on HN.
Re: How to train your own large language models
#40Did we ever get any resolution about what happened after this company threatened to sue their intern for making a side project that supposedly stole all their great ideas? I would like to know before I ever consider anything from them again.
Shh, another post on the home page is about them hiring (YC W18), don’t interfere with the business model!