Live data from Hacker News

How to train your own large language models

blog.replit.com

31–40 of 63 posts

Re: How to train your own large language models

#31
post #28

Earlier quoted context omitted.

> Imagine this, an AI able to operate some specific API in a deterministic / reliable way. It doesn’t seem like LLMs are going to be able to do this, unless the application has a high tolerance for mistakes.

Maybe something new needs to be invented? Or utilize something we already have? Anyway I think this is the next leap in AI. Operative Language Models. Where the model is trained to do specific tasks very well. In a reliable way. I don't anticipate them to be 'large' and expensive also. Then we have lots of these models and some kind of 'orchestration' layer that makes them all work together. This I believe will be th…

The next leap will be decided by what someone is able to actually implement, not by what anyone thinks it should be.

Re: How to train your own large language models

#32
post #31

Earlier quoted context omitted.

Maybe something new needs to be invented? Or utilize something we already have? Anyway I think this is the next leap in AI. Operative Language Models. Where the model is trained to do specific tasks very well. In a reliable way. I don't anticipate them to be 'large' and expensive also. Then we have lots of these models and some kind of 'orchestration' layer that makes them all work together. This I believe will be th…

The next leap will be decided by what someone is able to actually implement, not by what anyone thinks it should be.

Yes! This is where I'm at!

Re: How to train your own large language models

#34
post #2

Did we ever get any resolution about what happened after this company threatened to sue their intern for making a side project that supposedly stole all their great ideas? I would like to know before I ever consider anything from them again.

The founder admitted his mistake and the ex-intern's site is back up and running https://riju.codes/ . I'm personally a fan of both Amjad's (CEO) and Radon's (intern) and realize that everyone makes mistakes. It's not a reason to discount the hard work of the people at replit.

Thanks for linking this. This is actually a superior offering to replit. They recently removed the ability to access a simple repl without logging in. Now you a) have to login and b) have to deal with this obtuse IDE-in-a-browser project creation shit. It's so many extra steps before I can run code.

I just want a URL in which I can run some code. https://riju.codes/ is literally that. Thanks!

Re: How to train your own large language models

#35
post #2

Did we ever get any resolution about what happened after this company threatened to sue their intern for making a side project that supposedly stole all their great ideas? I would like to know before I ever consider anything from them again.

This CEO also has a history of punching down on Twitter. A very bad look.

Re: How to train your own large language models

#36
post #5

Earlier quoted context omitted.

Story for those who didn't see it: https://intuitiveexplanations.com/tech/replit/

That's weird. I would never do anything even remotely similar to what my (ex) employer does. CEO sounds like a douchebag tho.

I've seen terms/clauses here in AU for full time employment, depending on the industry/niche, where you can't jump to the same industry within X months.

Re: How to train your own large language models

#37
post #7

Ghostwriter is notably worse than GPT-4, so while it may be true in a sense that "Training a custom model allows us to tailor it to our specific needs and requirements", the reality is they'd be getting better results just using OpenAI right now. Probably true for almost every other use case. That said, I am patiently waiting and champing at the bit for the day this isn't true anymore. Cool to see the groundwork bein…

Stable Diffusion 1.5 is not SOTA, but in reality the sea of augmentations makes SD kinda unbeatable, if you are willing to put in the work to use them.

I think LLMs could end up the same way, if the comminity consolidates around a good one.

Re: How to train your own large language models

#38
post #4

How expensive is it? My understanding is that it's not reasonable to train an LLM from scratch by yourself, and that if you want one that isn't just very stupid then you need to spend between hundreds of thousands and hundreds of millions of dollars. But if you don't want to train from scratch then you can fine-tune existing models for cheaper.

Disclaimer: I work for MosaicML (MosaicML is the creator of the training platform used by Replit).

Training these models from scratch on your domain specific data is not as expensive as one might think. We have provided some cost estimates in our blogs.

https://www.mosaicml.com/blog/mosaicbert

https://www.mosaicml.com/blog/training-stable-diffusion-from...

https://www.mosaicml.com/blog/gpt-3-quality-for-500k

Re: How to train your own large language models

#39

Earlier quoted context omitted.

The founder admitted his mistake and the ex-intern's site is back up and running https://riju.codes/ . I'm personally a fan of both Amjad's (CEO) and Radon's (intern) and realize that everyone makes mistakes. It's not a reason to discount the hard work of the people at replit.

That’s a very generous interpretation of what happened because it wasn’t a “mistake” when he threatened the intern, it was something he purposefully and intentionally did, and doubled down on, even after having significant time to reconsider. Only when there was widespread public criticism of his actions did he backpedal. I’m curious what he’s said or done to make you a fan?

Seems like you're just arguing about the definition of the word "mistake". Intent has nothing to do with it. From Google (Oxford Dictionary):

> an action or judgment that is misguided or wrong.

So to admit that you made a mistake just means you were "misguided or wrong" which he definitely made clear he was. You're claiming that there was significant time to reconsider but the reality is that this all went down in a matter of hours from when the intern published his article. Sometimes people have bad judgement and the public and especially one's peers can help to see the error of their ways and improve for the better. If he was a repeat offender and this happened several times then I could totally understand it but there's no reason to keep bringing this up every single time anything related to Replit is mentioned on HN.

Re: How to train your own large language models

#40
post #29
post #2

Did we ever get any resolution about what happened after this company threatened to sue their intern for making a side project that supposedly stole all their great ideas? I would like to know before I ever consider anything from them again.

Shh, another post on the home page is about them hiring (YC W18), don’t interfere with the business model!

That's just a coincidence. The job posts go into a queue and it's semi-random when they get placed. The current submission appearing at the same time is unrelated. At least I presume it is.
Post reply on HN