Show HN: AskHN
11–20 of 139 posts
Re: Show HN: AskHN
#12Re: Show HN: AskHN
#13Is there a way to opt out of one's comments being used for this?
Re: Show HN: AskHN
#14I'm a little surprised that Hacker News comments weren't already in the GPT-3 training set. I just assumed that OpenAI had vacuumed up most of the web already.
I am guessing they already were? But this is 100% pure, concentrated HN not contaminated with nonsense from the rest of the web :)
Re: Show HN: AskHN
#15Re: Show HN: AskHN
#16Is there a way to opt out of one's comments being used for this?
Re: Show HN: AskHN
#17Re: Show HN: AskHN
#18Re: Show HN: AskHN
#19> I trained on a corpus of over 6.5 million Hacker News comments How long did it take to scrape them and train the "corpus" on this content?
Re: Show HN: AskHN
#20Earlier quoted context omitted.
I am guessing they already were? But this is 100% pure, concentrated HN not contaminated with nonsense from the rest of the web :)
I have to assume that targeted/curated LLM training sets will have a tendency to be less accurate than very general, just by the very nature of how they work. (edited for clarity)
This surprised me, I thought it wouldn't do much better, but I wasn't expecting that specializing it on my target data would reduce performance! I had fewer examples than the minimum OpenAI recommends, so maybe it was a case of overfitting or something like that.