Live data from Hacker News

OpenAI's new open-source model is basically Phi-5

seangoedecke.com

51–60 of 233 posts

Re: OpenAI's new open-source model is basically Phi-5

#51
post #37

Earlier quoted context omitted.

Porn is always the frontier. It's a well-understood self-contained use-case without many externalities and simple business models. What more, with porn, the medium is the product probably more than the content. Having it on home-media in the 80s was the selling point. Getting it over the 1-900 phone lines or accessing it over the internet ... these were arguably the actual product. It might have been a driver of earl…

1, Porn. 2 Military.

The firmer is a lot more nimble and the procurement processes of your customers are easier to navigate.

Re: OpenAI's new open-source model is basically Phi-5

#52

Earlier quoted context omitted.

There's something Freudian about the idea that the more you can customize porn, the more popular it is. That, despite the impression that "all men want one thing", it turns out that men all want very different and very oddly specific things. Imbuing somrthing with a "magical" quality that doesnt exist is the origin of the term "fetish". Its not about the raw attractive preference for a particular hair color; its a be…

oh it's wildly different. About 15 years ago I worked on a porn recommendation system. The idea is that you'd follow a number of sites based on likes and recommendations and you'd get an aggregated feed with interstitial ads. So I started with scraping and cross-reference, foaf, doing analysis. People's preferences are ... really complex. Without getting too lewd, let's say there's about 30-80 categories with non-mar…

> Also, these days people have Reddit accounts reserved for porn where they do exactly this. So it was built after all.

Didn't reddit remove porn?

Re: OpenAI's new open-source model is basically Phi-5

#53
post #20

I saw a bunch of people complaining on Twitter about how GPT-OSS can't be customized or has no soul and I noticed that none of them said what they were trying to accomplish. "The main use-case for fine-tuning small language models is for erotic role-play, and there’s a serious demand." Ah.

My use case has been trying to remove the damn "apologies for this" and extraneous language that just waste tokens for no reason. GPT has always always always been so quick to waffle. And removing the chat interface as much as possible. Many benchmarks are better with text completion models, but they keep insisting on this horrible interface for their models. Fine tuning is there to ensure you get the output format y…

> I swear they have tuned their models to waste tokens.

Which seems a bit weird, because the customers of the chat interface (ie non-API customers) don't pay per token.

Re: OpenAI's new open-source model is basically Phi-5

#54
post #19

> for instance, they have broad general knowledge about science, but don’t know much about popular culture That seems like a good focus. Why learn details that can change within days of it being released? Instead, train the models to have good general knowledge, and be really good at using tools, and you won't have to re-train models from scratch just because some JS library now has a different API, instead the model…

Why would anything change?

You feed the model approximately all the text you have ever. And some things like 'popular culture of 2025' won't change, just because the calendar changed to 2026. Just like the popular culture of the 1980s is what it was, and won't change.

Re: OpenAI's new open-source model is basically Phi-5

#55
post #7

If a model is trained only on synthetic data, is it still possible it will output things like this? https://x.com/elder_plinius/status/1952958577867669892

By definition, a model can't "know" things that are not somewhere in its training set, unless it can use a tool to query external knowledge. The problem is that the size of the training set required for a good model is so large, that's really hard to make a good model without including almost all known written text available.

> By definition, a model can't "know" things that are not somewhere in its training set, unless it can use a tool to query external knowledge.

Well, it could also make inferences. Like, it could find a new mathematical proof, even if that's never in the training set.

Re: OpenAI's new open-source model is basically Phi-5

#56
post #22
post #19

> for instance, they have broad general knowledge about science, but don’t know much about popular culture That seems like a good focus. Why learn details that can change within days of it being released? Instead, train the models to have good general knowledge, and be really good at using tools, and you won't have to re-train models from scratch just because some JS library now has a different API, instead the model…

Yeah, it always seemed like a sad commentary on our world that AIs are devoting their weights to encyclopedic knowledge of Harry Potter, Pokemon, and Reddit trolling.

Why? You gotta provide what your customers want.

And it's far from sad that we have so many resources, we can give everyone a supercomputer in their pocket just to take selfies and talk about Pokemon. Why would our AIs be any different?

Re: OpenAI's new open-source model is basically Phi-5

#57
post #52

Earlier quoted context omitted.

oh it's wildly different. About 15 years ago I worked on a porn recommendation system. The idea is that you'd follow a number of sites based on likes and recommendations and you'd get an aggregated feed with interstitial ads. So I started with scraping and cross-reference, foaf, doing analysis. People's preferences are ... really complex. Without getting too lewd, let's say there's about 30-80 categories with non-mar…

> Also, these days people have Reddit accounts reserved for porn where they do exactly this. So it was built after all. Didn't reddit remove porn?

No. Not at all. You must be thinking of a different site. Tumblr did and onlyfans did for a hot minute and then backtracked.

Neither of them intended to be porn sites. It's kind of a natural occurrence on UGC sites . Look at Civitai...

Credit card processors are kinda weary of it for some legal reasons I'm not qualified to enough to really understand.

Re: OpenAI's new open-source model is basically Phi-5

#58
post #20

I saw a bunch of people complaining on Twitter about how GPT-OSS can't be customized or has no soul and I noticed that none of them said what they were trying to accomplish. "The main use-case for fine-tuning small language models is for erotic role-play, and there’s a serious demand." Ah.

Porn is always the frontier. It's a well-understood self-contained use-case without many externalities and simple business models. What more, with porn, the medium is the product probably more than the content. Having it on home-media in the 80s was the selling point. Getting it over the 1-900 phone lines or accessing it over the internet ... these were arguably the actual product. It might have been a driver of earl…

even if it is victim free, it can affect mental health in a way that a consumer will be more compelled to do a criminal act and create a real victim.

let's say you publish a Steam game how to be a school shooter and shoot kids, wouldn't that lead to real school shootings ?

who can definitely say that computer generated content about criminal behavior, won't lead to real crime with real victims?

https://en.wikipedia.org/wiki/Active_Shooter

Re: OpenAI's new open-source model is basically Phi-5

#59
post #20

I saw a bunch of people complaining on Twitter about how GPT-OSS can't be customized or has no soul and I noticed that none of them said what they were trying to accomplish. "The main use-case for fine-tuning small language models is for erotic role-play, and there’s a serious demand." Ah.

My use case has been trying to remove the damn "apologies for this" and extraneous language that just waste tokens for no reason. GPT has always always always been so quick to waffle. And removing the chat interface as much as possible. Many benchmarks are better with text completion models, but they keep insisting on this horrible interface for their models. Fine tuning is there to ensure you get the output format y…

The jargon to google here is "length bias"

It turns out if you generate two LLM responses and ask a judge to choose which is better, many judges have a bias in favour of long answers full of waffle.

Re: OpenAI's new open-source model is basically Phi-5

#60
post #54
post #19

> for instance, they have broad general knowledge about science, but don’t know much about popular culture That seems like a good focus. Why learn details that can change within days of it being released? Instead, train the models to have good general knowledge, and be really good at using tools, and you won't have to re-train models from scratch just because some JS library now has a different API, instead the model…

Why would anything change? You feed the model approximately all the text you have ever. And some things like 'popular culture of 2025' won't change, just because the calendar changed to 2026. Just like the popular culture of the 1980s is what it was, and won't change.

We don't feed the model all the text ever. They are still trained on less than 1% of the entire Internet corpus.
Post reply on HN