Live data from Hacker News

OpenAI's new open-source model is basically Phi-5

seangoedecke.com

151–160 of 233 posts

Re: OpenAI's new open-source model is basically Phi-5

#151
post #20

I saw a bunch of people complaining on Twitter about how GPT-OSS can't be customized or has no soul and I noticed that none of them said what they were trying to accomplish. "The main use-case for fine-tuning small language models is for erotic role-play, and there’s a serious demand." Ah.

It's not just erotic role play that the censorship affects. My life involves a lot of sexual discussions and that means that everyday talk, chat summaries, email rewrites or translations will cause the model to shut down. I do the latter a lot especially to find colloquialisms because Google translate is often too literal. It's so annoying. Right now I'm using abliterated llama 3.1. I have no need for vision but I wa…

I found DeepSeek R1 (better for questions) and V3 (better for prose) to be very willing to discuss sex with a simple system prompt, as well as being very pleasant in articulation. I guess, I prefer them because they are almost SOTA and very large.

Not through the official interface though. Needs to be hosted by a third party. OpenRouter has a generous free tier for both.

I just saw that there is an abliterated version as well. Not sure how to try it though.

Re: OpenAI's new open-source model is basically Phi-5

#152
post #52

Earlier quoted context omitted.

> Also, these days people have Reddit accounts reserved for porn where they do exactly this. So it was built after all. Didn't reddit remove porn?

No. Not at all. You must be thinking of a different site. Tumblr did and onlyfans did for a hot minute and then backtracked. Neither of them intended to be porn sites. It's kind of a natural occurrence on UGC sites . Look at Civitai... Credit card processors are kinda weary of it for some legal reasons I'm not qualified to enough to really understand.

> for some legal reasons

For moralizing activist reasons. It's nothing to do with legality. With any luck eventually they'll inadvertently trample a sacred cow of whichever party is currently in power and we'll finally get sane legislation outlawing their overbearing nonsense.

Re: OpenAI's new open-source model is basically Phi-5

#153
post #139
post #55

Earlier quoted context omitted.

> By definition, a model can't "know" things that are not somewhere in its training set, unless it can use a tool to query external knowledge. Well, it could also make inferences. Like, it could find a new mathematical proof, even if that's never in the training set.

But how, it's not like it's thinking, it's just spitting the next likely token

This can generate new text. If the abilities generalise somewhat (and there is lots of evidence they DO generalise on some level), then there is no obstacle to generating new proofs, although the farther away they are from the training data, the less likely it becomes.

For an obvious example of generalisation: the models are able to write more code than there is in the dataset. If you ask it to write some specific, though easy, function, it is very unlikely it is present verbatim in the dataset, and yet the model can adapt.

Re: OpenAI's new open-source model is basically Phi-5

#154

From the article: > For the same reason that Microsoft probably continued to train Phi-style models: safety. Releasing an open-source model is terrifying for a large organization. Once it’s out there, your name is associated with it forever, and thousands of researchers will be frantically trying to fine-tune it to remove the safety guardrails. I don't think this is really an issue in practice. Llama 2 and 3 were unc…

If I think about Llma, I think about uncensored. not that I ever used one, but there were not many use cases for censored llama when others were so much better at other things

Re: OpenAI's new open-source model is basically Phi-5

#155

Earlier quoted context omitted.

Here's a more concrete example where GPT-OSS 20B performed very well IMHO. I tested it against Gemma 3 12B, Phi 4 Reasoning 14B, Qwen 2.5-coder 14B. The prompt is modeled as a part of an agent of sorts, and the "human" question is intentionally ill-posed to emulate people saying the wrong thing. The prompt begins with asking the model to convert a question into matlab code, add any assumptions as comments at the star…

Why Qwen2.5 and not Qwen3-30B-A3B-Thinking-2507 or Qwen3-Coder-30B-A3B-Instruct?

Mostly because I had it downloaded already and I'm mostly interested in models that fit on my 16GB GPU. But since you asked, I ran the same questions through both 30B models in the q4_k_m variant, as GPT-OSS 20B is also quantized to about q4.

First the ill-posed question:

Qwen 3 Coder gave very similar answer to Phi 4, though included a more long-winded explanation in the comments. So not bad, but not great either.

Qwen 3 Thinking thought for a good minute before deciding the question was ill-posed and return the hash marks. However the following explanation was not as good as GPT-OSS, IMHO: The question is unclear because an LC circuit (without resistance) does not have a "cutoff frequency"; cutoff frequency applies to filter circuits like RC or RLC. Additionally, the inductance (L) value is missing for calculating resonant frequency in an RLC circuit. The given R and C values are insufficient without L.

Sure, an unloaded LC filter doesn't have a cutoff frequency, but in all normal cases the load is implied[1] and so the LC filter does have a cutoff frequency. So more thinking to get to a worse answer.

The SQL question:

Qwen 3 Coder did identify the same pitfall as GPT-OSS, however didn't flag it as clearly as GPT-OSS, mostly because it also flagged some unnecessary stuff so got drowned. It did make the same assumption about evenly dividing, and overall the answer was about as good. However the speed on my computer was roughly half the number of tokens per second as GPT-OSS, at just ~9 tokens/second.

Qwen 3 Thinking thought for 3 minutes, yet managed to miss the key aspect, thus giving everyone the pizza. And it did so at the same slow pace as Qwen 3 Coder.

The SQL question requires a somewhat large context due to the large table definitions, and being a larger model it required pushing more layers to the CPU, which I assume is the major factor in the speed drop.

So overall Qwen 3 Coder was a solid contender, but on my PC much slower. If it could run entirely on GPU I'd certainly try it a lot more. Interestingly Qwen 3 Thinking was just plain worse. Perhaps not tuned to other tasks besides coding?

[1]: https://www.ti.com/lit/an/slaa701a/slaa701a.pdf section 3.3 page 9

[2]: https://github.com/ollama/ollama/issues/11772

Re: OpenAI's new open-source model is basically Phi-5

#156
post #20

I saw a bunch of people complaining on Twitter about how GPT-OSS can't be customized or has no soul and I noticed that none of them said what they were trying to accomplish. "The main use-case for fine-tuning small language models is for erotic role-play, and there’s a serious demand." Ah.

why everyone keep pretending very hard that this entire AI summer was not started exclusively by pioneers trying to perfect virtual girlfriend? it's a fact.

Re: OpenAI's new open-source model is basically Phi-5

#157
post #20

I saw a bunch of people complaining on Twitter about how GPT-OSS can't be customized or has no soul and I noticed that none of them said what they were trying to accomplish. "The main use-case for fine-tuning small language models is for erotic role-play, and there’s a serious demand." Ah.

We use OpenAI API at work, it fails when translating children stories, the reason is violence. Either the model safety is shit or the AI companies are pushed by some extremists groups to censor shit that is acceptable for children in Europe (Romania). But the most bullshit is when you give it a safe prompt, the mdoel generates a response and the safety checker kicks in and blocks the response because it thinks the model was too naughty.

Re: OpenAI's new open-source model is basically Phi-5

#158
post #11

Does anyone know how synthetic data is commonly generated? Do they just sample the model randomly starting from an empty state, perhaps with some filtering? Or do they somehow automatically generate prompts and if how? Do they have some feedback mechanism, e.g. do they maybe test the model while training and somehow generate data related to poorly performing tests?

I have done that at meta/FAIR and it is published in the Llama 3 paper. You usually start from a seed. It can be a randomly picked piece of website/code/image/table of contents/user generated data, and you prompt the model to generate data related to that seed. After, you also need to pass the generated data through a series of verifiers to ensure quality.

Re: OpenAI's new open-source model is basically Phi-5

#159
post #54
post #19

> for instance, they have broad general knowledge about science, but don’t know much about popular culture That seems like a good focus. Why learn details that can change within days of it being released? Instead, train the models to have good general knowledge, and be really good at using tools, and you won't have to re-train models from scratch just because some JS library now has a different API, instead the model…

Why would anything change? You feed the model approximately all the text you have ever. And some things like 'popular culture of 2025' won't change, just because the calendar changed to 2026. Just like the popular culture of the 1980s is what it was, and won't change.

> Why would anything change?

It's not that facts change across time, but the relevancy of the details change. For example, it would be great if we could teach LLMs all the APIs all React versions have ever had, but if we do that for everything, there will be no limit to the weight's weight, and we'd need new weights every quarter if not more often. That seems very unsustainable.

So the information that corresponds to "What is the current React API for X" changes whenever the API changes, but "What is the React v5 API for X" remains the same. Having the model being able to look up those things via external channels would let us use the same models for way longer, if you need "up to date data" about things.

Re: OpenAI's new open-source model is basically Phi-5

#160
post #20

I saw a bunch of people complaining on Twitter about how GPT-OSS can't be customized or has no soul and I noticed that none of them said what they were trying to accomplish. "The main use-case for fine-tuning small language models is for erotic role-play, and there’s a serious demand." Ah.

Maybe I'm just not seeing it, but is that use case really real and not just prudish hallucinations? The market for NSFW novels is smaller than even cyberpunk paperbacks, there's no way everyday people build up addiction for an interactive version of it.

Not quite. The market for romance (which can, in fact, get arbitrary degrees of spicy) is by far the largest literary market. (Source: https://bookadreport.com/book-market-overview-authors-statis...)

LLMs are also hilariously bad at what makes erotica hot in the first place.

Post reply on HN