Live data from Hacker News

OpenAI's new open-source model is basically Phi-5

seangoedecke.com

81–90 of 233 posts

Re: OpenAI's new open-source model is basically Phi-5

#81

Earlier quoted context omitted.

It’s not unreasonable to suspect that engaging in high-fidelity simulations of these behaviors will further entrench and worsen paraphilias. This is pretty evident with the progression of many pornography addictions that don’t include these sorts of things that still follow the pattern of increasing novelty seeking leading to increasingly deviant stuff. I am at a principled level uneasy with what’s fundamentally a so…

Are there any actual studies on this? Does access to simulations of illegal or objectionable material make pedophiles, rape fetishists, etc. more or less likely to try to access the real thing (or even worse, to try to commit crimes in the real world)? Because both possibilities are plausible, it’s hard to know which is correct.

Even if there are I and likely you lack the qualifications of making any clinical takeaways from them.

I'd really defer to experts.

I try to make tools in good faith and hope they're used responsibly to make the world a better place.

I'm not a clinical psychologist nor can I pretend to understand medical literature like someone with a PhD

Re: OpenAI's new open-source model is basically Phi-5

#83

Earlier quoted context omitted.

There's nothing wrong with it, but you have to understand the differences between different user groups to know which limitations are relevant to your own use cases. "It doesn't follow instructions" could mean "it won't pretend to be a horny elf" or "it hallucinates fields outside the JSON schema I specified"; the latter is much more of a problem for my uses.

{ "race": "elf", "horny": false ^^^^^^^^^^^^^^ Unsupported value.

Really, if you want a fey creature with horns, a satyr is probably a better bet than an elf.

Re: OpenAI's new open-source model is basically Phi-5

#84
post #60
post #54

Earlier quoted context omitted.

Why would anything change? You feed the model approximately all the text you have ever. And some things like 'popular culture of 2025' won't change, just because the calendar changed to 2026. Just like the popular culture of the 1980s is what it was, and won't change.

We don't feed the model all the text ever. They are still trained on less than 1% of the entire Internet corpus.

You are right, though on the other hand feeding it a selection of 1% of the entire corpus is already pretty close to 'all the text' (if you assume exponential growth in training over time).

Even multiplying that to approximately 100% of that corpus plus adding lots of non-internet text, will pale in comparison to all the non-text training data we will (or are) feeding our coming (and existing) multi-modal models.

If I may go out an a limb here: either we will see continuous great progress on text-based LLMs alone, or multi-modal models will become the next big focus. (Or both.)

That's because people are hungry for progress, and going multi-modal is the obvious thing to try to focus on, if text alone proves infeasible to drive progress.

Just to be clear: I make no prediction here on whether multi-modal will lead to progress, just that people will obviously try it and try it hard, if the focus on text starts to stall.

Re: OpenAI's new open-source model is basically Phi-5

#86
post #20

I saw a bunch of people complaining on Twitter about how GPT-OSS can't be customized or has no soul and I noticed that none of them said what they were trying to accomplish. "The main use-case for fine-tuning small language models is for erotic role-play, and there’s a serious demand." Ah.

My use case has been trying to remove the damn "apologies for this" and extraneous language that just waste tokens for no reason. GPT has always always always been so quick to waffle. And removing the chat interface as much as possible. Many benchmarks are better with text completion models, but they keep insisting on this horrible interface for their models. Fine tuning is there to ensure you get the output format y…

[deleted]

Re: OpenAI's new open-source model is basically Phi-5

#87
post #20

I saw a bunch of people complaining on Twitter about how GPT-OSS can't be customized or has no soul and I noticed that none of them said what they were trying to accomplish. "The main use-case for fine-tuning small language models is for erotic role-play, and there’s a serious demand." Ah.

it's not erotic role-play, but I have a use case of making an AI-powered NetHack clone. specifically, to generate dungeon layouts, dialog for NPCs and to fill in the boatloads of minutae and interactions which NetHack is famous for.

you kind of need soul for that, and a lot of background knowledge on mythology/fantasy lore, but also tool use to work the world systems.

Re: OpenAI's new open-source model is basically Phi-5

#88
post #53

Earlier quoted context omitted.

My use case has been trying to remove the damn "apologies for this" and extraneous language that just waste tokens for no reason. GPT has always always always been so quick to waffle. And removing the chat interface as much as possible. Many benchmarks are better with text completion models, but they keep insisting on this horrible interface for their models. Fine tuning is there to ensure you get the output format y…

> I swear they have tuned their models to waste tokens. Which seems a bit weird, because the customers of the chat interface (ie non-API customers) don't pay per token.

I've heard the theory a few times lately that AI businesses will increasingly move towards usage models over subscription models, so while it is probably accidental, it could also be a longer term strategy to normalize excessive token usage.

Re: OpenAI's new open-source model is basically Phi-5

#89
post #26

Earlier quoted context omitted.

what's the problem with that? we have erotic texts dating back thousands of years, basically as old as the act of writing itself https://en.wikipedia.org/wiki/Istanbul_2461

The pro-porn side has zero PR because respectable public figures don't see pro-porn advocacy as a good career move. At most, you'll get some oblique references to it. Meanwhile, the anti-porn side has a formidable alliance: Right-wing, religiously-motivated anti-porn activists. Left-wing, feminism-motivated anti-porn activists. Big corporate types with lots of $$$$ to spend who want their customer support chatbot to…

on the other hand, Musk et al are building AI-powered thirst traps, like Grok's "Ani", or the accursed Replika bots (whose user base went on suicide watch when the company abruptly decided to digitally neuter their "companions.")

erotic roleplay, imo, is much less harmful than using LLMs as surrogate partners. porn and sex workers have existed for millenia. they're an outlet for sexual tension. they don't alleviate feeling lonely or provide an alternative to human companionship.

I'm worried we'll produce a generation of hikkikomoris, who eschew human connection for sycophantic machines that always listen and never breaks their heart.

Re: OpenAI's new open-source model is basically Phi-5

#90
post #26

Earlier quoted context omitted.

what's the problem with that? we have erotic texts dating back thousands of years, basically as old as the act of writing itself https://en.wikipedia.org/wiki/Istanbul_2461

The pro-porn side has zero PR because respectable public figures don't see pro-porn advocacy as a good career move. At most, you'll get some oblique references to it. Meanwhile, the anti-porn side has a formidable alliance: Right-wing, religiously-motivated anti-porn activists. Left-wing, feminism-motivated anti-porn activists. Big corporate types with lots of $$$$ to spend who want their customer support chatbot to…

Maybe you have a porn test suite for LLM’s? See which ones are fine with or capable of talking about specific topics? I believe there was something similar for willingness to discuss sciency stuff.
Post reply on HN