Live data from Hacker News

OpenAI's new open-source model is basically Phi-5

seangoedecke.com

201–210 of 233 posts

Re: OpenAI's new open-source model is basically Phi-5

#201
I think its likely that soon the "AI" aspect of the "models" get worse and worse with each iteration. This is because:

1. data pollution/dilution -- with each subsequent generation more and more LLM-produced content will make up the bulk of the web this means it either gets into get into the data set, which makes the model "spikey" or reduces the relevant knowledge available which dumbs the model down 2. LLMs were sort of a freak breakthrough in deep learning and while its remarkable, tuning them only makes them marginally better. Diminishing returns apply here. Its a new LLM but its still an LLM -- years old technology now

However, despite these two realities, the total utility/productivity/societal gain from a product (note I didn't say model) like ChatGPT can still increase by orders of magnitude. That is because companies like OpenAI, and many, many, other technology corporations and startups are figuring out how to leverage LLMs to do stuff much more powerful than just answer questions from the corpus (ie: perform quality research, reevaluate itself, computer controls, etc, etc).

Consider for example, flat screen displays were pioneered in like the 50's and arguably, they didn't become disruptively useful until the advent of smart phones. So yeah, the model may get sort of worse or maybe more brittle, but it almost doesn't matter if they are figuring out what to do with the model and making the model more useful for actual tasks. Sure its cute to talk to an AI persona and ask it questions but that is probably the least important aspect of these type of models. Microsoft word had Clippy, and yeah that was cool. But productivity gain came from word processing, the '.doc' filetype, filesharing, editing, etc. Clippy is just a meme now and that's a likely future scenario for the "chat" features in LLM products IMHO.

Re: OpenAI's new open-source model is basically Phi-5

#202

I gave it a random sci-fi novel and made it translate a chapter, which is something I do with all models. It refused to discuss minors in sexualized contexts. I was like W.T.F.?! and started bisecting the book, trying to find the piece that triggers this. Turns out there was some absolutely innocent, two sentence long romantic remark involving two secondary 17 years old characters in an unrelated place. Another issue…

The 20B refused to acknowledge that he gave me wrong informations. Usually models just apologize after I insist 2 or 3 times. So it was the shortest LLM I tried, I honestly can't trust such models for anything.

Isn't an apology a bad metric for evaluating models?

Without understanding much, it seems to be more an indication of the type of content the model was trained on, rather than an indicator of how good or bad a model is, or how much it knows. It would probably be easy to create bad model that constantly outputs wrong information, but always apologizes when corrected.

Re: OpenAI's new open-source model is basically Phi-5

#203

Earlier quoted context omitted.

it's not erotic role-play, but I have a use case of making an AI-powered NetHack clone. specifically, to generate dungeon layouts, dialog for NPCs and to fill in the boatloads of minutae and interactions which NetHack is famous for. you kind of need soul for that, and a lot of background knowledge on mythology/fantasy lore, but also tool use to work the world systems.

I've been experimenting with using various LLMs as a game master for a Lovecraft-inspired role-playing game (not baked into an application, just text-based by prompting). While the LLMs can generate scenarios that fit the theme, they tend to be very generic. I've also noticed that the models are extremely susceptible to suggestion. For example, in one scenario, my investigator was in a bar, and when I commented to an…

> on, 'Hey, doesn't the barkeeper look a little strange?', the LLM immediately seized on that and turned the barkeeper into an evil, otherworldly creature.

Though making it an evil otherwordly creature is a bit extreme, it's at least similar to what a flexible GM can do. In my DMing days, I would often develop new paths that integrated into the whole inspired by things my players noticed/suspected.

Re: OpenAI's new open-source model is basically Phi-5

#204

Earlier quoted context omitted.

No. Not at all. You must be thinking of a different site. Tumblr did and onlyfans did for a hot minute and then backtracked. Neither of them intended to be porn sites. It's kind of a natural occurrence on UGC sites . Look at Civitai... Credit card processors are kinda weary of it for some legal reasons I'm not qualified to enough to really understand.

> for some legal reasons For moralizing activist reasons. It's nothing to do with legality. With any luck eventually they'll inadvertently trample a sacred cow of whichever party is currently in power and we'll finally get sane legislation outlawing their overbearing nonsense.

I doubt it's moralizing reasons, it's probably because once you disseminate porn, you become a vehicle for child porn, which is a legal and PR disaster.

Re: OpenAI's new open-source model is basically Phi-5

#205
post #11

Does anyone know how synthetic data is commonly generated? Do they just sample the model randomly starting from an empty state, perhaps with some filtering? Or do they somehow automatically generate prompts and if how? Do they have some feedback mechanism, e.g. do they maybe test the model while training and somehow generate data related to poorly performing tests?

Common synthetic data generation methods include distillation (teacher-student), self-improvement via bootstrapping (model improves its own outputs), instruction-following synthesis, and controlled sampling with filtering for quality/alignment.

Re: OpenAI's new open-source model is basically Phi-5

#206
post #80
post #58

Earlier quoted context omitted.

even if it is victim free, it can affect mental health in a way that a consumer will be more compelled to do a criminal act and create a real victim. let's say you publish a Steam game how to be a school shooter and shoot kids, wouldn't that lead to real school shootings ? who can definitely say that computer generated content about criminal behavior, won't lead to real crime with real victims? https://en.wikipedia.o…

Grand Theft Auto 5 sold over 200 million copies, and military/crime/shooter games have always been incredibly popular. Yet crime has been decreasing over the past few decades in the United States, where both cars and guns are easily accessible.

The most popular activity in GTA V isn't stealing cars. It's getting a 5-star wanted level as a mass shooter.

And to this day, military recruiters use the AC130 mission in CoD to convince people to become aerial gunners.

Re: OpenAI's new open-source model is basically Phi-5

#207

Earlier quoted context omitted.

The 20B refused to acknowledge that he gave me wrong informations. Usually models just apologize after I insist 2 or 3 times. So it was the shortest LLM I tried, I honestly can't trust such models for anything.

Isn't an apology a bad metric for evaluating models? Without understanding much, it seems to be more an indication of the type of content the model was trained on, rather than an indicator of how good or bad a model is, or how much it knows. It would probably be easy to create bad model that constantly outputs wrong information, but always apologizes when corrected.

Well if the model can't accept it got an information wrong, how can he help to tweak anything? or give something accurate?

Re: OpenAI's new open-source model is basically Phi-5

#208

Earlier quoted context omitted.

> for some legal reasons For moralizing activist reasons. It's nothing to do with legality. With any luck eventually they'll inadvertently trample a sacred cow of whichever party is currently in power and we'll finally get sane legislation outlawing their overbearing nonsense.

I doubt it's moralizing reasons, it's probably because once you disseminate porn, you become a vehicle for child porn, which is a legal and PR disaster.

If you disseminate user uploaded porn then moralizing activists can certainly accuse you of that. It's performative hand wringing though. The admins assuredly don't want to distribute it and the goal of anyone publicly uploading that on the clearnet is to harass and disrupt rather than to disseminate.

Anyway "child porn" as well as the broader "legal reasons" fails to explain the US payment processors' moves to block all sorts of content and products over the years. Even including porn that isn't user uploaded (and thus has proper records keeping).

Re: OpenAI's new open-source model is basically Phi-5

#209

Earlier quoted context omitted.

> researchers will be frantically trying to fine-tune it to remove the safety guardrails. It is a really weak excuse. They are more likely to take a reputation hit for having silly guardrails than for having someone to remove them. Imagine if Bill Gates decided not to release MS Paint in 1985 because someone could have drawn something offensive with it.

Or imagine if Bill Gates decided to not release Comic Sans in the 90s cause someone could have written something offensive with it.... oh wait that wouldn't have been too bad (/S)

I like it. Just block anything that uses it, and you have cleaned up a big part of your internet and lost nothing of value.

Re: OpenAI's new open-source model is basically Phi-5

#210
post #170

Earlier quoted context omitted.

i think god did a fairly hack job overall and i’ll gladly en masse commit acts that please me and fail to please his non-existent ass. i’d even turn gay but that’s a bit out of my comfort zone

nickpsecurity's post is currently flagged so I can't reply directly, but I think it's important to clarify that traditional Christian sexual morality aims to maximize population growth. The full list of sexual sins includes one that nickpsecurity missed: sodomy. Sodomy traditionally includes oral sex. The only logical reason to ban oral sex within marriage is to encourage reproduction. Although the Bible does not exp…

[flagged]
Post reply on HN