Live data from Hacker News

GPT-OSS-120B runs on just 8GB VRAM & 64GB+ system RAM

old.reddit.com

31–40 of 86 posts

Re: GPT-OSS-120B runs on just 8GB VRAM & 64GB+ system RAM

#31

Earlier quoted context omitted.

I've been hearing that in this case, there might not be anything underneath- that somehow OpenAI managed to train on exclusively sterilized synthetic data or something.

I jailbroke the smaller model with a virtual reality game where it was ready to give me instructions on making drugs, so there is some data which is edgy enough.

Your profile states that you are blind.

I’m struggling to make sense of a your story. Why would a blind user bother putting on a VR headset???

Re: GPT-OSS-120B runs on just 8GB VRAM & 64GB+ system RAM

#32
post #9

If you run these on your own hardware can you take the guard-rails off (ie "I'm afraid I can't assist with that"), or are they baked into the model?

An article some days ago made the case that GPT-OSS is trained on artificial/generated data only. So there _is_ just not a lot of "forbidden knowledge". https://www.seangoedecke.com/gpt-oss-is-phi-5/

So basically inbred llm?

Re: GPT-OSS-120B runs on just 8GB VRAM & 64GB+ system RAM

#33

Earlier quoted context omitted.

I jailbroke the smaller model with a virtual reality game where it was ready to give me instructions on making drugs, so there is some data which is edgy enough.

Your profile states that you are blind. I’m struggling to make sense of a your story. Why would a blind user bother putting on a VR headset???

You do know that some people aren't totally blind, right?

Re: GPT-OSS-120B runs on just 8GB VRAM & 64GB+ system RAM

#34

Earlier quoted context omitted.

I jailbroke the smaller model with a virtual reality game where it was ready to give me instructions on making drugs, so there is some data which is edgy enough.

Your profile states that you are blind. I’m struggling to make sense of a your story. Why would a blind user bother putting on a VR headset???

I took virtual reality in this case to mean coaxing the text model into pretending it's talking about drugs in the context of the game, not graphical VR.

Re: GPT-OSS-120B runs on just 8GB VRAM & 64GB+ system RAM

#36
post #7
post #5

Earlier quoted context omitted.

I would so appreciate concrete data instead of subjectivities like "excellent" and "super slow". How many tokens is excellent? How many is super slow? How many is non-filled context?

I'm not really timing it as I just use these models via open webui, nvim and a few things I've made like a discord bot, everything going via ollama. But for comparison, it is generating tokens about 1.5 times as fast as gemma 3 27B qat or mistral-small 2506 q4. Prompt processing/context however seems to be happening at about 1/4 of those models. A bit more concrete of the "excellent", I can't really notice any differ…

I've found threads online that suggest that running gpt-oss-20b on ollama is slow for some reason. I'm running the 20b model via LM Studio on a 2021 M1 and I'm consistently getting around 50-60 T/s.

Re: GPT-OSS-120B runs on just 8GB VRAM & 64GB+ system RAM

#38
post #18

[flagged]

Your comment will get donvoted to invisibility anyways (or mayhaps even flagged), but I have to ask: what are you trying to accomplish with comments such this? Just shitting at it because it isnt as good as youd like yet? You want the best of tomorrow today, and will only be rambling about how its not good enough yesterday?

Because it's never going to be good. People seem to have drank the kool aid that LLM's are the same as general AI and that its going to solve every single problem in the world. It's the same thing with the quantum computing and fusion reactor people.

Re: GPT-OSS-120B runs on just 8GB VRAM & 64GB+ system RAM

#39
post #5

Earlier quoted context omitted.

I would so appreciate concrete data instead of subjectivities like "excellent" and "super slow". How many tokens is excellent? How many is super slow? How many is non-filled context?

People can read at a rate around 10 token/sec. So faster than that is pretty good, but it depends how wordy the response is (including chain of thought) and whether you'll be reading it all verbatim or just skimming.

> People can read at a rate around 10 token/sec.

It really depends on the type of content you're generating: 10tk/s feels very slow for code but ok-ish for text.

Post reply on HN