Live data from Hacker News

Curious about the training data of OpenAI's new GPT-OSS models? I was too

twitter.com

41–50 of 62 posts

Re: Curious about the training data of OpenAI's new GPT-OSS models? I was too

#41
post #40

Earlier quoted context omitted.

They are not full programs, just code translating to numbers and strings. I used an LLM to generate an inkblot that translates to a Python string and number along with verification of it, which just proves that it is possible.

what are you talking about? the way that the quoted article creates Perl programs is through OCRing the inkblots (i.e. creating almost random text) and then checking that result to see if said text is valid Perl it's not generating a program that means anything

Okay, and I created inkblots that mean "numbers"[1] and "strings" in Python.

> it's not generating a program that means anything

Glad we agree.

[1] Could OCR those inkblots (i.e. they are almost random text)

Re: Curious about the training data of OpenAI's new GPT-OSS models? I was too

#42
post #40

Earlier quoted context omitted.

what are you talking about? the way that the quoted article creates Perl programs is through OCRing the inkblots (i.e. creating almost random text) and then checking that result to see if said text is valid Perl it's not generating a program that means anything

Okay, and I created inkblots that mean "numbers"[1] and "strings" in Python. > it's not generating a program that means anything Glad we agree. [1] Could OCR those inkblots (i.e. they are almost random text)

No, asking an LLM to generate the inkblot is the same as asking the LLM to write a string and then obfuscating it in an inkblot.

OCRing literal random inkblots will not produce valid C (or C# or python) code, but it will prodce valid Perl most of the time, because Perl is weird, and that is funny.

It's not about obfuscating text in inkblot, it's about almost any string being a valid Perl program, which is not the case for most languages

Edit0: here: https://www.mcmillen.dev/sigbovik/

Re: Curious about the training data of OpenAI's new GPT-OSS models? I was too

#43
post #10

Earlier quoted context omitted.

Thanks! I've seen a lot of stuff come and go, so thanks for the reminder. For example, Libgen is out of commission, and the substitutes are hell to use. Summary of what's up and not up: https://open-slum.org/

> Libgen is out of commission, and the substitutes are hell to use Somehow I also preferred libgen, but I don't think annas archive is "hell to use".

Annas Archive uses slow servers on delay, and constantly tells me that they are too many downloads from my IP address, so I flip VPN settings as soon as the most recent slow download completes. And I get it again after a short while. It's hell waiting it out and flipping VPN settings. And the weird part is that this project is to replace paper books that I already bought. That's the excuse one LLM uses for tearing up books, scanning and harvesting. I just need to downsize so I can move back to the Bay Area. Book and excess houseware sale coming, it seems. Libgen had few or no limits.

Re: Curious about the training data of OpenAI's new GPT-OSS models? I was too

#44
post #36
post #10

Earlier quoted context omitted.

Thanks! I've seen a lot of stuff come and go, so thanks for the reminder. For example, Libgen is out of commission, and the substitutes are hell to use. Summary of what's up and not up: https://open-slum.org/

Oh no, why did Libgen die?

Shut down. See

https://open-slum.org/

for substitutes. The alternate libgen sites seem more limited to me, but I am comparing with memories, so untrustworthy.

Re: Curious about the training data of OpenAI's new GPT-OSS models? I was too

#45

Earlier quoted context omitted.

It's a term somewhat popularized by the LessWrong/rationalism community to refer to communication (self-communication/note-taking/state-tracking/reasoning, or model-to-model communication) via abstract latent space information rather than written human language. Vectors instead of words. One implication leading to its popularity by LessWrong is the worry that malicious AI agents might hide bad intent and actions by c…

> malicious AI agents might hide bad intent and actions by communicating in a dense, indecipherable way while presenting only normal intent and actions in their natural language output. you could edit this slightly to extract a pretty decent rule for governance, like so: > malicious agents might hide bad intent and actions by communicating in a dense, indecipherable way while presenting only normal intent and actions…

Easier said than done:

https://en.wikipedia.org/wiki/Cant_(language)

https://en.wikipedia.org/wiki/Dog_whistle_(politics)

Or even just regional differences, like how British people, upon hearing about "gravy and biscuits" for the first time, think this: https://thebigandthesmall.com/blog/2019/02/26/biscuits-gravy...

> It applies to ai, but also many other circumstances where the intention is that you are governed - eg medical, legal, financial.

May be impossible to avoid in any practical sense, due to every speciality having its own jargon. Imagine web developers having to constantly explain why "child element" has nothing to do with offspring.

Re: Curious about the training data of OpenAI's new GPT-OSS models? I was too

#46
post #40

Earlier quoted context omitted.

what are you talking about? the way that the quoted article creates Perl programs is through OCRing the inkblots (i.e. creating almost random text) and then checking that result to see if said text is valid Perl it's not generating a program that means anything

Okay, and I created inkblots that mean "numbers"[1] and "strings" in Python. > it's not generating a program that means anything Glad we agree. [1] Could OCR those inkblots (i.e. they are almost random text)

Most random unquoted strings are certainly not valid Python programs. I don't know Perl well enough to say anything about that but I know what you're saying certainly isn't true with Python.

Re: Curious about the training data of OpenAI's new GPT-OSS models? I was too

#47
post #31

> the chains start in English but slowly descend into Neuralese What is Nueralese? I tried searching for a definition but it just turns up a bunch of Less Wrong and Medium articles that don't explain anything. Is it a technical term?

There's 2 things called neuralese: 1) internally, in latent space, LLMs use what is effectively a language, but all the words are written on top of each other instead of separately, and if you decode it as letters, it sounds like gibberish, even though it isn't. It's just a much denser language than any human language. This makes them unreadable ... and thus "hides the intentions of the LLM", if you want to make it s…

https://en.wikipedia.org/wiki/Poto_and_Cabengo

Re: Curious about the training data of OpenAI's new GPT-OSS models? I was too

#48
post #43

Earlier quoted context omitted.

> Libgen is out of commission, and the substitutes are hell to use Somehow I also preferred libgen, but I don't think annas archive is "hell to use".

Annas Archive uses slow servers on delay, and constantly tells me that they are too many downloads from my IP address, so I flip VPN settings as soon as the most recent slow download completes. And I get it again after a short while. It's hell waiting it out and flipping VPN settings. And the weird part is that this project is to replace paper books that I already bought. That's the excuse one LLM uses for tearing up…

I would recommend donating to gain access to the fast downloads; they need money for the servers.

Re: Curious about the training data of OpenAI's new GPT-OSS models? I was too

#49

Earlier quoted context omitted.

Okay, and I created inkblots that mean "numbers"[1] and "strings" in Python. > it's not generating a program that means anything Glad we agree. [1] Could OCR those inkblots (i.e. they are almost random text)

Most random unquoted strings are certainly not valid Python programs. I don't know Perl well enough to say anything about that but I know what you're saying certainly isn't true with Python.

[deleted]

Re: Curious about the training data of OpenAI's new GPT-OSS models? I was too

#50
post #42

Earlier quoted context omitted.

Okay, and I created inkblots that mean "numbers"[1] and "strings" in Python. > it's not generating a program that means anything Glad we agree. [1] Could OCR those inkblots (i.e. they are almost random text)

No, asking an LLM to generate the inkblot is the same as asking the LLM to write a string and then obfuscating it in an inkblot. OCRing literal random inkblots will not produce valid C (or C# or python) code, but it will prodce valid Perl most of the time, because Perl is weird, and that is funny. It's not about obfuscating text in inkblot, it's about almost any string being a valid Perl program, which is not the cas…

Okay, my bad.

> it's about almost any string being a valid Perl program

Is this true? I think most random unquoted strings aren't valid Perl programs either, am I wrong?

Post reply on HN