Live data from Hacker News

OpenAI Releases Largest GPT-2 Text Generation Model

openai.com

121–130 of 166 posts

Re: OpenAI Releases Largest GPT-2 Text Generation Model

#121
post #106

Earlier quoted context omitted.

Just an afterthought, but you could probably cobble something together that asks someone for some basic prompts, like names, locations, etc, then uses a series of templates to generate a chapter of text that's reasonably coherent and tailored to an individual.

So. 1. Feed in all the Game of Throne books 2. Generate final book 3. Profit?

Feeding in the names of GOT characters seems to produce some entertaining alternate-universe stuff, looks like they're trying to generate something similar to a Wiki page:

Ned Stark (The Lord of Winterfell) is a Stark from the North of the Seven Kingdoms.

Contents show]

Biography

Background

Ned is the father of Arya, Brandon, Robb, Rickon, and Hodor. He had married Catelyn Tully after her father died in the Battle of the Whispering Wood. During the War of the Five Kings he served as Hand of the King under Lord Eddard Stark after his father's death. Ned has two brothers, Willem and Torrhen Stark. Ned, his siblings and their children were all killed during the Red Wedding

Re: OpenAI Releases Largest GPT-2 Text Generation Model

#122
post #115

Just like various chat bot incarnations, it's reasonably OK at generating generic sentences without real understanding. I tried several attempts with this leading fragment: "Is borrow checking unique to Rust? Probably not. But". The results are not pretty (the best one is as follows). Not impressed at all with this giant model :| "Is borrow checking unique to Rust? Probably not. But the Rust library is written by and…

Wow! That is meaningless but difficult to distinguish at a glance, especially when read by someone not familiar with the subject.

Are we going to see such auto generated content take over as the primary pillar of the SEO content farms?

... Kinda like an automated (text) version of Siraj Raval videos :-P

Re: OpenAI Releases Largest GPT-2 Text Generation Model

#123

Earlier quoted context omitted.

So. 1. Feed in all the Game of Throne books 2. Generate final book 3. Profit?

You're joking, but I hope we can see this in RPGs of the future.

Just playing with a few prompts in the format:

It was [setting] in the [location]. The [creature] [action], and [intent] said "

And here's some of the results of the first complete sentence for a few iterations of each prompt (the intention being the prompts can be generated at runtime)

---

It was night in the dark dungeon. The Orc came out of nowhere, and threateningly said "Who are you? Where are you coming from?"

It was night in the dark dungeon. The Orc came out of nowhere, and threateningly said "We are not going to allow the Dark Lord to rule this world. We are going to rule this world from the top. You should go back to the shadows, or you can just go to my dungeon."

It was night in the dark dungeon. The Orc came out of nowhere, and threateningly said "What are you doing?"

---

It was twilight in the enchanted forest. The white elf suddenly appeared, and invitingly said "Good evening, we are ready for our banquet."

It was twilight in the enchanted forest. The white elf suddenly appeared, and invitingly said "Hey! I've been waiting for you! This is the first time I've seen you."

It was twilight in the enchanted forest. The white elf suddenly appeared, and invitingly said "Welcome, My Lady! I am here to serve you!"

It was twilight in the enchanted forest. The white elf suddenly appeared, and invitingly said "I am Tui-Yuan. Come down and meet my parents."

---

It was damp in the filthy sewer. The mutated rat crept up, and cunningly said "I will tell you everything."

It was damp in the filthy sewer. The mutated rat crept up, and cunningly said "I have an idea" in a voice so high that all the other rats in the sewer turned pale

It was damp in the filthy sewer. The mutated rat crept up, and cunningly said "I am the rat."

It was damp in the filthy sewer. The mutated rat crept up, and cunningly said "You'll die soon".

---

It was humid in the abandoned brothel. The policeman barged in, and brusquely said "I'm a policeman".

It was humid in the abandoned brothel. The policeman barged in, and brusquely said "go to hell".

---

It was frigid in the abandoned space station. The xenomorph burst in, and acerbically said "It's warm on the other side."

It was frigid in the abandoned space station. The xenomorph burst in, and acerbically said "Hello" while it slowly closed in.

It was frigid in the abandoned space station. The xenomorph burst in, and acerbically said "I hate this cold."

---

I think there's a lot of merit to this idea hey. Some of the responses are left field but could be woven into the charm. I guess the algorithm is pretty processor intensive though - is it worth it for "flavour"? It could work for a low fidelity or text based game I think.

Edit: I think it would work better if the prompt is not displayed, you just see the bit following the quote.

Re: OpenAI Releases Largest GPT-2 Text Generation Model

#125
post #96

Earlier quoted context omitted.

> Can anybody who is not into climbing even tell this is all fake? Hm, I'm pretty sure it's hard to climb five 8000 meter peaks in 24 hours :)

Yeah-yeah, and "Lomonosov Ridge" is a weird name for a peak in China (and you might even happen to know the Lomonosov Ridge is underwater). Please don't be distracted by details. Maybe the author meant 24 h per peak, maybe his typing is really spurious. Is it really obvious to somebody who doesn't know anything about climbing that this is not true? Come on. What's really fascinating is how truth is mixed up with made…

>8,000 meters is 24,064 ft

It's 26,247 ft.

Re: OpenAI Releases Largest GPT-2 Text Generation Model

#126
post #33

You can try it at: http://textsynth.org

Here's another one, seed text in italics. Clearly some Star Trek used in the training data. "Engage", said Captain Picard from the bridge of the Enterprise. "But sir", Commander Data began , "this must have been part of something called the Borg." "You know I didn't tell them that," said Picard, "I have a very hard time talking to them." - Picard, Data, and Worf while the ship is under attack by Borg drones at Federa…

> 'Borgs eat Borges. Borges eat Borges.'

That is amusingly appropriate.

https://en.wikipedia.org/wiki/Jorge_Luis_Borges

https://en.wikipedia.org/wiki/The_Library_of_Babel

Re: OpenAI Releases Largest GPT-2 Text Generation Model

#127
post #115

Just like various chat bot incarnations, it's reasonably OK at generating generic sentences without real understanding. I tried several attempts with this leading fragment: "Is borrow checking unique to Rust? Probably not. But". The results are not pretty (the best one is as follows). Not impressed at all with this giant model :| "Is borrow checking unique to Rust? Probably not. But the Rust library is written by and…

Love the fake link to github... Which model was this? Was it trained on software type discussion?

Re: OpenAI Releases Largest GPT-2 Text Generation Model

#128
post #98
post #94

Earlier quoted context omitted.

Well, it gives two different heights and locations for a single mountain

Yeah, I consider it the most serious fuck up. Yet, I'm not sure I would be really troubled even by that one if I knew nothing about that and casually came across this piece. I mean, maybe K2 is a name for 2 different peaks (it sure sound generic!), maybe something else I don't understand. Who knows!

I think the bigger problem is that we as content consumers are more and more looking for quantity rather than quality. While reading this climbing text I was thinking that I should probably carefully check for inconsistencies but I felt I can't be bothered to. I am not sure if that is only my feeling but it feels like the content matters less and less as long as I get entertained and with the neural nets getting better and better and I getting less picky we will meet somewhere in the middle where my brain keeps scanning for buzzwords and the net optimizing for it. Which feels this way with youtube, netflix et al. as well and I think this won't become a future we really like.

Re: OpenAI Releases Largest GPT-2 Text Generation Model

#130
post #20

Earlier quoted context omitted.

From my observation, even the largest GPT-2 model has difficulty retaining any long-range relationship information. In the "unicorn" writing example that was published originally, the model 'forgets' where the researchers are (climbing a mountain versus being beside a lake iirc) after just a few sentences. Because of this, it's hard to imagine models of this type being able to write long-form coherent papers. Now if…

Maybe the problem is that most of these models seem to rely on sequential information (even the transformer needs this for forward generation of text) to encode long range information. But I can’t remember the last time I relied on sequentially remembering the ordering of tokens in order to complete an essay or hell even reply to an email. Structurally we retain some kind of hierarchical information (topic, places, n…

I'm interested on this as well.

I have been trying to fine-tune GPT-2 on genre fiction to work as a sort of "fiction replicator". Stylistically it actually seems to do quite reasonably, but it lacks narrative cohesion. This problem, as you point out, is corpus agnostic.

I thought of trying to keep track of characters and key interactions outside of the model, but I haven't figured out how to make these two models interact reliably -- outside of just having the first component generate prompts for the second model in a kind of cooperative setting.

Is there a known way to set up transformer to do infix generation? That is: give it a start and end prompt, and an estimated number of tokens to fill in between. That seems like it should be doable and could improve things, but I haven't found any work on this problem yet and haven't had the time (and potentially don't have the skills) to look deeply myself yet.

Post reply on HN