Live data from Hacker News

The computers are getting better at writing

newyorker.com

131–140 of 165 posts

Re: The computers are getting better at writing

#131

Earlier quoted context omitted.

True, but OP was asking for an AI tool to help in this. I think the language bias on such a tool would be enormous, and it would contribute to flatten a piece of writing. Being effective by saying things the right way, or picking the right words to express a specific context, is definitely important. My point was this: once you get to a good-enough point, what you're looking for is an editor (of sorts, not necessaril…

>> One clear example: attributive nouns. Specific examples, please? I'm curious :)

In a case where a noun is a modifier of the following word, Italians tend to prefer specificity , therefore falling for the “noun + of…” construct. A chicken soup bowl is easy to understand, but for an Italian that’s a “piatto di brodo di pollo”, so an Italian native would tend to feel like “a bowl of chicken soup” is the better choice. In this case there’s no real mistake (and for “bowl of” you might argue there’s almost no meaning difference) but most of the time you end up with convoluted sentences, especially if a genitive is lurking around —- “Andy’s chicken soup bowl” vs “the bowl of chicken soup of Andy”. Understandable, more familiar to an Italian speaker, yet kind of wrong.

Re: The computers are getting better at writing

#133
Does anybody know what the author means when he asserts that "nearly all writers he knows don't like writing"?

Seems contradictory to everything I've heard about the iteration loop between practice, play, and discovery that characterizes mastery in any craft, be it programming, writing, music, sports, and so on.

Re: The computers are getting better at writing

#134

Oh wow, just seeing this post made it to the front page. Founder of Sudowrite here, happy to answer questions.

I don't mean to put on you the spot, but I have been wondering about this: Do you think that such tools will have the eyes of the regulators on top of them due to copyright issues, since the data used could have a lot of copyrighted material? I haven't yet made my mind about this: is the output original content, or just shuffled proprietary content?

(Caveat: IANAL) As far as I know the issue of copyright and AI generated content hasn’t been tested in the courts yet. There are two issues:

1) The author inadvertently uses output from GPT-3 that happens to be close enough to copyrighted material to be a problem.

This is why we recommend to authors do two things: a) enter as much of their own original work to prompt Sudowrite and b) edit and incorporate the AI output organically into their work. The chance that GPT-3 outputs verbatim copyright work in that case is small. As we grow, we also plan to incorporate plagiarism checks in the background.

2) Someone could challenge the author’s copyright on a piece of work that was primarily written by GPT-3.

This has not been tested in the courts, and I suspect it will come to head in the next few years. For me, the question of who owns copyright is dependent on how much work the author has done to the AI output (there are some parallels to the Monkey Selfie case [2]). For example, Stephen Marche (the author of the new yorker post from OP) also wrote a short story [1] using Sudowrite, in which he claims that 17.1% of the final words were from the AI. But of course, he has done non-trivial amounts of work to edit and incorporate those words and made it his own. So in this case, he def deserves ownership. All of our Sudowrite authors use the AI similarly in their process.

However, what if an author uses vanilla GPT-3 with no humans int the loop to generate hundreds of biographies of famous people based on wikipedia entries? I don’t have a good answer to this. It shouldmay come down to how much the author has done to guide GPT-3 and/or pre-post-process it.

Here’s further reading:

http://rutgerslawreview.com/wp-content/uploads/2017/07/Rober...

https://core.ac.uk/download/pdf/77599763.pdf

[1] https://lareviewofbooks.org/article/the-thing-on-the-phone/

[2] https://www.npr.org/sections/thetwo-way/2017/09/12/550417823...

Re: The computers are getting better at writing

#135

See also my experiments with GPT-3 on sane prompts, which have wildly varying quality even after generating them in bulk: https://github.com/minimaxir/gpt-3-experiments Creative writing hasn't been one of the super-hyped use cases by OpenAI for the OpenAI API outside of AI Dungeon, surprisingly. For just random generation, the necessary curation can detract from the time-savings advantages. (as an aside, the API is a…

Generating M:tG cards is interesting to me. Can we talk about this a bit? :) My degree dissertation was an M:tG expert system that included a hand-crafted parser and generator for (a small subset of) ability text [1]. My Masters' thesis was a grammar induction system trained on M:tG cards [2]. And I finally managed to sneak some M:tG grammar induction in my papers as a PhD student [3]. I had a quick look at the colab…

The outputs on pageload are just a random uncurated set I ran. I never claimed smaller models like the MtG models were immune to the same curation issues as larger text generation models.

As someone who did try to reproduce RoboRosewater with LSTMs years ago, I can say using GPT-2 overall has a higher signal-to-noise ratio in terms of generating interesting cards which is a win.

Re: The computers are getting better at writing

#136
post #45

GPT-3-generated writing is a combination of comically absurd and incredibly deep. The weak infants of the woman next door who had to be picked up with the tongs and thrown into the dustbin is quite silly but then there's He wanted to crawl away from it, but there was no place to go. which is a surprisingly deep statement about the revulsion he feels towards himself. That a computer can generate that says very uncomfo…

I was actually taken in by the infants-in-the-dustbin image. I thought it was gruesome, then I reflected on how infant mortality has decreased and what a better time we're living in. This article was uncanny valley for me.

Writers use all kinds of tools to "cheat", up to and including not actually writing. Mining computer-generated nonsense or coherent-but-empty prose for the odd accidentally-excellent phrase or sentence might end up being just another. Hell, it may already be.

Re: The computers are getting better at writing

#137

Earlier quoted context omitted.

Generating M:tG cards is interesting to me. Can we talk about this a bit? :) My degree dissertation was an M:tG expert system that included a hand-crafted parser and generator for (a small subset of) ability text [1]. My Masters' thesis was a grammar induction system trained on M:tG cards [2]. And I finally managed to sneak some M:tG grammar induction in my papers as a PhD student [3]. I had a quick look at the colab…

The outputs on pageload are just a random uncurated set I ran. I never claimed smaller models like the MtG models were immune to the same curation issues as larger text generation models. As someone who did try to reproduce RoboRosewater with LSTMs years ago, I can say using GPT-2 overall has a higher signal-to-noise ratio in terms of generating interesting cards which is a win.

>> The outputs on pageload are just a random uncurated set I ran. I never claimed smaller models like the MtG models were immune to the same curation issues as larger text generation models.

Are you sure? I reloaded the page once, before commenting above, but the output didn't change. I keep reloading now and the same set of cards is displayed. I reloaded with Shift-F5 to be sure ish.

>> As someone who did try to reproduce RoboRosewater with LSTMs years ago, I can say using GPT-2 overall has a higher signal-to-noise ratio in terms of generating interesting cards which is a win.

I don't know what an "interesting" card is, when we're talking about automatically generated cards. If a card looks too much like an existing card (e.g. if it has abilities that you can find on cards in a real-world set) then it's not very interesting. If it's an ability that hasn't been seen before but is only a variant of an existing ability ("Destroy target donkey") it's still not very interesting. To generate really "interesting" cards a generator must go beyond what's in the M:tG corpus, but not so far out that the abilities don't make sense anymore because that's not "interesting", just "random". The ones generated by your project are not far out enough to be what I'd call "interesting" but I think if you tried to make them more interesting they'd also become less coherent than they are currently (looking at the same few results I keep getting anyway).

(This is not very harsh criticism I hope.)

Re: The computers are getting better at writing

#138

Earlier quoted context omitted.

>> One clear example: attributive nouns. Specific examples, please? I'm curious :)

In a case where a noun is a modifier of the following word, Italians tend to prefer specificity , therefore falling for the “noun + of…” construct. A chicken soup bowl is easy to understand, but for an Italian that’s a “piatto di brodo di pollo”, so an Italian native would tend to feel like “a bowl of chicken soup” is the better choice. In this case there’s no real mistake (and for “bowl of” you might argue there’s a…

Thanks for the example!

Wouldn't "piatto di brodo di pollo" translate to "a bowl of soup of chicken", as a more unnatural English sentence?

Greek is similar in that respect and I catch myself sometimes lapsing into such more micro-managed speaking, and I also noticed it in other Greeks (perhaps a few Italians also).

Re: The computers are getting better at writing

#139
post #69

Earlier quoted context omitted.

If you want to echo the last paragraph of the article, the system also doesn’t know when to stop. Which leads to a simple improvement: teach the model when to stop, what sections to remove. There’s already great ML-based summary engines, and systems able to answer questions from descriptions.

I wonder if that could be a viable approach, having a separate "writer" and "editor" AI

GPT-3 is the editor. The hard part of writing is getting a decent idea onto the page. GPT-3 can't do that, it doesn't think for itself. It is good at the second part of writing, creating a coherent string of sentences describing an idea. But since GPT-3 doesn't have decent ideas, the final product doesn't end up being interesting either. The most interesting use cases for GPT-3 involve pairing it with a human who will contribute the actual thinking to the project, with GPT being there to come up with interesting turns of phrase and nice conjoining sentences. Having an AI in the driver seat is still a long way off.

Re: The computers are getting better at writing

#140

Earlier quoted context omitted.

I don't mean to put on you the spot, but I have been wondering about this: Do you think that such tools will have the eyes of the regulators on top of them due to copyright issues, since the data used could have a lot of copyrighted material? I haven't yet made my mind about this: is the output original content, or just shuffled proprietary content?

(Caveat: IANAL) As far as I know the issue of copyright and AI generated content hasn’t been tested in the courts yet. There are two issues: 1) The author inadvertently uses output from GPT-3 that happens to be close enough to copyrighted material to be a problem. This is why we recommend to authors do two things: a) enter as much of their own original work to prompt Sudowrite and b) edit and incorporate the AI outpu…

Thank you for the insightful reply! Will definitely read more into it!
Post reply on HN