Live data from Hacker News

AniSora: Open-source anime video generation model

komiko.app

211–220 of 232 posts

Re: AniSora: Open-source anime video generation model

#211
post #192

Earlier quoted context omitted.

That makes no sense, neither legally nor philosophically. > Language developed independently of translators. And it also developed independently of writers and poets. > Artists shape the space in which they’re generating output. Not writers and poets, apparently. And so maybe not even artists, who typically mostly painted book references. Color perception and symbolism developed independently of professional artists,…

If you think memes are art too and we lack shared consumption of art due to personalization, you clearly don't have kids into YouTube or Minecraft or Frozen, or ...

I don't get what you are trying to say here? Yes, memes are arts however foreign it might be to older folks. To your second point, you know about Frozen because everyone else also watches that. We are about to lose that if there are 1 million variations of "Frozen"-esque movie that people can watch.

I don't think have an AI partner that is trained from zero from childhood to adulthood with goals such as "make me laugh" is too far fetched. The problem is you will never be able to connect with this child because the AI is feeding it insanely obscure, highly specific videos that matches the neurons of the kid perfectly.

Re: AniSora: Open-source anime video generation model

#213
post #96

Earlier quoted context omitted.

I think the “paper rock cross blade” short films by Corridor is absolute great and can by all accounts be called art and if they make a 3rd they will probably use this model. In terms of losing styles, that is already been happening for ages. Disney moved to xeroxing instead of inking, changed the style because inking was “too hard”. In the late 90s/early 2000s we saw a burst of cartoons with a flash animation style…

I disagree with the positive characterisation. Those videos have a funny schtick of exaggerating anime tropes for a couple of minutes and that’s the extent of it. The animation is all over the place, reactions, expressions, mouth movements often fail, style changes from frame to frame. It maybe kind of works precisely because it’s a short exaggerated parody and we have a high tolerance for flaws in comedy, but even t…

I think it’s a successful creative endeavor for two reasons:

They took the weaknesses of last years style transfer models and used them as a style, working around and with it’s shortcomings and weaknesses. That is a far cry from “type a prompt and be do e with it”.

Secondly I think the story is fun and the whole thing is fun, not in a will smith eats spaghetti kind of why but fun as in an actually fun short film.

I think it shows that AI can be a tool that empowers creativity and creative work and more than a power point stock photo generator.

Re: AniSora: Open-source anime video generation model

#214
post #209

Earlier quoted context omitted.

A translation is absolutely under copyright. It is a creative process after all. This means a book can be in public domain for the original text, because it's very old, but not the translation because it's newer. For example Julius Caesar's "Gallic War" in the original latin is clearly not subject to copyright, but a recent English translation will be.

So if a machine was to do the translation, should that also be considered a creative work? If not, that would put pressure on production companies to use machines so they don’t have to pay future royalties

> So if a machine was to do the translation, should that also be considered a creative work?

No, but it will be derived work covered by the same copyright as original.

The quality of human translation is better, for now.

Re: AniSora: Open-source anime video generation model

#215
post #129

Earlier quoted context omitted.

The issue is whether the artists creating things for love of the game will be crowded out even further by studios churning out slop (or in HN terms, Minimal Viable Products) for cash. There are probably 15 disposable reality TV shows created for every scripted sitcom or drama that needs good writers, set designers and directors.

They already are ; have been for decades now. AI is amplifying this, true, but art done for love and for money are already pretty much disjoint ventures, and in areas where they mix (like TV shows), it's an uphill battle for the artist - and they're not always right, either; a good show is more than just great writing or beautiful art.

I'd argue that better graphical genai is a solution for this.

It's a fact of life that creative production companies will always attempt to optimize costs -- which means most efficiently using any human labor.

In the 00s/10s, animation studios especially tried to do this with... mixed results (coughToeicough)

More capable models should allow better keyframe-to-keyframe animation.

Re: AniSora: Open-source anime video generation model

#216
post #99

Earlier quoted context omitted.

Ambiguities are not a good argument against laws that still have positive outcomes. There are very few laws that are not giant ambiguities. Where is the line between murder, self-defense and accident? There are no lines in reality. (A law about spectrum use, or registered real estate borders, etc. can be clear. But a large amount of law isn’t.) Something must change regarding copyright and AI model training. But it d…

> There are very few laws that are not giant ambiguities. Where is the line between murder, self-defense and accident? There are no lines in reality. These things are very well and precisely defined in just about every jurisdiction. The "ambiguities" arise from ascertaining facts of the matter, and whatever some facts fits within a specific set of set rules. > Something must change regarding copyright and AI model tr…

> These things are very well and precisely defined in just about every jurisdiction.

Yes, we have lots of wording attempting to be precise. And legal uses of terms are certainly more precise by definition and precedent than normal language.

But ambiguities about facts are only half of it. Even when all the facts appear to be clear, human juries have to use their subjective human judgement to pair up what the law says, which may be clear in theory, but is often subjective at the borders, vs. the facts. And reasonable people often differ on how they match the two up in many borderline cases.

We resolve both types of ambiguities case-by-case by having a jury decide, which is not going to be consistent from jury to jury but it is the best system we have. Attorneys vetting prospective jurors are very much aware that the law comes down to humans interpreting human language and concepts, none of which are truly precise, unless we are talking about objective measures (like frequency band use).

---

> it is the question of what constitutes a derivative

Yes, the legal side can adapt.

And the technical side can adapt too.

The problem isn't that material was trained on, but that the resulting model facilitates reproducing individual works (or close variations), and repurposing individual's unique styles.

I.e. they violate fair use by using what they learn in a way that devalues other's creative efforts. Being exposed to copyrighted works available to the public is not the violation. (Even though it is the way training currently happens that produces models that violate fair use.)

We need models that one way or another, stay within fair use once trained. Either by not training on copyrighted material, or by training on copyrighted material in a way that doesn't create models that facilitate specific reproduction and repurposing of creative works and styles.

This has already been solved for simple data problems, where memorization of particular samples can be precluded by adding noise to a dataset. Important generalities are learned, but specific samples don't leave their mark.

Obviously something more sophisticated would need to be done to preclude memorization of rich creative works and styles, but a lot of people are motivated to solve this problem.

Re: AniSora: Open-source anime video generation model

#217
post #209

Earlier quoted context omitted.

A translation is absolutely under copyright. It is a creative process after all. This means a book can be in public domain for the original text, because it's very old, but not the translation because it's newer. For example Julius Caesar's "Gallic War" in the original latin is clearly not subject to copyright, but a recent English translation will be.

So if a machine was to do the translation, should that also be considered a creative work? If not, that would put pressure on production companies to use machines so they don’t have to pay future royalties

Well that's the real question, isn't it?

Our current best technology, LLMs, are good enough for translating an email or meeting transcript and getting the general message across. Anything more creative, technical, or nuanced, and they fall apart.

Meaning for anything of value like books, plays, movies, poetry, humans will necessarily be part of the process: coaxing, prompting, correcting...

If we consider the machine a tool, it's easy, the work would fall under copyright.

If we consider the machine the creator, then things get tricky. Are only the parts reworked/corrected under copyright? Do we consider under copyright only if a certain portion of the work was machine generated? Is the prompt under copyright, but not its output?

Without even getting into the issue of training data under copyright...

There is some movement regarding copyright of AI art, legislation being drawn up and debated in some countries. It's likely translations would be impacted by those decisions.

Re: AniSora: Open-source anime video generation model

#218
post #182

Earlier quoted context omitted.

how did you come up with it?

Let's say I used an AI. Actually I browsed unrelated word lists in Onelook - I think I was on synonyms for "confusion" - until I remembered the word atelier because it turns up in fantasy anime a lot. But that's a kind of machine assistance, so let's pretend it was an LLM, if that helps with where you're taking this. Now what?

it's exactly where I was going with it. Just because you used tools or have some mechanical process that you used along the way doesn't distract from the fact that you might be able to come up with a good pun. So if you have a computer model that can predict whether people will like a given pun (or a given response for any other domain), then the work produced will have the same effect on the audience. You could try to ascribe the prime mover title to the human that says "write me a pun, a good one!", but ultimately the machine is producing art.

The machine could also just produce lots of examples and test them on a large number of humans - in which case none of them individually is the artist, but the art is still being produced.

Re: AniSora: Open-source anime video generation model

#219
post #92

Earlier quoted context omitted.

While LLMs are pretty good, and likely to improve, my experience is OpenAI's offerings *absolutely* make stuff up after a few thousand words or so, and they're one of the better ones. It also varies by language. Every time I give an example here of machine translated English-to-Chinese, it's so bad that the responses are all people who can read Chinese being confused because it's gibberish. And as for politics, as Gr…

> While LLMs are pretty good, and likely to improve, my experience is OpenAI's offerings absolutely make stuff up after a few thousand words or so, and they're one of the better ones. That's not how you get good translations from off-the-shelf LLMs! If you give a model the whole book and expect it to translate it in one-shot then it will eventually hallucinate and give you bad results. What you want is to give it a s…

>> While LLMs are pretty good, and likely to improve, my experience is OpenAI's offerings absolutely make stuff up after a few thousand words or so, and they're one of the better ones.

> That's not how you get good translations from off-the-shelf LLMs! If you give a model the whole book and expect it to translate it in one-shot then it will eventually hallucinate and give you bad results.

You're assuming something about how I used ChatGPT, but I don't know what exactly you're assuming.

> What you want is to give it a small chunk of text to translate, plus previously translated context so that it can keep the continuity

I tried translating a Wikipedia page to support a new language, and ChatGPT rather than Google translate because I wanted to retain the wiki formatting as part of the task.

LLM goes OK for a bit, then makes stuff up. I feed in a new bit starting from its first mistake, until I reach a list at which point the LLM invented random entries in that list. I tried just that list in a bunch of different ways, including completely new chat sessions and the existing session, it couldn't help but invent things.

> In an open ended questions - sure. But that doesn't apply to translations where you're not asking the model to come up with something entirely by itself, but only getting it to accurately translate what you wrote into another language.

"Only" rather understates how hard translation is.

Also, "explain this in Fortnite terms" is a kind of translation: https://x.com/MattBinder/status/1922713839566561313/photo/3

Post reply on HN