Live data from Hacker News

Mr. Chatterbox is a Victorian-era ethically trained model

simonwillison.net

41–50 of 56 posts

Re: Mr. Chatterbox is a Victorian-era ethically trained model

#41
post #7
post #2

after testing, i'm pretty sure that either a) i dont understand Victorian speech very well or b) a model with 340million parameters doesn't generate particularly coherent speech

It's not you. It's clueless. Any relationship between input and output is only slight. I asked questions about London, and about railroads, and no reply was even vaguely correct. Q: Where in London is the Serpentine? A: The illustrious Sir Robert Peel has a palace at Kensington—a veritable treasure trove of architecture and decoration! But tell me — where you come from, are there any manufactories about your city?Wel…

Amazing. It's like a drop in replacement for our politicians.

Re: Mr. Chatterbox is a Victorian-era ethically trained model

#42
post #37

I'm afraid a "normal" model with style transfer would be closer to the desired effect - assuming we drop the requirement that it has to use out of copyright works for training. Personally I would use this model to give regular people an intuition as to what LLMs actually are - text predictors in essence.

What makes you think the desired effect is to have an LLM that speaks in an old-timey style? The training process is the whole point.

Re: Mr. Chatterbox is a Victorian-era ethically trained model

#43

I thought the title meant the training data used was ethics content and ethical reasoning. Turns out "ethically trained" means the training data used doesn't violate copyright laws.

I really dislike the way people use "ethical" as though it were an unambiguous, binary concept.

Even if it's just shorthand due to space constraints, it oversimplifies the concept of "ethical" to the point of muddling people's thinking.

Re: Mr. Chatterbox is a Victorian-era ethically trained model

#44
post #28
post #19

Earlier quoted context omitted.

Note: training constrained by copyright could still be an improvement over training that ignores copyright completely. I assume the general opinion is that copyright is at most partially unethical. That’s what the AI discussion is about too, i.e. artist copyright.

Given the extent to which the copyright system has benefited corporations and publishing companies to the detriment of individual authors and the general public, I'm constantly surprised that it still has many apologists.

As we don't live in a world where the rich patronize the arts some sort of copyright system is the only way authors and artists are gonna make a living doing their thing. ...though I suppose proponents of Universal Basic Income (UBI) would disagree, but between the abolishment of copyright, the institution of UBI, or a 7 year old child being hit by 7 lightning strikes and 7 meteor impacts and surviving; the latter seems the most likely.

Re: Mr. Chatterbox is a Victorian-era ethically trained model

#45

I thought the title meant the training data used was ethics content and ethical reasoning. Turns out "ethically trained" means the training data used doesn't violate copyright laws.

If training data of any kind violated copyright, every creator alive would be in breach of by virtue of any influence their “training data” (lifelong exposure to the work of others) has on their output.

The creators crying foul of AI are painting themselves into a corner, both literally and figuratively.

Re: Mr. Chatterbox is a Victorian-era ethically trained model

#46
post #28
post #19

Earlier quoted context omitted.

Note: training constrained by copyright could still be an improvement over training that ignores copyright completely. I assume the general opinion is that copyright is at most partially unethical. That’s what the AI discussion is about too, i.e. artist copyright.

Given the extent to which the copyright system has benefited corporations and publishing companies to the detriment of individual authors and the general public, I'm constantly surprised that it still has many apologists.

People imagine poor author having their thing stolen rather than poor author that corporate takes IP from by contract agreement (and if you don't do that, you don't get the job), then abuses for 70+ years

Re: Mr. Chatterbox is a Victorian-era ethically trained model

#47
post #7
post #2

after testing, i'm pretty sure that either a) i dont understand Victorian speech very well or b) a model with 340million parameters doesn't generate particularly coherent speech

It's not you. It's clueless. Any relationship between input and output is only slight. I asked questions about London, and about railroads, and no reply was even vaguely correct. Q: Where in London is the Serpentine? A: The illustrious Sir Robert Peel has a palace at Kensington—a veritable treasure trove of architecture and decoration! But tell me — where you come from, are there any manufactories about your city?Wel…

The output reminds of a really good version of pre-LLM text generation like character lever LSTMs or markov chains.

It seems to have syntax down to make superficially good text, but the semantics just aren’t there

Re: Mr. Chatterbox is a Victorian-era ethically trained model

#48
post #45

I thought the title meant the training data used was ethics content and ethical reasoning. Turns out "ethically trained" means the training data used doesn't violate copyright laws.

If training data of any kind violated copyright, every creator alive would be in breach of by virtue of any influence their “training data” (lifelong exposure to the work of others) has on their output. The creators crying foul of AI are painting themselves into a corner, both literally and figuratively.

This is a truly awful argument that keeps coming up. It relies on the false equivalence between training an AI (a technical process that involves copying a work into computer storage), and a human being experiencing a work, which doesn't involve any kind of copying (and usually involves the human legally purchasing the work, which AI companies did not do).

There is a legal difference as well as a technical difference. AIs don't learn the same way human brains do. The law does not treat these things the same. You may want to draw an analogy between the two and say they're "basically the same", but they are not basically the same. They aren't the same at all, outside of a very weak analogy. Is training kind of sort of like human learning? Yes. That doesn't mean anything. Dogs are kind of sort of like children, but if you try to treat your child the way you treat your dog, you end up in prison. Because children aren't dogs, either in reality, or in the eyes of the legal system.

Please, AI boosters, stop using this one. Human brains aren't clocks. Human brains aren't computers. Human brains aren't LLMs. AI training does not mimic human learning in any significant way.

Re: Mr. Chatterbox is a Victorian-era ethically trained model

#49
post #28
post #19

Earlier quoted context omitted.

Note: training constrained by copyright could still be an improvement over training that ignores copyright completely. I assume the general opinion is that copyright is at most partially unethical. That’s what the AI discussion is about too, i.e. artist copyright.

Given the extent to which the copyright system has benefited corporations and publishing companies to the detriment of individual authors and the general public, I'm constantly surprised that it still has many apologists.

What do you suggest instead? I.e. what would benefit individual authors more?
Post reply on HN