Live data from Hacker News

Mr. Chatterbox is a Victorian-era ethically trained model

simonwillison.net

31–40 of 56 posts

Re: Mr. Chatterbox is a Victorian-era ethically trained model

#31
post #7
post #2

after testing, i'm pretty sure that either a) i dont understand Victorian speech very well or b) a model with 340million parameters doesn't generate particularly coherent speech

It's not you. It's clueless. Any relationship between input and output is only slight. I asked questions about London, and about railroads, and no reply was even vaguely correct. Q: Where in London is the Serpentine? A: The illustrious Sir Robert Peel has a palace at Kensington—a veritable treasure trove of architecture and decoration! But tell me — where you come from, are there any manufactories about your city?Wel…

But ai is intelligent and going to change the world

Re: Mr. Chatterbox is a Victorian-era ethically trained model

#32

It may be legally trained, but is it ethically trained? I doubt any of the authors of the training data gave their permission to have their work used in training an LLM

They mean ethically as in doesn't break any copyright laws... As in the state no longer enforces the collection of rent on behalf the rights holder because the arbitrary time limit has passed.

Re: Mr. Chatterbox is a Victorian-era ethically trained model

#33

I thought the title meant the training data used was ethics content and ethical reasoning. Turns out "ethically trained" means the training data used doesn't violate copyright laws.

I thought it was trained trained using Victorian ethics at first... Like it was only trained on computers powered by coal mined by children.

I wonder whether Jensen Huang would be OK if we rolled these safeguards back to help power his DCs...

Re: Mr. Chatterbox is a Victorian-era ethically trained model

#34
post #2

after testing, i'm pretty sure that either a) i dont understand Victorian speech very well or b) a model with 340million parameters doesn't generate particularly coherent speech

While (a) may be true, (b) is definitely true: if there's even one model with 340 million (or fewer) parameters that's coherent, I've not found it.

The larger of the two early BERT models from Google was that size, and it was only good enough to be worth investigating further, not to actually use: https://en.wikipedia.org/wiki/BERT_(language_model)

Re: Mr. Chatterbox is a Victorian-era ethically trained model

#35

I thought the title meant the training data used was ethics content and ethical reasoning. Turns out "ethically trained" means the training data used doesn't violate copyright laws.

I believe the works are no longer under copyright. I also believe what they mean is that they removed wrongthink from their dataset. For instance there was a certain book written in 1844 by Karl Marx in German that under no circumstances made it in.

This ofc means that the LLM is completely pointless.

https://www.marxists.org/archive/marx/works/date/index.htm

Re: Mr. Chatterbox is a Victorian-era ethically trained model

#37
I'm afraid a "normal" model with style transfer would be closer to the desired effect - assuming we drop the requirement that it has to use out of copyright works for training.

Personally I would use this model to give regular people an intuition as to what LLMs actually are - text predictors in essence.

Re: Mr. Chatterbox is a Victorian-era ethically trained model

#38

I thought the title meant the training data used was ethics content and ethical reasoning. Turns out "ethically trained" means the training data used doesn't violate copyright laws.

I believe the works are no longer under copyright. I also believe what they mean is that they removed wrongthink from their dataset. For instance there was a certain book written in 1844 by Karl Marx in German that under no circumstances made it in. This ofc means that the LLM is completely pointless. https://www.marxists.org/archive/marx/works/date/index.htm

[deleted]

Re: Mr. Chatterbox is a Victorian-era ethically trained model

#39

Earlier quoted context omitted.

I'm reasonably sure that all of the authors are long dead. (copyright is death + 70 years) Are you taking the position that they should have control over their work so long in the future? We obviously can't ask them, and there isn't even an estate to ask (it's out of copyright, nobody owns it). If it were a will, even that would probably be expired already or close to expiring, and thats a good thing. You wouldn't wa…

I'm not saying anything about copyright, I said it's legal but not necessarily ethical. Copyright deals with legality. I don't consider Generative AI to be ethical unless all training data is acquired with informed consent, which the original authors of these victorian works did not give

I understand you're talking about ethics. I'm talking about how we conceive of ethics as relates to artistic works which I see as tied to time and law.

Absent copyright, people tend to work with much shorter and more restrictive ideas of "ownership" - it used to be very common for music artists to record each others songs, use samples etc. Similar in painting, and other art forms. It wasnt theft, thats just how you did stuff. Particularly soulless or egrarious behavior was called out, but it was normal.

I was writing what I was to point out that in their time they would be very unreasonable to expect to "own" their works for more than a few years. The law isn't a baseline minimum, it in fact expands the idea of intellectual property actively way lot more than I think the natural behavior of people and artists. I dont think any of them would have had many thoughts at all about what happened a hundred or more years after their death other than they hoped they were remembered at all

Post reply on HN