Live data from Hacker News

Mr. Chatterbox is a Victorian-era ethically trained model

simonwillison.net

51–56 of 56 posts

Re: Mr. Chatterbox is a Victorian-era ethically trained model

#52

I thought the title meant the training data used was ethics content and ethical reasoning. Turns out "ethically trained" means the training data used doesn't violate copyright laws.

It would be interesting to talk with a victorian-era chatbot, including victorian-era ethics. would be interesting to see how much divergence from modern era ethics it would have.

Re: Mr. Chatterbox is a Victorian-era ethically trained model

#54
post #20

I wonder if you could generate synthetic Victorian-era training data.

Certainly – use a bigger general purpose model to create more works 'in the style of'.

I don't just mean stylistically, I mean with clearly constrained knowledge, etc.

Re: Mr. Chatterbox is a Victorian-era ethically trained model

#55

Earlier quoted context omitted.

Do you know what public domain is?

Yes. As I said, it's legally trained, if all the data is in the public domain, but legal != ethical. I think the current legal defence of modern LLMs is that it's transformative so copyright doesn't apply, and I certainly wouldn't call them ethical

Do you think the concept of public domain is not ethical?

Re: Mr. Chatterbox is a Victorian-era ethically trained model

#56
post #7
post #2

after testing, i'm pretty sure that either a) i dont understand Victorian speech very well or b) a model with 340million parameters doesn't generate particularly coherent speech

It's not you. It's clueless. Any relationship between input and output is only slight. I asked questions about London, and about railroads, and no reply was even vaguely correct. Q: Where in London is the Serpentine? A: The illustrious Sir Robert Peel has a palace at Kensington—a veritable treasure trove of architecture and decoration! But tell me — where you come from, are there any manufactories about your city?Wel…

From the author's writeup:

>the final pre-trained model came out to about 340 million parameters, and had a final validation bpb of 0.973. The pretraining process took about five hours on-chip, and cost maybe $35. I had my pretrained model, trained in 6496 steps. Things were proceeding swiftly, and cheaply!

GPT-3 had 175,000 million parameters. The smallest of the Gemma 4 models released today clock in at 5,000 million parameters, and I would bet that Google trained them for more than five hours. Just too small and not trained for enough time. A fun art project but not a functional LLM.

Post reply on HN