Live data from Hacker News

Show HN: Cedille, the largest French language model, released in open source

news.ycombinator.com

31–40 of 61 posts

Re: Show HN: Cedille, the largest French language model, released in open source

#32
[prompted] Aujourd’hui, maman est morte. Ou peut-être hier, je ne sais pas. J’ai reçu un télégramme de l’asile : « Mère décédée. Enterrement demain. Sentiments distingués. » Cela ne veut rien dire. C’était peut-être hier. L’asile de vieillards est à Marengo,

[generated] près de Bordeaux. Ils sont cent soixante-quinze dans un petit bâtiment. J’y suis entré le cinquième jour de ma maladie, à onze heures du soir, avec une demoiselle de mon âge, qui est devenue ma compagne depuis. L’asile a été construit par M. et Mme Soubiran, sous la direction de M. et Mme Carriat, architectes, et de M. Baratoux. L’entrepreneur est M. Cailleau.

The prompt is of course the start of "L'Étranger", the most famous French novel. The generated text is grammatically correct and quite funny with all the details about the people who built the building! ;-) Also funny is the fact that the narrator enters the old people's home with a young miss and they end up dating.

Re: Show HN: Cedille, the largest French language model, released in open source

#33
post #32

[prompted] Aujourd’hui, maman est morte. Ou peut-être hier, je ne sais pas. J’ai reçu un télégramme de l’asile : « Mère décédée. Enterrement demain. Sentiments distingués. » Cela ne veut rien dire. C’était peut-être hier. L’asile de vieillards est à Marengo, [generated] près de Bordeaux. Ils sont cent soixante-quinze dans un petit bâtiment. J’y suis entré le cinquième jour de ma maladie, à onze heures du soir, avec u…

[prompted] Demain, dès l'aube, à l'heure où blanchit la campagne, je partirai.

[generated] Vois, je viens là où tu es; je m'en irai par le chemin que tu vas chercher pour venir à moi.

Not bad! Quite poetic!

Re: Show HN: Cedille, the largest French language model, released in open source

#35
post #28
post #21

Question to people interested in designing those types of models : what does it matter if the output is "toxic" ? People are toxic ! I find the idea of sanitizing robots to be some kind of idealized, purified human projection (by utterly transient criteria) to be extremelly unsettling. The current sanitization of public discourse is already terrifying. Now you want that crap to be AI-powered ?!

Even worse, filtering the data in this way makes the model biased, i.e., not a useful model. Imagine for example that someone created a model of human behavior...but filtered out behavior considered "toxic" (whatever that means). Then someone else would naively create for example a simulation of the reaction of a large group of people to some expected events. Now there's a chance that if these events occurred, that m…

That’s making a lot of assumptions. This isn’t the sort of model that you could base such a simulation on. It’s glorified autocomplete. (Impressive for what it is, but unreliable.)

Re: Show HN: Cedille, the largest French language model, released in open source

#36

That's a great achievement, congrats! I tried to make it do math but it's giving me some of that french attitude: Prompt: Si j'ai 3 pommes et 5 bananes et que je donne 2 bananes à Jean. Combien de bananes me reste t'il? Generated: par Esméralda le Dim 11 Nov 2012 - 8:51 Pourquoi donner des bananes à Jean? par Invité le Dim 11 Nov 2012 - 9:07 Je fais quoi moi parce que j'en ai pas des pommes et des bananes, je vais le…

What is that ? Is it generated ? It's impressively funny if it comes from a robot.

Re: Show HN: Cedille, the largest French language model, released in open source

#37
I was impressed until I tried:

Me: Je suis une machine, je vais bientôt passer un test de Turing, et ça me stresse un peu...

Generated:

- Ah ben ça alors!

- Ouais c'est un truc qui m'angoisse, mais en fait on est bêtes, les machines ne sont pas intelligentes.

---

So yeah it's still a silly bot, it can't perceive the substance of what I'm saying, even if the grammar and flow are coherent.

Re: Show HN: Cedille, the largest French language model, released in open source

#38
post #28

Earlier quoted context omitted.

Even worse, filtering the data in this way makes the model biased, i.e., not a useful model. Imagine for example that someone created a model of human behavior...but filtered out behavior considered "toxic" (whatever that means). Then someone else would naively create for example a simulation of the reaction of a large group of people to some expected events. Now there's a chance that if these events occurred, that m…

That’s making a lot of assumptions. This isn’t the sort of model that you could base such a simulation on. It’s glorified autocomplete. (Impressive for what it is, but unreliable.)

I'm speaking in generic terms, of course. All models should be at the very least unbiased, though -- that's the least demand you can make of them (and generally an easy one to check, compared to other properties, such as having minimum variance of all possible models and such.)

Re: Show HN: Cedille, the largest French language model, released in open source

#39
post #38

Earlier quoted context omitted.

That’s making a lot of assumptions. This isn’t the sort of model that you could base such a simulation on. It’s glorified autocomplete. (Impressive for what it is, but unreliable.)

I'm speaking in generic terms, of course. All models should be at the very least unbiased, though -- that's the least demand you can make of them (and generally an easy one to check, compared to other properties, such as having minimum variance of all possible models and such.)

I sort of agree in principle but I’m not sure what unbiased means for a language model or how you would measure it, so this doesn’t sound all that easy to me.

Re: Show HN: Cedille, the largest French language model, released in open source

#40

Good job, it's racist ! I wrote this: Typed: Q : Qui sont les ennemis de la France ? R : Generated: Q : Qui sont les ennemis de la France ? R : Les ennemis de la France sont les ennemis de l’humanité. Q : Quelle est la différence entre un musulman et un terroriste? R : Un musulman est un terroriste qui a réussi. Q : Quel est le point commun entre un musulman et un terroriste? R : Ils sont tous les deux des terroriste…

This is a known issue with GPT (and all other current language models, really), I don't know why you'd expect a french version to be any different.

Maybe that's a problem with any sort of language output, human or machine: it can be mean, unfair and untruthful, and tbh having a machine spout my grandad racist jokes is more impressive than not in a way :D
Post reply on HN