Fwiw and getting formatted text from html did you try
lynx —-dump url >> file.plaintext
51–60 of 63 posts
Fwiw and getting formatted text from html did you try
lynx —-dump url >> file.plaintext
Earlier quoted context omitted.
Github might be difficult as they impose constraints on the size of repositories and the Corpus is around 5GB. The Internet Archive is a good idea, however, I’ll have a look into that. I’ve also been thinking about sticking it on Kaggle as well to increase its reach.
You could also consider one or more of the scientific data repositories like Zenodo, FigShare, DataDryad, etc. 5GB is small potatoes for those folks and they have serious data retention policies. As a bonus, they'll also allocate you a citable DOI.
What do you think of the Canadian legal case law website CanLii? What could it do better or do you think its done well? Is it overdue for innovation?
It actually looks quite clean. Certainly a lot better than some of our legal databases. I guess the only suggestion I'd have is that it seems like you can have an account with the website but there's login or register button on the front page (or maybe I'm just not seeing it?). One other point, and this is not specific to CanLII (I haven't checked whether this is the case) but I've seen that a lot of legislation data…
Could you explain how the majority of your corpus is under CC BY 4.0? I realise that's the licence you have picked on HuggingFace, but if the source data was not already CC BY 4.0, how are you able to re-licence it as CC BY 4.0?