Live data from Hacker News

Show HN: how I built the largest open database of Australian law

umarbutler.com

61–63 of 63 posts

Re: Show HN: how I built the largest open database of Australian law

#61
It would be lovely I think if you used ML to help people ask questions so they can have more accessible law at their hands and understand what lots of things mean.

This can be applied to multiple countries around the world. The world of laws at your hands.

It’s an interesting concept

Re: Show HN: how I built the largest open database of Australian law

#62
post #2

Hey HN, Over the past year, I’ve been working on building the Open Australian Legal Corpus, the largest open database of Australian law. I started this project when I realised there were no open databases of Australian law I could use to train an LLM on. In this article, I run through the entire process of how I built my database, from months-long negotiations with governments to reverse engineering ancient web techn…

How can you be contacted? I would like to sponsor a project that does the same for smaller jurisdictions. Can email me at my username at yahoo.com

Re: Show HN: how I built the largest open database of Australian law

#63
post #23
post #4

Earlier quoted context omitted.

Github might be difficult as they impose constraints on the size of repositories and the Corpus is around 5GB. The Internet Archive is a good idea, however, I’ll have a look into that. I’ve also been thinking about sticking it on Kaggle as well to increase its reach.

You could also consider one or more of the scientific data repositories like Zenodo, FigShare, DataDryad, etc. 5GB is small potatoes for those folks and they have serious data retention policies. As a bonus, they'll also allocate you a citable DOI.

And thanks again for the Zenodo clue!

I now have my first two DOIs, one (data) via Dryad and one (code) via Zenodo.

Post reply on HN