Live data from Hacker News

The British Library puts 1M newspaper pages online for free

ianvisits.co.uk

21–30 of 60 posts

Re: The British Library puts 1M newspaper pages online for free

#21

This is great. Historical newspapers are one of the largest corpora of information that has yet to be adequately brought on line. In the U.S. the Library of Congress has digitized a fair number, but at the state and local level it's really hit or miss. Some states such as California and New York have put quite a bit on line, but many others rely on individual towns and historical societies. Different pay services cov…

I wonder what ever happened to all the newspapers that were fed into services like CompuServe in the 80's.

I dated a newspaper reporter during that era, and all of her stuff went into the online services. But her newspaper's current online archive only goes back to about 2005, even for subscribers.

Re: The British Library puts 1M newspaper pages online for free

#22
Do anyone know any existing effort on converting these scanned image to text corpus ( probably a new OCR model needed to be developed on these old text ) ? I think it would be more usable if they are in text form in terms of search and research purpose.

Re: The British Library puts 1M newspaper pages online for free

#23

Earlier quoted context omitted.

Do they really have automated systems searching for 50+ year old newspapers? Probably a lot of them haven't even been digitised before, so it would be impossible to search for them in an automated way.

From personal experience, I can tell you that there are automated systems searching for such pieces of art. As for newspapers, all it takes is one copyright troll to realize it's happening, and suddenly there's lawsuit settlements everywhere, making him rich.

For that to happen they'd have to be the owners of the Copyright, they're not exactly patent trolls who can use the law creatively.

The Newspapers are largely regional (and mostly defunct) papers. Who at the Derby Evening Telegraph is going to waste lawyers fees getting a Library to remove 100year old content?

Re: The British Library puts 1M newspaper pages online for free

#24
post #2

> The British Library keeps to a ‘safe date’ when determining when a newspaper can be considered to be entirely out-of-copyright, which is 140 years after the date of publication. It's depressing that copyright has been extended so far it's now longer then any single individual's possible lifetime.

Oh. It's worse than that. Copyright apparently extends back to the middle ages:

https://stiobhart.net/2021-04-8-bubble-trouble/

Re: The British Library puts 1M newspaper pages online for free

#25
post #13

Direct link to the site: https://www.britishnewspaperarchive.co.uk/

Thanks, it's nicely parametrised and easy to browse; shame they decided to require free registration in order to actually view anything though.

You only get to see three pages for free. The original submission title is really misleading:

  >Try the British Newspaper Archive for FREE

  >View 3 pages FREE when you register to help you get started

Re: The British Library puts 1M newspaper pages online for free

#26
post #2

> The British Library keeps to a ‘safe date’ when determining when a newspaper can be considered to be entirely out-of-copyright, which is 140 years after the date of publication. It's depressing that copyright has been extended so far it's now longer then any single individual's possible lifetime.

This needs to change, big time. There is almost no cash value to an article one day later, yet we completely impoverish the public domain for its sake. Not only are creators of valuable works is usually pretty distant from direct ownership anyway, there's no possible way for them to profit directly from this work. The only way a spotify-like deal works is because copyright ownership is conglomerated. IMO, public (esp…

> There is almost no cash value to an article one day later, yet we completely impoverish the public domain for its sake.

We're not imposing publishing restrictions on past works in order to preserve their cash value. We're imposing publishing restrictions on past works in order to stop them from competing with present works. It keeps the cash value of present works up.

Re: The British Library puts 1M newspaper pages online for free

#27

This is great. Historical newspapers are one of the largest corpora of information that has yet to be adequately brought on line. In the U.S. the Library of Congress has digitized a fair number, but at the state and local level it's really hit or miss. Some states such as California and New York have put quite a bit on line, but many others rely on individual towns and historical societies. Different pay services cov…

> Historical newspapers are one of the largest corpora of information that has yet to be adequately brought on line.

Not just information, but works of art as well!

A few years ago, I trained an image classifier to help me find Krazy Kat comics in newspaper archives. In the process of doing that, I came across a shocking amount of other comics and artwork. I was honestly surprised to see how many amazing illustrations and comics are just sitting in newspaper archives, waiting to be rediscovered.

Re: The British Library puts 1M newspaper pages online for free

#29

Earlier quoted context omitted.

This needs to change, big time. There is almost no cash value to an article one day later, yet we completely impoverish the public domain for its sake. Not only are creators of valuable works is usually pretty distant from direct ownership anyway, there's no possible way for them to profit directly from this work. The only way a spotify-like deal works is because copyright ownership is conglomerated. IMO, public (esp…

> There is almost no cash value to an article one day later, yet we completely impoverish the public domain for its sake. We're not imposing publishing restrictions on past works in order to preserve their cash value. We're imposing publishing restrictions on past works in order to stop them from competing with present works. It keeps the cash value of present works up.

That makes even less sense, though. Yesterday's news does not really compete with today's news, they're effectively in separate markets.

Re: The British Library puts 1M newspaper pages online for free

#30

Earlier quoted context omitted.

> There is almost no cash value to an article one day later, yet we completely impoverish the public domain for its sake. We're not imposing publishing restrictions on past works in order to preserve their cash value. We're imposing publishing restrictions on past works in order to stop them from competing with present works. It keeps the cash value of present works up.

That makes even less sense, though. Yesterday's news does not really compete with today's news, they're effectively in separate markets.

So? Yesterday's books, movies, and music compete with today's books, movies, and music. Nobody cares about the news one way or the other. So the news gets treated just like everything else.
Post reply on HN