> The system ingests raw newspaper scans and uses a multi-step LLM pipeline to generate the daily edition This is neat! But I wonder about longevity of the project if it relies on scanning newspapers. Do you have an endless suply? Perhaps there is some digital archive you could use?
My Wikipedia Library membership gives me access to some cool resources that might be of interest: - Arcanum is the largest and continuously expanding digital periodical database from Eastern Europe, which contains scientific and specialized journals, encyclopaedias, weekly and daily newspapers and more - NewspaperARCHIVE.com is an online database of digitized newspapers, with over 2 billion news articles; coverage ex…
My first edit was 20 years ago this month and at my current pace I'll be able to access that in another 588 years.
Is there some other way to pay [Wikipedia/WMF] for access to that bundle?