Live data from Hacker News

Roman Letters

romanletters.org

11–20 of 21 posts

Re: Roman Letters

#12
post #9

The website looked as any LLM ("AI") generated one, usually via Claude, considering the design that model frequently uses. And it is (300,755++ lines from Claude): https://github.com/CraigVG/roman-letters-network Here, I am sorry, but I just cannot consider it serious nor accountable, since I just cannot trust its data. If all the information there is valid and verified, every single letter and the authors' word afte…

The design is good. It is unoriginal but not every project needs to use an original design.

serious_angel is not contending with you that the design is bad, or that it is bad because it is unoriginal. In fact, they are not even specifically calling out the design.

They have noticed the design, recognized it as the output of an LLM, then proceeded to discover that an LLM was involved in much of the creation of the project. This is an academic project. Whatever the pedigree of the researcher is, this implies to the grandparent that the final result of the work may be amateurish or worse, to an extent generated. Therefore, he's concerned that it puts the legitimacy of the research outcomes (e.g. completeness, contents of letters, classification, maybe even hallucinations in the thesis proper).

Preemptive arguments:

1. "The author's a researcher, not a programmer; therefore it's fine to use an LLM. It is preposterous to ask each researcher to learn web development to publish their research." You are right, but given the amount of vibe-coded websites we see, and them all having the default (Astro?) style, the grandparent all the same has the right to associate that style with untrustworthy crap. I'm not saying that this academic website is necessarily crap. However, I think it's useful for the grandparent to share their sentiment, because the researcher might not know.

2. "A lot of pages have links to sources; you could verify the legitimacy yourself". perhaps, but doubting the veracity of research is a bad first impression, isn't it?

It's a bit sad, because the website is non-trivial, and would have taken quite a bit of effort without an LLM. But it is difficult to separate webdev enablement with the rest of the LLM baggage.

Re: Roman Letters

#13

A few things about AI-led projects like this come to my mind — first, it’s cool to see all this pulled together. I’m sure the design will read “Claude 2026” soon, but that’s fine - it’s clean and generally has reasonable UX. There are some real rough spots - for instance, the Latin texts are generated via OCR from scanned documents directly; they’re not from some other scholarly corpus that’s been checked. I only loo…

[deleted]

Re: Roman Letters

#14
post #2

What a cool project, I like this one where Pliny the Younger complains about a no-show at his dinner party: https://romanletters.org/letters/pliny_younger/1015/

Had to look up "sow's matrices." > A "sow's matrix" (or vulva in Latin) is a dish from ancient Rome consisting of the uterus of a sow (a female pig), often specifically from one that has never farrowed or that was slaughtered shortly after farrowing. It was considered a delicacy among the wealthy elite and was a common dish served at lavish Roman banquets and dinner parties, often used as a sign of luxury, wealth, an…

[deleted]

Re: Roman Letters

#15

A few things about AI-led projects like this come to my mind — first, it’s cool to see all this pulled together. I’m sure the design will read “Claude 2026” soon, but that’s fine - it’s clean and generally has reasonable UX. There are some real rough spots - for instance, the Latin texts are generated via OCR from scanned documents directly; they’re not from some other scholarly corpus that’s been checked. I only loo…

I haven't checked any texts from the 500s. But I did some work with texts from the 1700s. Most of them had terrible transcriptions on archive.org, made using old tesseract versions. You could probably improve a lot with newer tesseract versions. I went for the nuclear option and just passed the image of each page (along with some context on how the previous page ended) to Qwen2.5vl:32b and got near-perfect transcriptions. And as you can tell by the old model that was months ago, vision models only got better.

Of course in some cases vision models are a liability for OCR because the errors they do make are replaced by plausible sounding replacements instead of alphabet soup. But if you only use the transcription as input for an LLM that doesn't matter. It only becomes an issue of how much compute you are willing to throw at it

Re: Roman Letters

#17

A few things about AI-led projects like this come to my mind — first, it’s cool to see all this pulled together. I’m sure the design will read “Claude 2026” soon, but that’s fine - it’s clean and generally has reasonable UX. There are some real rough spots - for instance, the Latin texts are generated via OCR from scanned documents directly; they’re not from some other scholarly corpus that’s been checked. I only loo…

I haven't checked any texts from the 500s. But I did some work with texts from the 1700s. Most of them had terrible transcriptions on archive.org, made using old tesseract versions. You could probably improve a lot with newer tesseract versions. I went for the nuclear option and just passed the image of each page (along with some context on how the previous page ended) to Qwen2.5vl:32b and got near-perfect transcript…

Yes, exactly. What could be durable is not the specific transcription as of today - until it’s perfect or at least ‘good enough’ - but the web site, comments, and process that can be run and turn into improved results - that part seems likely to be valuable to me.

Re: Roman Letters

#19

A few things about AI-led projects like this come to my mind — first, it’s cool to see all this pulled together. I’m sure the design will read “Claude 2026” soon, but that’s fine - it’s clean and generally has reasonable UX. There are some real rough spots - for instance, the Latin texts are generated via OCR from scanned documents directly; they’re not from some other scholarly corpus that’s been checked. I only loo…

Thank you so much for the feedback on this. I just implemented some updates that should better track our iterations and changes to the project. Obviously, GitHub does some tracking, but I've added some more formal updates.

Additionally, I completely agree on the OCR issues. Long-term, the goal is to use higher quality OCR and a broader set of data to make this even more accessible to people.

I'll note that the primary goal is as much scholarship as it is giving access to hobbyists or people interested in the data itself. If the data could get to the point where it is scholarly useful, of course that would be something I'd like to achieve.

Re: Roman Letters

#20
post #12
post #9

Earlier quoted context omitted.

The design is good. It is unoriginal but not every project needs to use an original design.

serious_angel is not contending with you that the design is bad, or that it is bad because it is unoriginal. In fact, they are not even specifically calling out the design. They have noticed the design, recognized it as the output of an LLM, then proceeded to discover that an LLM was involved in much of the creation of the project. This is an academic project. Whatever the pedigree of the researcher is, this implies…

Thanks for the feedback on this. I'll note that I am not an academic nor a researcher. I'm a hobbyist who wanted to bring together data from different letters and read them for myself. I don't aspire to be a researcher nor publish academic work.

The primary goal is to provide access to these letters to non-academics who would like to read what Romans were writing in language they can understand.

I take all the points about research outcomes and the quality of the data itself. That is going to be an ongoing process to continue to improve it alongside LLMs. I have a day job, and this is just a side project, but where it can provide value, I want to lean into that side of things.

Post reply on HN