Earlier quoted context omitted.
They are limited on how much they can output and there is generally an inverse relationship between the amount of tokens you send vs quality after the first 20-30 thousand tokens.
Are there papers on this effect? That quality of responses diminishes with very large inputs I mean. I observed the same.
I've experienced this problem but I haven't come across papers about it. For this context, it would be interesting to compare the accuracy of transcribing one page at a time to batches of n pages.