Earlier quoted context omitted.
How good are these models at summarization anyways? I tried uploading obscure books I've already read, to GPT4 and Claude 3 and asked them to summarize the plot and particular details, as well as asking how many times does a particular thing happen in the book, and the results have been hit and miss. I certainly would not trust these models to create comprehensive and correct summaries of highly sensitive records.
Asking "how many times does a particular thing happen in the book" is always going to be hard, because LLMs are notoriously bad at counting.
his may of been a context/chunking issue (if that particular section doesn't name the character performing an action), but maybe its better now.