Earlier quoted context omitted.
Well you can train AIs to uncover image fraud and detect manipulation that is well known (imagetwin does this). Presumably the same can be done for tabular data and genomic or "omics" data. At the very least statistical techniques can be used. I imagine high throughput imaging modalities will be the main target. That being said the only way the -omics data will have utility is via AI models which are trained on some…
The OP paper is easily handled by GOFAI from the 1980s; it's just detecting similar images. You don't need a language model. The effort is in collecting and chopping the data to find snippets to compare. NNs can help with that.
Newer fraud will probably use generative diffusion AIs to make "realistic western plots" on demand .. there's probably a paper in that!