Live data from Hacker News

Extract-0: A specialized language model for document information extraction

arxiv.org

61–63 of 63 posts

Re: Extract-0: A specialized language model for document information extraction

#61

It's wild to me how many people still think that fine-tuning doesn't work. There was just a thread the other day with numerous people arguing that RL should only be done by the big labs. There is so much research that shows you can beat frontier models with very little investment. It's confusing that the industry at large hasn't caught up with that

the subject of this news likely doesn't generalise to random documents You need some serious resources to do this properly, think about granite docling model by IBM. For LLM: Finetuning makes sense for light style adjustments with large models (eg. customize a chat assistant to sound a certain way) or to teach some simple transformations (eg. a new output format). You get away with 100-1000 samples. If you want to te…

The pragmatic case is always prompt engineering. I’m speaking to the common idea that fine tuning doesn’t work, which if you need it, and have the capital (which isn’t all that much) it’s helpful
Post reply on HN