I think that part of the beauty of LLMs is their versatility in so many different scenarios. When I build my agentic pipeline, I can plug in any of the major LLMs, add a prompt to it, and have it go off to do its job. Specialized, fine-tuned models sit somewhere in between LLMs and traditional procedural code. The fine-tuning process takes time and is a risk if it goes wrong. In the meantime, the LLMs by major provid…
Are they though? Or are they just getting better at gaming benchmarks?
Subjectively, there has been modest progress in the past year, but I'm curious to hear other anecdotes from people that aren't firmly invested in the hype.