I ran my custom agentic SQL debugging benchmark against it and I'm impressed. Results: 8 passed, 0 failed, 17 errored out of 25 That puts it right between Qwen3.5-4B (7/25) and Nanbeige4.1-3B (9/25) for example, but it took only 200 seconds for the whole test. Qwen3.5 took 976 seconds and Nanbeige over 2000 (although both of these were on my 1070 so not quite the same hardware) Granite 7B 4bit does the test in 199 se…
I have been using @freakynit's runpod as well all be it, I like making working pomodoro apps as my own custom test, and although its not good for it (none of the prototypes work), I feel like it can be good within a specific context like Sql as you mention. I imagine this being used as sub-agents with some sota models directing them but I wasn't really able to replicate it personally (I had asked Claude to create a d…
Yes I think very constrained task: known data universe, well known language etc should be the best possible place for small language models to play