Hi HN! We worked at OpenAI and Anthropic and believe we can provide much higher quality code generation by fine-tuning an LLM on your codebase compared to Sonnet-3.5 or o1 but not fine-tuned. Let me know if you are interested and we can fine-tune for you for free to test.
Talk is cheap, benchmarks please. Also why did you decide for LLama? AFAIK deepseek always had a slight edge over llama when it comes to coding performance, or is this no longer the case?
We fine-tuned Llama and got 4.2x Sonnet 3.5 accuracy for code generation
71–79 of 79 posts
Re: We fine-tuned Llama and got 4.2x Sonnet 3.5 accuracy for code generation
#72How do you measure code generation accuracy? Are there some base tests and if so how can I ensure the models aren't tuned for those tests only the same way vw cheated the emissions tests on their diesels?
Re: We fine-tuned Llama and got 4.2x Sonnet 3.5 accuracy for code generation
#73Is the source code available for inspection somewhere? It's not really clear from the landing page.
Re: We fine-tuned Llama and got 4.2x Sonnet 3.5 accuracy for code generation
#744.2x doesn't mean anything if you don't tell me what "accuracy" Sonnet 3.5 had.
Re: We fine-tuned Llama and got 4.2x Sonnet 3.5 accuracy for code generation
#75You make a bold marketing claim, 4.2x Sonnet, but viewing your website, I can see no data or test results to back this up.
Thanks for calling this out, even if it just gets OP to comment with some details/data. Was hoping this would be a shallow or deep dive into the results, but looks like it’s just a marketing post to a marketing page to support a PH launch.
Re: We fine-tuned Llama and got 4.2x Sonnet 3.5 accuracy for code generation
#76Was the comparison done with or without code context (as obtained using RAG or letting Sonnet ask for files)?
In the absense of other information, looks like a cherry-picked example to me.
Re: We fine-tuned Llama and got 4.2x Sonnet 3.5 accuracy for code generation
#77Maybe ask this model to create a better landing page?
Re: We fine-tuned Llama and got 4.2x Sonnet 3.5 accuracy for code generation
#78I like it and it makes sense, but from a business perspective I wonder what keeps the upstream LLM providers (all trying to generate profits) from offering the same fine-tuning service quickly ? Edit: OK, right it's olama, so I assume you can download your own model. (Assuming it's downloadable?) I think openAI already offers fine-tuning with custom data for some of their models, but maybe not specific to coding task…
Re: We fine-tuned Llama and got 4.2x Sonnet 3.5 accuracy for code generation
#792023: Our tiny model blah blah blah beats GPT4! 2024: Our tiny model blah blah blah beats Claude! 2025: Our tiny model blah blah blah beats ???