AccountingBench: Evaluating LLMs on real long-horizon business tasks
11–20 of 154 posts
Re: AccountingBench: Evaluating LLMs on real long-horizon business tasks
#12Re: AccountingBench: Evaluating LLMs on real long-horizon business tasks
#13* https://en.wikipedia.org/wiki/Financial_Modeling_World_Cup
* https://www.cbc.ca/radio/asithappens/2024-excel-world-champi...
Can't wait for this to start having 'e-sports' tournaments. :)
Re: AccountingBench: Evaluating LLMs on real long-horizon business tasks
#14Re: AccountingBench: Evaluating LLMs on real long-horizon business tasks
#15Re: AccountingBench: Evaluating LLMs on real long-horizon business tasks
#16Yes, LLMs have and will continue to improve. But it's that initial "holy shit, this thing is basically as good as a real accountant" without any understanding that it can't sustain it which leaves many with an overinflated view of their current value.
Re: AccountingBench: Evaluating LLMs on real long-horizon business tasks
#17We've been on this train of not caring about the details for so long but AI just amps it up. Non-deterministic software working on things that have extremely precise requirements is going to have a bad outcome A company may be OK with an AI chatbot being so bad it results in 5-20% of customers getting pissed off and not having a 5-star experience. The SEC and DOJ (and shareholders) are not going to be happy when the…
Re: AccountingBench: Evaluating LLMs on real long-horizon business tasks
#18I don't think you'll find many sane CFOs willing to send the resulting numbers to the IRS based on that. That's just asking to get nailed for tax fraud.
It is coming for the very bottom end of bookkeeping work quite soon though, especially for first draft. There are a lot of people doing stuff like expense classification. And if you give an LLM an invoice it can likely figure out whether it's stationary or rent with high accuracy. OCR and text classification is easier for LLMs than numbers. Things like concur can basically do this already.
Re: AccountingBench: Evaluating LLMs on real long-horizon business tasks
#19I guess having access to tools / running Python would make all the difference.