Earlier quoted context omitted.
Did you benchmark the competition and can we see?
The problem even attempting to develop a tool for the frontier model space is that the cost to run a statistically significant benchmark is almost certainly going to be over $100 - for a single model. Unless something is like 25%+ more cost effective on Gemini for a task, I would not assume those savings are going to transfer to GPT. If you need to run a test this expensive and slow for every release, hobbiests aren'…
On the previous large benchmark run, i proved 40-50% cost reduction per correct answer.
I'm not sure why the vendors aren't using token filtering/compression more in their tooling, but perhaps they don't mind users feeding them more data and using more data.