Live data from Hacker News

Show HN: Oqoqo – build evals and custom benchmarks for real-world tasks

oqoqo.ai

1–4 of 4 posts

Show HN: Oqoqo – build evals and custom benchmarks for real-world tasks

#1
Most benchmarks today exist in curated environments and do not translate well to the real world.

We built Oqoqo to bridge this gap. Oqoqo makes it super simple to build realistic evals and custom benchmarks for tasks users actually care about.

Oqoqo can:

- Reliably measure how agent friendly your product surfaces are against Codex, Claude Code, OpenClaw, Hermes, Pi, Opencode, Cursor, GitHub Copilot - Regression test MCP, CLI, skills, SDK, and any agent facing interface (we are continuously using Oqoqo to dogfood and improve our own MCP/CLI) - Create and share custom benchmarks for how agents discover and use your product - Compare models and harnesses for domain specific tasks - See whether new versions improve agent experience

Would love to hear feedback and thoughts on what kind of evals you are running today, if you are benchmarking agent facing interfaces, and whether you have published a custom benchmark.

Show HN: Oqoqo – build evals and custom benchmarks for real-world tasks
oqoqo.ai

Re: Show HN: Oqoqo – build evals and custom benchmarks for real-world tasks

#2
This will become extremely relevant. Sterile evaluation environments are a waste of time. If agents are the new software interface, we need a way to evaluate them as they actually work. Building a realistic eval harness is too complicated, especially if you want to make it reproducible at scale without cross contamination

Re: Show HN: Oqoqo – build evals and custom benchmarks for real-world tasks

#4

Real-world evals are more useful than benchmark scores. How do you handle false positives when a workflow passes but produces a plausible wrong result?

users have complete control of what a good output looks like. we have seen that adding criteria that is specific to what the expected output is, including at trajectory level, helps a lot to avoid false positives