Live data from Hacker News

Anthropic's original take home assignment open sourced

github.com

71–80 of 394 posts

Re: Anthropic's original take home assignment open sourced

#72

I consider myself rather smart and good at what I do. It's nice to have a look at problems like these once in a while, to remind myself of how little I know, and how much closer I am to the average than to the top.

disagree. nobody has a monopoly on what metric makes someone good. I don't understand all this leet code optimization. actually i do understand it, but it's a game that will attract game optimizers.

the hot take is, there are other games.

Re: Anthropic's original take home assignment open sourced

#73

I suspect this was released by Anthropic as a DDOS attack on other AI companies. I prompted 'how do we solve this challenge?' into gemini cli in a cloned repo and it's been running non-stop for 20 minutes :)

Which Gemini model did you use? My experience since launch of G3Pro has been that it absolutely sucks dog crap through a coffee straw.

> sucks dog crap through a coffee straw.

That would be impressive.

Re: Anthropic's original take home assignment open sourced

#75

Naively tested a set of agents on this task. Each ran the same spec headlessly in their native harness (one shot). Results: Agent Cycles Time ───────────────────────────────────────────── gpt-5-2 2,124 16m claude-opus-4-5-20251101 4,973 1h 2m gpt-5-1-codex-max-xhigh 5,402 34m gpt-5-codex 5,486 7m gpt-5-1-codex 12,453 8m gpt-5-2-codex 12,905 6m gpt-5-1-codex-mini 17,480 7m claude-sonnet-4-5-20250929 21,054 10m claude-…

Very interesting thanks! I wonder what would happen if you kept running Gemini in a loop for a while. Considering how much faster it ended it seems like there is a lot more potential.

Re: Anthropic's original take home assignment open sourced

#76
post #73

Earlier quoted context omitted.

Which Gemini model did you use? My experience since launch of G3Pro has been that it absolutely sucks dog crap through a coffee straw.

> sucks dog crap through a coffee straw. That would be impressive.

New LLM benchmark incoming? I bet once it's done, people will still say it's not AGI.

Re: Anthropic's original take home assignment open sourced

#77

Naively tested a set of agents on this task. Each ran the same spec headlessly in their native harness (one shot). Results: Agent Cycles Time ───────────────────────────────────────────── gpt-5-2 2,124 16m claude-opus-4-5-20251101 4,973 1h 2m gpt-5-1-codex-max-xhigh 5,402 34m gpt-5-codex 5,486 7m gpt-5-1-codex 12,453 8m gpt-5-2-codex 12,905 6m gpt-5-1-codex-mini 17,480 7m claude-sonnet-4-5-20250929 21,054 10m claude-…

I do wonder how Grok would compare, specifically their Claude Code Fast model.

Re: Anthropic's original take home assignment open sourced

#78

Naively tested a set of agents on this task. Each ran the same spec headlessly in their native harness (one shot). Results: Agent Cycles Time ───────────────────────────────────────────── gpt-5-2 2,124 16m claude-opus-4-5-20251101 4,973 1h 2m gpt-5-1-codex-max-xhigh 5,402 34m gpt-5-codex 5,486 7m gpt-5-1-codex 12,453 8m gpt-5-2-codex 12,905 6m gpt-5-1-codex-mini 17,480 7m claude-sonnet-4-5-20250929 21,054 10m claude-…

Could you make a repo with solutions given by each model inside a dir/branch for comparison?

Re: Anthropic's original take home assignment open sourced

#79
post #44

Earlier quoted context omitted.

They wrote: > If you optimize below 1487 cycles, beating Claude Opus 4.5's best performance at launch, email us at performance-recruiting@anthropic.com with your code (and ideally a resume) so we can be appropriately impressed and perhaps discuss interviewing. That doesn’t seem snarky to me. They said if you beat Opus, not their best solution. Removing “perhaps” (i.e. MAYBE) would be worse since that assumes everyone…

That paraphrases to "do better than we have publicly admitted most of humanity can do, and we may deign to interview you" It sounds incredibly condescending, if not snarky, but I would classify those adjectives as mostly synonymous.

I understand how it can be interpreted as snarky, but how could it have been written better? It's a hard path to walk and recruiting/interviewing is inherently sensitive it seems.

Re: Anthropic's original take home assignment open sourced

#80
post #73

Earlier quoted context omitted.

> sucks dog crap through a coffee straw. That would be impressive.

New LLM benchmark incoming? I bet once it's done, people will still say it's not AGI.

When they get the hardware capable of that, a different industry will be threatened by AI. The oldest industry.
Post reply on HN