Live data from Hacker News

/architect: Reduce Fable tokens by 80%, Fable orchestrates/reviews, Codex builds

github.com

21–30 of 47 posts

Re: /architect: Reduce Fable tokens by 80%, Fable orchestrates/reviews, Codex builds

#21
Reducing token usage is this year's "one weird trick". It doesn't make sense on the face of it.

Even if one discovered something that millions (billions?) of dollars of AI compute and the best statisticians in the world was not able to find via exhaustive research, domain search and training... what do you think are the chances this won't be folded into the next update of every model, making the rigmarole moot?

Extraordinary claims require extraordinary evidence and technology-shattering innovations in AI are not know to come from a markdown.

Re: /architect: Reduce Fable tokens by 80%, Fable orchestrates/reviews, Codex builds

#22

DESIGN.md: > Each rule below is enforced mechanically by the skill, not left to vibes. > R1. Repo docs are the memory; not in HANDOFF.md = didn't happen SKILL.md: > Not in docs/HANDOFF.md = didn't happen. Refuse to judge results that exist only in conversation or builder chat output. "Mechnical enforcement" just means "prompting the LLM a bit extra" these days? It (still) amazes me how much effort and tokens we expen…

[dead]

Re: /architect: Reduce Fable tokens by 80%, Fable orchestrates/reviews, Codex builds

#24
post #21

Reducing token usage is this year's "one weird trick". It doesn't make sense on the face of it. Even if one discovered something that millions (billions?) of dollars of AI compute and the best statisticians in the world was not able to find via exhaustive research, domain search and training... what do you think are the chances this won't be folded into the next update of every model, making the rigmarole moot? Extra…

incentives aren’t aligned

Re: /architect: Reduce Fable tokens by 80%, Fable orchestrates/reviews, Codex builds

#25

DESIGN.md: > Each rule below is enforced mechanically by the skill, not left to vibes. > R1. Repo docs are the memory; not in HANDOFF.md = didn't happen SKILL.md: > Not in docs/HANDOFF.md = didn't happen. Refuse to judge results that exist only in conversation or builder chat output. "Mechnical enforcement" just means "prompting the LLM a bit extra" these days? It (still) amazes me how much effort and tokens we expen…

[dead]

Re: /architect: Reduce Fable tokens by 80%, Fable orchestrates/reviews, Codex builds

#26

DESIGN.md: > Each rule below is enforced mechanically by the skill, not left to vibes. > R1. Repo docs are the memory; not in HANDOFF.md = didn't happen SKILL.md: > Not in docs/HANDOFF.md = didn't happen. Refuse to judge results that exist only in conversation or builder chat output. "Mechnical enforcement" just means "prompting the LLM a bit extra" these days? It (still) amazes me how much effort and tokens we expen…

Agents are in a wacky state, which makes projects like this fall into a weird spot. Eg I vaguely expect my agent to do two disparate things: manage dependency injection for tools, prompt modifications, etc, but also be the sort of “brain trust” that controls the flow of execution (can we stop now, do we keep going, etc).

This project is meant to be the latter, but there’s not a clean way to integrate that into Claude Code or Codex because they expect to do both.

Pi can do it, but then your users can’t use their Claude subscriptions, so you have to cludgily try to do the same thing via LLM prompts.

Re: /architect: Reduce Fable tokens by 80%, Fable orchestrates/reviews, Codex builds

#28
post #18
post #15

Earlier quoted context omitted.

The problem is that there are a bunch of benchmarks, the model providers often don't even use the same benchmarks, a bunch of them have known problems, and it's expensive to do your own benchmarks. I am a GPT 5.x booster since to me it just feels smarter, and I generally felt like the benchmarks backed me up, but it's not every benchmark, so sadly we're mostly arguing about vibes. SWEBench-Pro was a big one, though a…

I find it fascinating that every time this kind of discussion comes up, people talk about night and day experiences between Claude and Codex, in both directions. I’m really wondering what people are doing to get such different outcomes. I’m currently working on two projects/clients one using Claude, one using Codex. I have a strong preference for the latter, but not because I think it is much more intelligent or writ…

I think I like Codex for the same reason tbh. I think it's just general misanthropy or autism or something lol. Most people seem to prefer Claude.

For me, I think Codex was visibly smarter than Claude until 4.8 came out, it would regularly do better debugging and IMO write better code. 4.8 I think is close.

I think Claude is widely regarded to have a big lead in front-end, which I do not work on.

Claude's Ultrathink is pretty cool, though it eats up tokens like nothing else obviously.

Post reply on HN