I compared AOCI + Codex with native Codex in a production repository of roughly 360,000 lines. I saw no efficiency gain in this test: elapsed time and uncached input tokens increased in all four scenarios. One task was blocked by the onboarding validation gate.

In the most comparable completed full-repository pair, both answers scored 6/6. AOCI took 2.93× as long and used 3.40× the uncached input tokens. That excludes nearly four hours of one-time indexing.

So far, I have not found a way for AOCI to pay off in my own work.

My principle: every added layer needs evidence of its benefit

I prefer giving the agent the problem and the necessary tools, then letting it choose how to use them.

Indexes, workflows and other scaffolding are worth trying. Keeping them depends on a concrete question: for the same task and comparable quality, is the result faster, cheaper or more reliable?

GitHub stars and recommendations help me discover tools. Comparisons on real tasks determine whether I use them. As models improve, I want clearer evidence that an added layer earns its place.

Test details

A code index organizes a repository in advance to help an agent find relevant material. Think of it as a book’s table of contents. Building and reading that table also costs something.

Both setups used gpt-6.1-sol / low. Each task started in a fresh session. Two tasks covered an eight-file slice; two covered the full repository.

Task Seconds: native → AOCI Uncached input tokens: native → AOCI Outcome
Call-chain trace, 8 files 69 → 105 36,479 → 44,068 Both answered
Bug diagnosis, 8 files 26 → 58 9,445 → 28,210 Both fixes passed 15 tests
Source-null trace, full repository 106 → 230 60,803 → 169,731 AOCI gate blocked; no business answer
Forecast to release, full repository 111 → 326 61,495 → 209,236 Both answers scored 6/6

AOCI versus native Codex benchmark

One-time full-repository indexing took about 3 hours 55 minutes and used 6.07 million uncached input tokens, including failed attempts and recovery. Task times exclude indexing. Uncached input tokens are not a direct billing estimate.

Possible explanations

  • A large repository can still contain small tasks. With a clear entry point, failing example or call chain, native Codex can locate the relevant code through search and a few source reads.
  • The index adds a fixed onboarding cost. Both full-repository tasks loaded all 22 index chunks and performed delivery confirmation and cognitive challenges. Precise answers still required source verification, potentially adding index reading to code reading.
  • Reuse remains untested. Fresh sessions repeated the onboarding cost. Several tasks in one session might produce different results.

These are possible explanations. I did not separately time index loading, validation and business reasoning, so I cannot attribute the slowdown precisely.

These were single runs, with only one completed full-repository pair. Long-session reuse and amortization remain untested. The results support my current usage decision, without establishing how AOCI performs across all repositories and tasks.

If you have measured a clear gain, I would like to know which task, how you integrated the tool, and where the benefit came from.