Recently, I ran into a noticeable drop in agent performance.
A development task that should have taken 20 to 30 minutes ended up taking almost three hours. The agent was not stuck on one especially difficult problem. It kept burning time on reading, deciding, validating, and then validating again.
So I started taking apart the environment it had been asked to work in.
The environment got heavy before the agent got slow
I found three kinds of problems.
The first was repository bloat. The repo contained quite a few source files longer than 3,000 lines, along with a temporary evidence file of roughly 110,000 lines. The .git directory had grown to 72 MB.
Those numbers alone do not prove how much delay each item caused. The real warning sign was the growing amount of noise the agent had to cross whenever it searched, reconstructed dependencies, or worked out what a file was responsible for. When a temporary artifact looks like an authoritative source, the agent also has to spend time deciding whether it can be trusted.
The second problem was the Skills I had previously asked AI to generate. They had become longer and longer, packed with step-by-step instructions and exceptions accumulated from historical incidents. Each addition made sense in isolation. Together, they had started to micromanage the agent.
The third was safety machinery. To reduce the chance of mistakes, I had added more and more gates to development and CI/CD. The validation chain grew longer, with more places to fail. The agent often ended up fixing secondary problems created by the validation machinery instead of solving the original problem.
All three shared the same pattern: whenever we encountered a risk, we added a little more. Code, documentation, and gates each grew incrementally until all of them sat directly in the agent’s working path.
I started by subtracting
For the repository, I began by inspecting the distribution of lines of code, identifying unusually large files, and deciding case by case whether each one should be refactored, archived, deleted, or retained.
A large file is not automatically a bad file. First I check its actual consumers, authoritative source, and recovery path. Temporary evidence can be deleted. Historical facts may need to be archived. A large module only becomes a refactoring candidate when it carries clearly separable responsibilities. For code over 1,000 lines, I look for responsibilities that can be judged and tested independently.
I also began reviewing abstractions and mechanisms that had been added to defend against possible future risks. I recently installed Ponytail to see whether it can help me identify overengineering more consistently. A tool can only support the judgment. The real question remains: if I remove this layer, which concrete risk comes back?
The bigger change happened in Skills and AGENTS.md.
From how-to toward what, where, and why
I used to tell the agent:
- which command to run first;
- which file to read next;
- which branch to take when it encountered a particular state;
- which sequence to follow for final validation.
That was usually correct when I wrote it. But how-to instructions are highly sensitive to external state: tools get upgraded, directories move, branches change, and runtime state shifts. Yesterday’s correct path can easily become today’s detour.
Now I want a Skill to answer four kinds of questions:
- What is there: What entities exist in this domain, and what is each one responsible for?
- Where are they: Where do authoritative facts, runtime state, and outputs live?
- What you might miss: Which boundaries, exceptions, and failure states are easy to overlook?
- Why: Why do these constraints exist, and what do they protect?
A session-level Goal then defines the outcome for this task, its acceptance criteria, and the limits of the agent’s authority. Once the agent has those relatively stable facts, it can work out the how from the current environment.
Goal
What must be achieved in this task, and what counts as done
Skill / Context
What exists -> where it is -> what is easy to miss -> why it matters
Agent
Read live state -> decide how to proceed this time -> validate the outcome
This does not mean eliminating operational guidance altogether. Dangerous, irreversible, or consistency-critical actions still need explicit rules. The difference is that those rules describe invariants and decision boundaries. A runbook appears only when there really is one correct path.
This is not just my impression
I later found that the industry has been describing a similar shift under different names.
Anthropic calls it Context Engineering: context is a finite attention budget, and the aim is to find the smallest set of high-signal information that reliably produces the desired behavior. Hard-coding too much agent behavior creates brittle prompts that are difficult to maintain. The better level of abstraction is specific enough to guide the agent while preserving room for judgment.
Anthropic also used a minimal scaffold for its SWE-bench agent: give the model the task, the repository, and a small set of basic tools, then let it decide how to proceed instead of encoding the workflow as a strict state machine.
OpenAI recently described the repository-side practice as Harness Engineering. One principle captures what I am trying to do: enforce boundaries centrally while preserving autonomy within them. Humans own goals, priorities, and acceptance. The repository makes architecture, tools, and facts discoverable. The agent owns the concrete execution path.
So “write less how-to” is only half the answer. Effective subtraction requires two things at once:
- remove process controls that become stale easily;
- make environmental facts, responsibility boundaries, and acceptance conditions easier to discover.
Otherwise, a short Skill is simply an underspecified one.
Early results
I removed the how-to sections from a batch of Skills and kept their responsibilities, locations, easy-to-miss details, and rationale. The Skill bodies became more than 50% shorter.
I do not yet have rigorous enough data to claim a specific productivity improvement. My subjective experience is that agents spend less time mechanically following outdated steps and complete tasks faster. The next step is to record task duration, wasted tool calls, repeated validation, and human interventions so that this impression can become a credible conclusion.
The unresolved problem: safety gates
Code and Skills can gradually be slimmed down through deletion, archiving, and refactoring. Safety gates are harder. Their waiting cost is visible; the incidents they prevent usually never happen and are therefore much harder to measure.
I am changing how I review them. For each gate, I no longer ask only, “Does this make the system safer?” I ask:
- Which identified risk does it protect against?
- Does it constrain outcomes and invariants, or prescribe the implementation process?
- How many real defects has it caught, and how many false positives and delays has it created?
- Can it run in parallel, activate only for relevant risks, or move closer to the final boundary?
- If it is removed, can existing tests, permissions, or rollback mechanisms cover the same risk?
OpenAI’s practical guide to building agents recommends adding guardrails incrementally in response to risks found in practice, while optimizing both safety and user experience. To me, that means safety mechanisms also need evidence and a lifecycle. Calling something “safety” should not exempt it from review forever.
My current view of working with agents is:
Give the agent a goal, reality, and boundaries. Let it choose the path for this particular reality.
As agents become more capable, the most valuable human work may be writing fewer detailed steps and making the facts that matter easier to see—and the boundaries that matter harder to cross.