ousterhout-quality-program

by Rob ZappNo install yetNo like yetUpdated September 21, 2026Category: Code review

What it does

Use whenever code being written or reviewed creates or changes a boundary — a new module, class, component, helper, hook, service, or wrapper; any extraction or centralization of shared code; any "let's make this reusable" moment — and when explicitly reviewing, refactoring, or designing a module. Judges whether an abstraction earns its keep: module depth, whether to hide a design decision, whether duplicated code protects a shared invariant or merely rhymes, whether an interface is stable. Guards against mechanical SOLID/Clean Code that produces many shallow classes. Also defines the reader-cost test (code that is cheap for humans and agents to read and change) and the procedure for refactoring an existing codebase to this standard.

Install opens this entry in your AgentsRoom desktop app. If the app is not installed yet, you will be sent to the download page.

SKILL.md

---
name: ousterhout-quality-program
description: Use whenever code being written or reviewed creates or changes a boundary — a new module, class, component, helper, hook, service, or wrapper; any extraction or centralization of shared code; any "let's make this reusable" moment — and when explicitly reviewing, refactoring, or designing a module. Judges whether an abstraction earns its keep: module depth, whether to hide a design decision, whether duplicated code protects a shared invariant or merely rhymes, whether an interface is stable. Guards against mechanical SOLID/Clean Code that produces many shallow classes. Also defines the reader-cost test (code that is cheap for humans and agents to read and change) and the procedure for refactoring an existing codebase to this standard.
---

# Ousterhout Quality Program

## Overview

A module's job is to hide complexity behind a small interface. The core measure
is **depth**: a deep module offers a simple interface over substantial
functionality; a shallow module's interface is nearly as complex as its
implementation, so it pays for itself in nothing. Complexity is what you feel
when a change forces you to understand or touch code you did not expect to —
Ousterhout names two sources: **dependencies** (you can't change A without
changing B) and **obscurity** (the important information isn't obvious).

Ousterhout alone tells you what a good module *feels* like. It is strongest
combined with a few other lenses that tell you where the boundaries belong and
how to move toward them safely. This skill is that combined lens.

## Where Reviews Actually Go Wrong

The two failures this skill exists to correct — observed repeatedly in
agent-written code — are in the **remedy**, not the split/don't-split verdict
itself:

1. **The shallow fix.** Given six `as unknown as` casts, the unaided reviewer
   centralizes them into one generic `castRows<T>()` helper — tidier, but the
   obscurity survives. The deep fix is typed row→domain mappers with tests
   pinned first (applying Parnas: a cast is the smell of a missing boundary;
   Beck: prove the mapping before moving it). Neatening a smell is not
   removing it.
2. **The reflexive extraction.** Given the same update logic repeated in three
   sibling components, every unaided reviewer said "extract a shared helper" —
   the DRY reflex. This program's rule, extending Metz: wait for the
   invariant, not the third look-alike — centralize when the code protects a
   shared rule, not when it rhymes.

When you find yourself recommending a fix, run it through both: does it remove
the obscurity or just relocate it, and does the extraction protect an invariant
or just deduplicate a shape?

## Proportionality Gate

Skip the lens when a change adds no new exported/importable name, creates no
new module/class/component/helper/hook/service/wrapper, and centralizes
nothing. Pure renames, mechanical codemods, config/data edits, and
single-line fixes are exempt. When in doubt, run only the two core tests
(depth, invariant) and stop there.

## The Rule in Full

Every piece of generated or reviewed code passes the Ousterhout lens before the
task is called done — not only explicit design reviews — except changes below
the proportionality gate (no new boundary, no centralization: renames, codemods,
config edits). Two tests: (1) **Depth** — a new interface must hide
substantially more than it exposes; an interface as complex as what it wraps
pays for nothing. (2) **Invariant** — extract shared code only when it protects
a shared rule, never because three sites rhyme; and a fix must remove
obscurity, not relocate it (centralizing six casts into one helper is still six
casts). When a change creates or reshapes a boundary, first find how an
established product solves a problem of this shape and scale and adopt its
conventions unless there is a stated reason not to (a pattern recalled from
training is a claim, not a source), then run the checks below.

## When to Use

- Deciding whether a new class/function/hook is worth its interface, or is just a shallow pass-through.
- A file crosses a size threshold and you're deciding *how* to split it, not just that you should.
- Repeated code tempts you to extract a shared helper.
- Designing or reviewing a boundary around a business rule (an authorization scope check, a money/rounding rule, a state-machine transition guard, a data-retention rule).
- An interface is about to grow a parameter or a special case.
- Bringing an existing codebase to this standard — see "Refactoring an Existing Codebase to This Standard" below.

**Not for:** trivial mechanical edits, or when a project convention already dictates the structure — see the Proportionality Gate above. Defer to `karpathy-guidelines` for surgical-change discipline and a test-driven-development skill for the refactor safety net, when those are available.

## The Lenses

Each lens adds exactly one question. Ousterhout is the spine; the others correct its blind spots.

| Lens | The one question it adds | When it overrides |
|---|---|---|
| **Ousterhout** — deep modules | Does this interface hide more than it exposes? | Default spine. |
| **Parnas** — information hiding | What design decision (likely to change) does this module hide? | The *reason* a module should be deep. If it hides nothing that changes, depth is cosmetic. |
| **Brooks** — essential vs accidental | Does this remove accidental complexity, or just relocate essential domain complexity? | Kills "refactors" that move the mess without shrinking it. |
| **Evans** — Domain-Driven Design | Is this boundary named in domain language, not generic utility language? | Rename `utils`/`helpers` — name the boundary after the invariant this repo actually has. |
| **Fowler** — refactoring / smells | What is the smallest safe move toward the deeper design? | Turns "should be deeper" into concrete steps behind passing tests. |
| **Beck** — simple design, test-first | Have I proven current behavior before deepening the seam? | A brake on premature architecture. Make it work and tested first, then deepen the right seam. |
| **Hickey** — simple vs easy | Does this interleave unrelated concepts, or is it genuinely one concept? | A shallow helper is usually *easy* (nearby, quick), not *simple* (few interleaved concepts). Prefer simple. |
| **Metz** — duplication over wrong abstraction | Does this repeated code protect a shared invariant, or just look alike (this program's rule, extending Metz)? | Metz: duplication is cheaper than the wrong abstraction — inline a wrong abstraction back rather than bend it. This program extends her: do **not** centralize because it repeats; centralize only when it protects a real invariant. Tolerate duplication until the invariant reveals itself. |
| **Hyrum's Law** — observable behavior | Will callers depend on behavior beyond this interface's contract? | Argues for small, stable interfaces: every observable behavior eventually becomes load-bearing. |

## The Combining Recipe

Apply in this order — later lenses only matter once the earlier ones pass:

1. **Metz — the admission gate.** Does this boundary/abstraction earn existence at all? This program's rule, extending Metz: extract only when the code protects a shared rule — three look-alikes are not a revealed invariant. If no, stop here.
2. **Parnas / Ousterhout** — Hide the volatile decision (authorization scope, rounding rule, transition guard, retention rule) behind a deep module.
3. **Evans** — Name that module in domain language, not `utils`.
4. **Beck / Fowler** — For existing code, pin current behavior with tests, then refactor toward it in small safe moves. For freshly generated code there is no current behavior to pin — write the test that defines intended behavior instead.
5. **Hickey** — Reject interfaces that mix unrelated concepts just because the workflows look similar.

## The Structural Anti-Pattern

**Mechanical SOLID / Clean Code produces shallow modules.** A dogmatic reading —
one class per responsibility, extract every function, keep everything tiny —
yields a swarm of classes whose interfaces are as complex as their bodies. When
a rule says "split this," ask what *decision* the split hides (Parnas) and
whether it hides more than it exposes (Ousterhout). If it hides nothing that
changes, don't split. This guard matters most under refactor pressure ("clean
this up", "this file is too big") — in calm analysis, reviewers already resist
it; mid-refactor, with a mandate to produce visible change, is when the swarm
of shallow files gets written.

## Common Mistakes

- **Splitting on size alone.** A 400-line query module that hides one coherent decision may be deeper than four 100-line modules that each leak the same joins.
- **Naming the split `helpers`/`utils`.** If you can't name it in domain language (Evans), the boundary is probably wrong.
- **Extracting on the second occurrence.** This program's rule, extending Metz: wait for the invariant, not the third look-alike.
- **Deepening before pinning behavior.** Beck: without a test proving current behavior, a "deepening" refactor is a rewrite.
- **Counting a pass-through as a module.** A wrapper that forwards its arguments adds an interface and hides nothing — shallow by definition.
- **Mistaking a rhyme for an invariant.** The best evidence of a shared invariant is co-change: the copies have been fixed or changed together in history (the same bug fixed in two places). Look-alikes that change independently are rhymes; leave them duplicated.
- **Neatening a smell instead of removing it.** Centralizing six casts into one generic cast helper is the tidy version of the same obscurity. The deep fix names the boundary the cast was papering over.

## Reader Cost: the Third Test

Depth and invariant decide whether a boundary should exist. Reader cost
decides whether the code around it is cheap to change. The next reader, human
or agent, pays for every line they must load to change something safely.
Agents pay in tokens and navigate by text search, partial reads and
typecheck/test loops, so the same defects cost them more. Ask:

- **Findable?** One name per concept, spelled the same everywhere, reachable by plain text search. Defects: names assembled from strings, wiring by import side effect, re-export chains that hide the definition, two names for one concept.
- **Can the reader stop early?** The contract sits at the top of the file or above the export: what it promises, what it hides, what it never does. Defect: the contract can only be derived by reading the body.
- **Machine-checkable?** Precise types in and out of every boundary, so a typecheck replaces reading callers. Defects: `any`, bare dictionaries, boolean flags whose meaning lives in the body.
- **Is coupling visible?** Places that must change together are enforced (a shared type, a test, a single source) or, failing that, marked at both sites. The evidence of hidden coupling is co-change in history that nothing in the code mentions.
- **Noise-free?** No comments that restate the code, no commented-out code, no dead branches, no change-history comments, no obsolete path kept beside its replacement.
- **Predictable?** Layout follows the repo's existing pattern; the test is where a reader will look for it and runs on its own.

File size is deliberately absent. A very large file is a reason to look for a
second hidden decision, never a reason to cut: readers can search and read a
range, and a split that hides nothing adds interfaces without removing load.

For in-code markers and a repo codemap, use `context-audit` where available:
its `AIDEV-NOTE:` anchor (one non-recoverable fact plus a provenance ref, at
most two lines, at the site) is the convention for coupling that cannot be
enforced.

## Refactoring an Existing Codebase to This Standard

A retrofit is judged the same way as new code; what differs is order and
restraint. Most of a codebase should be left alone.

1. **Census, read-only.** List the boundaries (modules, services, shared helpers). For each record: the decision it hides, or "none"; interface size against body; co-change partners from history; reader-cost defects. Change nothing yet.
2. **Rank by churn, not by ugliness.** Priority is how often the code changes times what it costs to read. Cold code that works stays as it is, however shallow. Essential domain complexity stays where it is (Brooks).
3. **Assign one remedy per finding:**
   - pass-through layer or wrapper that hides nothing: delete it, callers use what it wrapped;
   - wrong abstraction bent by flags and special cases: inline it back (Metz), then look for the real invariant;
   - shallow siblings that share one decision: merge them behind one interface;
   - leaked decision (callers know the format, the rule, the schema): pull it down into the module that owns it;
   - generic name (`utils`, `helpers`, `manager`): rename for the decision it hides, or dissolve it into its callers;
   - untyped boundary: type it, and replace casts with the mapper they were papering over;
   - hidden coupling: enforce it, or mark both sites;
   - noise: delete it.

   Rhymes that change independently get no remedy.
4. **Pin behavior first.** No remedy starts until a test proves the current behavior of the code it touches (Beck). Refactors are behavior-preserving; a behavior change is a separate commit.
5. **Cut the work into units one agent can finish alone.** One boundary per unit. Each unit names the files it owns, the contract it must preserve, and the command that proves it on its own. No two concurrent units write the same file; shared files (barrels, registries, route tables) get a single owner or wait for integration. Interface changes that several units depend on land first, as their own unit.
6. **Measure the result.** Pick a representative change before starting and count the files and lines a reader must load to make it; count again after. Exported names and total lines should fall or hold. A refactor that adds interfaces owes a stated reason.
7. **Stop** when what remains is cold, essential, or a rhyme.

Related skills, where available: `repo-review` (design type) produces the
census as an advise-only artifact; `design-cleanup` runs the fix-and-rescan
loop for accidental complexity; `context-audit` adds anchors and the codemap;
`ousterhout-build-deep` is the author-time checklist for the agents doing the
units.

## Where this sits

This skill is the review and judgment layer: use it to decide whether an abstraction is deep, named for the right decision, and worth extracting. `find-shared-code` uses it as its admission test when sweeping recent history for code worth sharing. The Appendix below gives each author's reasoning.

---

## Appendix: The Lenses in Depth

The failure mode each author catches, and the one move each one gives you. The
table above is the quick reference; this is the reasoning behind it.

### Ousterhout — Deep Modules (the spine)

*A Philosophy of Software Design.*

- **Depth** = benefit (functionality hidden) ÷ cost (interface complexity). A deep
  module hides a lot behind a little. A shallow module's interface is nearly as
  complex as its body, so it earns nothing.
- **Complexity** is anything about the system that makes it hard to understand or
  modify. Two sources:
  - **Dependencies** — you can't change one piece without touching another.
  - **Obscurity** — the important information isn't obvious from the code.
- **Symptoms:** change amplification (one decision, many edits), cognitive load
  (how much you must hold in your head), unknown unknowns (you can't tell which
  code a change will affect).
- **Key move:** pull complexity *downward* — the module absorbs the hard case so
  callers don't have to. Configuration parameters and pass-throughs push
  complexity *up* to the caller; that's shallowness.

Catches: interfaces that leak their implementation; helpers that don't help.

### Parnas — Information Hiding (why depth matters)

*On the Criteria To Be Used in Decomposing Systems into Modules (1972).*

- Decompose around **design decisions likely to change**, not around the steps of
  a computation. Each module hides one such decision.
- This is the direct ancestor of the deep module. A module is deep *because* it
  hides a decision that would otherwise ripple through callers.

Catches: a "module" that hides nothing volatile — its depth is cosmetic. Ask:
what changes behind this interface that callers never see? If the answer is
"nothing," the boundary is decoration.

### Brooks — Essential vs Accidental Complexity

*No Silver Bullet.*

- **Essential** complexity is inherent in the domain (appraisal really is this
  intricate). **Accidental** complexity is what our tools and structure impose.
- Only accidental complexity is removable. A refactor that "cleans up" by moving
  essential domain complexity from one file to another has done nothing.

Catches: reshuffles disguised as simplification. Ask: did total complexity drop,
or did it just move?

### Evans — Domain-Driven Design

*Domain-Driven Design.*

- Boundaries should be named in the **ubiquitous language** of the domain, not in
  generic utility terms. A module called `helpers` names nothing; a module called
  `AccessScope` or `PricingPolicy` names an invariant.
- Bounded contexts keep business invariants from leaking across seams.

Catches: correct decomposition with meaningless names. If you can't name the
module in domain language, you probably cut the boundary in the wrong place.

### Fowler — Refactoring and Code Smells

*Refactoring.*

- Gives the concrete, safe, named moves (Extract Function, Move Field, Replace
  Conditional with Polymorphism) to get from the current design to the deeper one.
- Every move is behavior-preserving and small, so it stays reversible.

Catches: the gap between "this should be deeper" and knowing the next commit.
Ousterhout sets the target; Fowler is the road.

### Beck — Simple Design, Test-First

*Test-Driven Development; XP.*

- Four rules of simple design, in Beck's published order: passes tests, no
  duplication, reveals intention, fewest elements. This program follows the
  later Fowler/Haines reordering — intent before duplication — because it
  serves this program's Metz-extending invariant rule (see Metz, below):
  don't act on duplication until you can name the intention it protects.
- Test-first is a brake against premature architecture. Make it work and prove
  behavior *first*, then deepen the seam that the tests now protect.

Catches: architecture built before behavior is pinned. Without a test proving the
current behavior, a "deepening" refactor is an unverified rewrite.

### Hickey — Simple vs Easy

*Simple Made Easy.*

- **Simple** = un-braided: one concept, not interleaved with others (objective).
- **Easy** = near at hand, familiar, quick to reach for (relative to you).
- The two are independent. A shallow helper is usually *easy* — fast to write,
  close by — but not *simple* if it braids unrelated concerns.

Catches: convenience masquerading as design. Prefer constructs that keep concepts
un-braided even when a braided one is quicker to type.

### Metz — Prefer Duplication Over the Wrong Abstraction

*"The Wrong Abstraction" (2016).*

- Duplication is far cheaper than the wrong abstraction. An abstraction extracted
  too early forces every future caller to bend around assumptions that were never
  true for all of them.
- When an abstraction turns out wrong, Metz's remedy is to inline it back and let
  the duplication return, rather than bend it to fit a case it was never built for.
- **This program's rule, extending Metz: do not centralize because code repeats.
  Centralize when it protects a real, shared invariant.** Until the invariant
  reveals itself, tolerate the duplication.

Catches: over-centralization — the shallow shared helper that everyone must now
work around. This is the counterweight to a mechanical "DRY at all costs."

### Hyrum's Law — Observable Behavior Becomes Contract

*"With a sufficient number of users, every observable behavior of your system
will be depended on by somebody."*

- Whatever an interface *happens* to do — ordering, timing, error text — someone
  will eventually rely on. So the surface you expose is larger than the surface
  you documented.
- This supports Ousterhout's preference for **small, stable interfaces**: the less
  you expose, the less can become load-bearing by accident.

Catches: wide interfaces that will ossify. Every extra observable becomes a
future constraint.

### How They Fit Together

- **Parnas → Ousterhout:** hide a volatile decision → the module is deep.
- **Brooks:** confirm the depth removed complexity rather than relocating it.
- **Evans:** name the boundary in domain language.
- **Beck → Fowler:** pin behavior, then refactor in small safe moves.
- **Metz:** resist centralizing until the invariant is real.
- **Hickey:** keep the interface to one concept.
- **Hyrum:** keep that interface small so it can stay stable.

The danger is mixing Ousterhout with a mechanical reading of SOLID or Clean Code:
that produces many tiny classes and functions with shallow interfaces — the exact
opposite of deep modules. Ousterhout, with Metz as counterweight, is the antidote.

Tags

designarchitecturereviewrefactoringousterhout