Two Models at Every Decision

Propel consults a second model automatically — at every gate, with no command to invoke and no way to forget. Codex advises. It never decides, and it never writes.

Why a Second Model

A model reviewing its own plan is a model grading its own homework. Claude's design proposal and Claude's review of that proposal come from the same training distribution, the same priors, and the same blind spots. The review will be thorough and it will miss exactly what the proposal missed.

A model trained on a different distribution disagrees in different places. Those disagreements are the signal — they mark the spots where a human should be spending attention. So Propel routes every major decision through OpenAI Codex before it reaches you.

A second opinion you have to request is one you get only when you already suspect you need it — which is the opposite of when it helps. So it isn't requested. It's scheduled.

There is no /c-codex, no flag to arm, nothing to remember. Earlier versions of Propel had a manual bridge command; it was the wrong design and it is gone.

Five steps at every gate: Claude drafts, Codex critiques, Claude verifies every file and line against the repository, an attributed card is produced, and you decide.
The five steps that run at every gate.

Where It Fires

Decision pointThe question Codex is asked
Gate 0 — scope statement"Is this scope self-contradictory, under-specified, or hiding a decision the human hasn't been asked about?"
Gate 1 — investigation findings"What should this investigation have checked and didn't? Name one thing."
Gate 2 — design proposal"Propose the single strongest alternative design, and name one failure mode this proposal doesn't address."
Gate 3 — component audit"Name bugs by file:line, or say 'no additional findings'."
Gate 4 — diagnosis"What else explains this evidence equally well?"
Bug classification"Argue against this classification."
3-strike break"Three attempts failed. What assumption do all three share?"
Retrospective"Which of these conclusions does the evidence not actually support?"

Notice that none of these questions can be answered with agreement. "Does this look right?" gets you a yes; "name the strongest counter-approach" gets you a second opinion. Every scheduled question is phrased so that a useful answer is a disagreement.

It does not fire on trivia. Not on Q0/Q1 (those are questions for you), not on mode selection, not on formatting or a three-line edit, not on anything confined to scratch/. Over-firing turns the ◆ Consulting Codex line into noise you scroll past — at which point the second model has stopped working even though it is still running and still costing you time.

Verification Is the Load-Bearing Step

Codex is a genuinely useful critic and a prolific fabricator of file paths, line numbers, and function names. Those two facts are not in tension — together they are the entire design constraint.

Every consult runs through the codex-bridge subagent, in its own context window. It runs the CLI in a read-only sandbox, extracts every checkable claim, opens the actual file for each one, and returns a three-way table:

BucketMeaning
VERIFIEDThe claim was confirmed against the repo, with the line quoted. A line number that was off by a few but right about the content counts as verified, with the location corrected.
DISPUTEDThe repo contradicts the claim. Reported as contradicted, not softened into uncertainty.
UNVERIFIEDThe claim names nothing checkable, or the target doesn't exist. Never promoted to a finding — and never silently dropped either, because a claim nobody could check is information about where the code is hard to reason about.

The raw transcript is discarded. This is deliberate: if Codex's unverified prose lands in the main conversation, every later turn reasons next to confident claims nobody checked. The subagent spends its own context on verification so your session doesn't have to.

Attribution Is Not Optional

Whenever Codex is used, Claude says so before the result, never after and never implicitly:

◆ Consulting Codex — Gate 2 (design): "Propose the strongest alternative
  and name the failure mode this design doesn't address."

And every finding in the resulting card carries its source:

TagMeaning
[claude]Claude's own finding
[codex]Codex raised it; Claude verified it against the repo
[both]Both models raised it independently
[codex · unverified]Codex raised it; Claude could not confirm it

A blended voice would destroy the entire point of having two models. If you cannot tell who said what, you cannot calibrate how much to trust it.

Agreement is weak evidence. Two models trained on overlapping internet text agreeing means very little, and Propel reports it that way rather than dressing consensus up as confirmation. When the two models fully agree, the card says so plainly and notes that the agreement is cheap.

How You Know It Actually Ran

The ◆ Consulting Codex line and the [codex] tags are good for reading a conversation, and they are not evidence. A model writes them. A model that forgets to announce — or that writes "Codex agrees" without calling anything — produces a transcript indistinguishable from the honest one. Forbidding fabrication is not the same as preventing it.

So there is a ledger. scripts/codex-consult.sh appends one line to .propel/codex-log.jsonl on every invocation, including the ones that fail and the ones skipped because the layer is off. It is a shell script, not a model, and it writes the entry as a side effect of actually running Codex.

A gate that claimed a Codex consult with no matching ledger entry did not consult Codex. That is the one signal in the whole system that cannot be forged, because nothing that could forge it is involved in writing it.

Reading it

propel codex log          # last 20 consults
propel codex log -q       # also show the question sent each time
propel codex log --all    # everything
propel codex status       # is it on, reachable, and being used?
  TIME      GATE                      OUTCOME       TOOK
  17:42:08  Gate 2 (design)           reply         8.1s
  17:51:20  Gate 3 (component 1)      no findings   4.3s
  18:04:57  Gate 4 (diagnosis)        UNAVAILABLE   0.1s
            ℬ codex CLI not found on PATH

  3 consults · 2 replied · 1 unavailable

Inside a session, /codex-log does the same thing and reads the ledger rather than answering from the conversation. Ask "was Codex involved in the encoder design?" and it looks for the entry, and tells you plainly if there isn't one.

Seeing it live, in the status line

The ledger answers "did it run?" after the fact. For "is it on, and did it just fire?" while you work, Propel ships a status-line segment:

[Opus 5] │ propel git:(master*) │ codex ● 3 · 7s
Context ░░░░░░░░░░ 0%
SegmentMeans
codex ● 3 · 7sEnabled. Three consults today; the last one 7 seconds ago — so it fired on the gate you just passed.
codex ● readyEnabled and reachable, nothing consulted yet today.
codex ✘ off/disable-codex is in effect for this project.
codex ⚠ no cliEnabled, but the Codex CLI isn't installed or on PATH — gates are running single-model.

Wire it up with:

propel codex statusline             # dry run — shows exactly what it would change
propel codex statusline --install  # apply (backs up settings.json first)
propel codex statusline --remove   # undo

It extends an existing claude-hud status line through its --extra-cmd hook rather than replacing it, so you keep your model, context and cost segments. In a project without .propel/ the segment prints nothing and is omitted entirely, so it never clutters unrelated work. It renders in about 20 ms and can't block: a status line that hangs is worse than one that's missing.

What it does not record

The question, never the material. Briefs carry source code, diffs and diagnoses; the ledger carries only the QUESTION block, the outcome, the duration and the decision-point label. What was sent was in the conversation — it is deliberately not written to disk. The file lives under .propel/, which propel init gitignores, and it is one short line per consult.

Guardrails

When Codex Isn't There

If the codex CLI is missing or not authenticated, Propel says so once per session, with the fix, and then proceeds single-model.

Codex is not available (not installed / not logged in), so this gate is
single-model. Run `propel` to set it up, or /disable-codex to stop
seeing this.

A missing second model never blocks a gate. Single-model review is worse than dual-model review; it is not worse than no review. The notice appears once and is not repeated — the state is recorded in .propel/codex.json.

Turning It Off

There is no invoke command, because there is nothing to invoke. The only two controls are switches:

CommandEffect
/disable-codexSingle-model for this project until re-enabled. Writes {"enabled": false} to .propel/codex.json.
/enable-codexTurn the dual-model layer back on, and re-check whether the CLI is actually reachable.

When disabled, gates, questioners and auditors all still fire. Only the second opinion goes away. Propel prints one ◇ Codex disabled — single-model gate. line at the first gate of each session, so you always know which regime you're in, and then stops mentioning it.

Reasons people turn it off are usually diagnostic and worth taking seriously: it's slow (Gate 3 fires most often — raise its threshold instead), it always agrees (that's a real failure of the layer), or the code can't leave the machine (completely legitimate, and the right call).

Setup

The easiest path is propel with no arguments. It detects whether Codex is installed and signed in, and installs it with OpenAI's native installer (a single binary, no Node). On a laptop the setup console then opens a terminal for the OAuth flow. On a cluster, the terminal setup runs codex login --device-auth inline, and you approve the code from any device. See Getting Started for both versions.

Manually:

curl -fsSL https://chatgpt.com/codex/install.sh | sh
codex login                    # on a cluster: codex login --device-auth

Propel drives codex exec directly through .claude/scripts/codex-consult.sh, which carries the timeout, the enabled-check, and the auth-error handling. There is no plugin layer in between.