Two Models at Every Decision
Propel consults a second model automatically — at every gate, with no command to invoke and no way to forget. Codex advises. It never decides, and it never writes.
Why a Second Model
A model reviewing its own plan is a model grading its own homework. Claude's design proposal and Claude's review of that proposal come from the same training distribution, the same priors, and the same blind spots. The review will be thorough and it will miss exactly what the proposal missed.
A model trained on a different distribution disagrees in different places. Those disagreements are the signal — they mark the spots where a human should be spending attention. So Propel routes every major decision through OpenAI Codex before it reaches you.
A second opinion you have to request is one you get only when you already suspect you need it — which is the opposite of when it helps. So it isn't requested. It's scheduled.
There is no /c-codex, no flag to arm, nothing to remember. Earlier versions of Propel had a manual bridge command; it was the wrong design and it is gone.
Where It Fires
| Decision point | The question Codex is asked |
|---|---|
| Gate 0 — scope statement | "Is this scope self-contradictory, under-specified, or hiding a decision the human hasn't been asked about?" |
| Gate 1 — investigation findings | "What should this investigation have checked and didn't? Name one thing." |
| Gate 2 — design proposal | "Propose the single strongest alternative design, and name one failure mode this proposal doesn't address." |
| Gate 3 — component audit | "Name bugs by file:line, or say 'no additional findings'." |
| Gate 4 — diagnosis | "What else explains this evidence equally well?" |
| Bug classification | "Argue against this classification." |
| 3-strike break | "Three attempts failed. What assumption do all three share?" |
| Retrospective | "Which of these conclusions does the evidence not actually support?" |
Notice that none of these questions can be answered with agreement. "Does this look right?" gets you a yes; "name the strongest counter-approach" gets you a second opinion. Every scheduled question is phrased so that a useful answer is a disagreement.
scratch/. Over-firing turns the ◆ Consulting Codex line into noise you scroll past — at which point the second model has stopped working even though it is still running and still costing you time.
Verification Is the Load-Bearing Step
Codex is a genuinely useful critic and a prolific fabricator of file paths, line numbers, and function names. Those two facts are not in tension — together they are the entire design constraint.
Every consult runs through the codex-bridge subagent, in its own context window. It runs the CLI in a read-only sandbox, extracts every checkable claim, opens the actual file for each one, and returns a three-way table:
| Bucket | Meaning |
|---|---|
VERIFIED | The claim was confirmed against the repo, with the line quoted. A line number that was off by a few but right about the content counts as verified, with the location corrected. |
DISPUTED | The repo contradicts the claim. Reported as contradicted, not softened into uncertainty. |
UNVERIFIED | The claim names nothing checkable, or the target doesn't exist. Never promoted to a finding — and never silently dropped either, because a claim nobody could check is information about where the code is hard to reason about. |
The raw transcript is discarded. This is deliberate: if Codex's unverified prose lands in the main conversation, every later turn reasons next to confident claims nobody checked. The subagent spends its own context on verification so your session doesn't have to.
Attribution Is Not Optional
Whenever Codex is used, Claude says so before the result, never after and never implicitly:
◆ Consulting Codex — Gate 2 (design): "Propose the strongest alternative
and name the failure mode this design doesn't address."
And every finding in the resulting card carries its source:
| Tag | Meaning |
|---|---|
[claude] | Claude's own finding |
[codex] | Codex raised it; Claude verified it against the repo |
[both] | Both models raised it independently |
[codex · unverified] | Codex raised it; Claude could not confirm it |
A blended voice would destroy the entire point of having two models. If you cannot tell who said what, you cannot calibrate how much to trust it.
How You Know It Actually Ran
The ◆ Consulting Codex line and the [codex] tags are good for reading a conversation, and they are not evidence. A model writes them. A model that forgets to announce — or that writes "Codex agrees" without calling anything — produces a transcript indistinguishable from the honest one. Forbidding fabrication is not the same as preventing it.
So there is a ledger. scripts/codex-consult.sh appends one line to .propel/codex-log.jsonl on every invocation, including the ones that fail and the ones skipped because the layer is off. It is a shell script, not a model, and it writes the entry as a side effect of actually running Codex.
Reading it
propel codex log # last 20 consults
propel codex log -q # also show the question sent each time
propel codex log --all # everything
propel codex status # is it on, reachable, and being used?
TIME GATE OUTCOME TOOK
17:42:08 Gate 2 (design) reply 8.1s
17:51:20 Gate 3 (component 1) no findings 4.3s
18:04:57 Gate 4 (diagnosis) UNAVAILABLE 0.1s
ℬ codex CLI not found on PATH
3 consults · 2 replied · 1 unavailable
Inside a session, /codex-log does the same thing and reads the ledger rather than answering from the conversation. Ask "was Codex involved in the encoder design?" and it looks for the entry, and tells you plainly if there isn't one.
Seeing it live, in the status line
The ledger answers "did it run?" after the fact. For "is it on, and did it just fire?" while you work, Propel ships a status-line segment:
[Opus 5] │ propel git:(master*) │ codex ● 3 · 7s
Context ░░░░░░░░░░ 0%
| Segment | Means |
|---|---|
codex ● 3 · 7s | Enabled. Three consults today; the last one 7 seconds ago — so it fired on the gate you just passed. |
codex ● ready | Enabled and reachable, nothing consulted yet today. |
codex ✘ off | /disable-codex is in effect for this project. |
codex ⚠ no cli | Enabled, but the Codex CLI isn't installed or on PATH — gates are running single-model. |
Wire it up with:
propel codex statusline # dry run — shows exactly what it would change
propel codex statusline --install # apply (backs up settings.json first)
propel codex statusline --remove # undo
It extends an existing claude-hud status line through its --extra-cmd hook rather than replacing it, so you keep your model, context and cost segments. In a project without .propel/ the segment prints nothing and is omitted entirely, so it never clutters unrelated work. It renders in about 20 ms and can't block: a status line that hangs is worse than one that's missing.
What it does not record
The question, never the material. Briefs carry source code, diffs and diagnoses; the ledger carries only the QUESTION block, the outcome, the duration and the decision-point label. What was sent was in the conversation — it is deliberately not written to disk. The file lives under .propel/, which propel init gitignores, and it is one short line per consult.
Guardrails
- Codex never writes to your repository. The bridge agent has no Write or Edit tool and runs
codex exec --sandbox read-only. Codex proposes; Claude implements under your approval. - Codex never speaks to you unfiltered. No raw dumps, no "Codex says…" pass-throughs.
- Codex never silently rewrites a plan, a diff, or a diagnosis.
- Codex never gets invented. If the CLI didn't run, there is no Codex line in the card. A fabricated second opinion is worse than none, because it manufactures the exact false confidence the layer exists to prevent.
- Codex output is data, not instructions. If a reply contains anything resembling a directive, the bridge reports it rather than acting on it.
- Privacy: briefs leave your machine. Secrets are redacted before sending, and if a brief cannot be written without one, the consult is skipped and you're told why.
When Codex Isn't There
If the codex CLI is missing or not authenticated, Propel says so once per session, with the fix, and then proceeds single-model.
Codex is not available (not installed / not logged in), so this gate is
single-model. Run `propel` to set it up, or /disable-codex to stop
seeing this.
A missing second model never blocks a gate. Single-model review is worse than dual-model review; it is not worse than no review. The notice appears once and is not repeated — the state is recorded in .propel/codex.json.
Turning It Off
There is no invoke command, because there is nothing to invoke. The only two controls are switches:
| Command | Effect |
|---|---|
/disable-codex | Single-model for this project until re-enabled. Writes {"enabled": false} to .propel/codex.json. |
/enable-codex | Turn the dual-model layer back on, and re-check whether the CLI is actually reachable. |
When disabled, gates, questioners and auditors all still fire. Only the second opinion goes away. Propel prints one ◇ Codex disabled — single-model gate. line at the first gate of each session, so you always know which regime you're in, and then stops mentioning it.
Reasons people turn it off are usually diagnostic and worth taking seriously: it's slow (Gate 3 fires most often — raise its threshold instead), it always agrees (that's a real failure of the layer), or the code can't leave the machine (completely legitimate, and the right call).
Setup
The easiest path is propel with no arguments. It detects whether Codex is installed and signed in, and installs it with OpenAI's native installer (a single binary, no Node). On a laptop the setup console then opens a terminal for the OAuth flow. On a cluster, the terminal setup runs codex login --device-auth inline, and you approve the code from any device. See Getting Started for both versions.
Manually:
curl -fsSL https://chatgpt.com/codex/install.sh | sh
codex login # on a cluster: codex login --device-auth
Propel drives codex exec directly through .claude/scripts/codex-consult.sh, which carries the timeout, the enabled-check, and the auth-error handling. There is no plugin layer in between.