Codex Consult
The automatic dual-model layer. Fires on its own at every major decision point — there is no command to invoke, and no way to forget.
Overview
The Codex page explains that a second model is consulted at every decision point, and why. This skill is how.
| Property | Details |
|---|---|
| Trigger | None. It is scheduled, not requested. |
| Active in | All four modes |
| Dispatches | codex-bridge |
| Governed by | core/CODEX.md and .propel/codex.json |
A command would put the burden on you at exactly the moments you're least able to carry it — mid-design, mid-diagnosis, head down in the problem. So there is no command.
Step 0 — Should this consult happen?
It fires at eight points and nowhere else:
| Point | Fires when |
|---|---|
| Gate 0 | The scope statement is written, before you see it |
| Gate 1 | Investigation findings are assembled |
| Gate 2 | The design proposal is complete |
| Gate 3 | A component diff is non-trivial (≥ 30 lines, or model / loss / data / training-loop code) |
| Gate 4 | A root cause is identified, before the fix is proposed |
| Classification | A bug is about to be called code / design / config |
| 3-strike | The third attempt at the same approach has failed |
| Retrospective | Conclusions are written, before they enter the registry |
Not at Q0/Q1 — those are questions for you, not for a model — nor on mode selection, reads, renames, formatting, trivial diffs, or anything in scratch/.
Step 1 — Announce, before dispatching
◆ Consulting Codex — Gate 2 (design): "Propose the strongest alternative
to the two-stage encoder and name the failure mode this design doesn't
address."
One line: the decision point, and the actual question being sent. You should be able to tell from this line alone whether it was a good question. It is not optional and it is not retroactive — if you only learn Codex was involved when the findings appear, the announcement has failed.
Step 2 — The brief
Five to fifteen lines. Facts first, position second, question last.
CONTEXT
Project: <one line>
Decision: <what is being decided right now>
Material: <the design / diff / diagnosis, inlined>
CLAUDE'S POSITION
<2-4 lines — omitted when Claude wrote the thing being reviewed>
QUESTION
<one sharp, falsifiable question>
CONSTRAINTS
Under 400 words. Cite file:line for every code claim. If you have no
additional finding, say "no additional findings" rather than restating
mine.
Brief-writing rules
- Don't bias the question. "Does this look right?" gets agreement. "Name the strongest counter-approach" gets a second opinion.
- Omit Claude's position on self-authored material — a Gate 3 diff, or a design Claude proposed. Anchoring Codex to the author's reading of their own work throws away most of the value.
- Inline the material. Codex doesn't share Claude's context, working memory, registry, or CLAUDE.md. A brief referencing "the design we discussed" produces a hallucinated review of a design Codex has never seen.
- Redact. Keys, tokens, unpublished data. If the brief can't be written without a secret, skip the consult and say why.
Step 3 — Dispatch through codex-bridge
Never run the CLI from the main context. The bridge subagent runs it, checks each claim against the repo, throws away the transcript, and returns a VERIFIED / DISPUTED / UNVERIFIED table.
The reason is context hygiene, and it's load-bearing: Codex's raw reply is long, confident, and partly wrong. In the main conversation it becomes background the model reasons next to for the rest of the session.
Step 4 — Judgment on top of verification
The bridge does the mechanical verification. Claude does the judgment.
For every VERIFIED finding: is it real? A confirmed line number is not a confirmed bug — Codex will correctly quote a line and draw the wrong conclusion from it. For every DISPUTED finding: was Codex wrong, or right about something in the wrong place?
UNVERIFIED claim to a finding. It goes in the card, labelled, or it doesn't appear. Silently dropping it is also wrong — a claim nobody could check is information about where the repo is hard to reason about.
Step 5 — The card
┌─ ◆ Codex consult — Gate 2 (design) ─────────────────┐
│ Asked: Propose the strongest alternative and name a failure │
│ mode this design doesn't address. │
│ │
│ Both models agree: │
│ • [both] The encoder must be frozen before stage 2 │
│ — weak evidence, two models agreeing is cheap │
│ │
│ They disagree: │
│ • Codebook init. [claude] k-means per §3.2. [codex] │
│ uniform, argues k-means leaks val statistics. │
│ Claude's read: codex is right about the leak, wrong │
│ that §3.2 requires uniform. This is a real decision. │
│ │
│ New from Codex: │
│ • [codex] EMA decay applied per-step, paper applies it │
│ per-epoch — vq.py:88 (verified) severity: high │
│ │
│ Codex claimed, Claude could not confirm: │
│ • [codex · unverified] "loss.py:210 detaches the codebook" │
│ — loss.py has 140 lines; no detach found anywhere. │
│ │
│ Your decision: │
│ The codebook init question is the one that changes │
│ results. Uniform, or k-means on a train-only sample? │
└─────────────────────────────────────────────────────┘
Then the gate proceeds as normal. The card is input to the gate, not a replacement for it. Codex is not a fifth reviewer whose approval advances the pipeline — you still answer the gate question.
Controls
| Command | Effect |
|---|---|
/disable-codex | Single-model for this project until re-enabled |
/enable-codex | Turn it back on, and re-check that the CLI is reachable |
If you ask for a second opinion mid-turn — not at a gate — Propel just runs the consult. Same procedure; the schedule is a floor, not a ceiling.
See also: Codex — Two Models, codex-bridge, think-deeply (the internal anti-sycophancy pass), and c-review, which folds this consult into a merged Gate 3 card.