Propel

Not an autonomous agent. A research coding assistant. Propel automates the tedious half of research engineering — tracing code, checking implementations against the paper, hunting silent bugs, remembering what already failed. It stops at every decision that is actually yours. You decide; the machines execute.

Kaiwen Bian · Yuer Tang

Overview

Propel is not an autonomous agent, and it is not trying to become one. It does not choose your research question, run experiments unattended, or decide what counts as a result. It automates the dirty work — reading the codebase, tracing data flow, checking every equation against the paper, hunting the bugs that don't crash, remembering what failed last month — and then it stops and asks you the questions only you can answer.

That split is deliberate, because it tracks what the technology is actually good at. Machines are excellent at the mechanical, repetitive, easily-botched work that researchers do badly precisely because it is boring. They are bad at judgment: what question is worth asking, which variant of the method matters, what threshold makes a result mean something, whether to believe a number. Automation that crosses that line doesn't save you work — it makes decisions you never saw, in code you will publish.

The reason structure is needed at all is that an unconstrained LLM writes the mean of its training data. Ask it to "add a diffusion policy head" and you get a blend of every diffusion codebase it has seen: a cosine noise schedule from one repo, ε-prediction from another, an EMA decay from a third. The code runs and may even train, but it is not the variant your paper describes. This matters most in research code, where mistakes are quiet: a broadcasting bug in a loss function doesn't crash, it just gives you a training run whose numbers are wrong — and a plot you believe.

Who Decides What

Every capability Propel adds falls on one side of a line. On the left, work that a machine does identically every time and a tired researcher does not. On the right, decisions that determine whether the experiment means anything.

Two columns. Machine executes: reading the codebase, checking code against the paper, hunting silent bugs, guarding against regressions, remembering across sessions, arguing with itself. You decide: what question is worth asking, which paper and which variant, what counts as a result, which trade-off to accept, when the evidence is enough, whether to believe the number.
Figure 1. Propel automates the left column aggressively and refuses to touch the right one.
The line between the columns is the product. An agent that crosses it is not more helpful — it is less auditable.

This is why Propel's gates are not a speed bump on the way to autonomy. They are the point. The framework is built so that the mechanical work happens without you having to ask for it, and the judgment calls cannot happen without you.

The Pipeline

Propel enforces five human-in-the-loop gates, two questioner checkpoints, automatic auditor dispatch after every code change, and an automatic second-model consult at every gate. Investigation happens before implementation, design decisions are explicit, and silent bugs are caught before they reach training. Each retrospective feeds a working memory of past experiments and designs that is loaded into the next session, so the agent checks what was already tried before proposing it again.

Propel pipeline: seven stages from intake to retrospective, with human gates G0 to G4, questioners Q0 and Q1, an automatic Codex consult at each gate, a per-component review loop, a working-memory loop into the next session, and the stages each mode runs.
Figure 2. The Propel pipeline: seven stages, five human gates, two questioners, and a Codex consult at each gate. The bars below show which stages each mode runs.
Intake → G0 → Q0 → Investigate → G1 → Q1 → Design & plan → G2 → Implement (G3 after each component) → Debug → G4 → Train → Retrospective → working memory for the next session

At each gate, the agent stops and asks structured questions that reveal design assumptions — never "shall I proceed?" but "should we [A] or [B]? A means [trade-off], B means [trade-off]." The questioners (Q0, Q1) ground the work in concrete reference implementations and specific implementation details before the agent acts.

Two Models at Every Decision

A model reviewing its own plan is a model grading its own homework. Propel consults a second model — OpenAI Codex — automatically, at every major decision point. There is no command to invoke and nothing to remember. A second opinion you have to request is one you get only when you already suspect you need it, which is the opposite of when it helps.

Five steps at every gate: Claude drafts, Codex critiques, Claude verifies every file and line against the repository, an attributed card is produced, and you decide. Below: what Codex never does — write to your repository, speak to you unfiltered, decide anything, or get invented.
Figure 3. Every gate runs the same five steps. Exactly one of the participants is a human, and that one decides.

Claude writes a short brief, a codex-bridge subagent runs the consult in an isolated context, and every file:line Codex cites is checked against your actual repository before you see it. Codex hallucinates paths and line numbers; that verification step is what makes its genuinely-good findings usable. Anything that cannot be confirmed is labelled [codex · unverified] — never promoted to a finding, never silently dropped.

Claude says so every time, with a ◆ Consulting Codex line naming the question being sent, and every finding in the resulting card is tagged [claude], [codex], [both] or [codex · unverified]. You can always tell which model said what — a blended voice would destroy the entire point of having two.

Where it fires

Gate 0 scope, Gate 1 findings, Gate 2 design, Gate 3 audit, Gate 4 diagnosis, bug classification, 3-strike breaks, and retrospective conclusions. Not on trivia — over-firing turns the announcement into noise, which is worse than not firing at all.

What it never does

Write to your repository. Speak to you unfiltered. Decide anything. Get invented — if the CLI did not run, there is no Codex line in the card, and Propel says it ran single-model.

Agreement is weak evidence

Two models trained on overlapping internet text agreeing means little, and Propel reports it that way rather than dressing consensus up as confirmation. The disagreements are what deserve your attention.

Don't want it? /disable-codex turns the layer off for the project. Gates, questioners and auditors all still fire — only the second opinion goes away. A missing or unauthenticated Codex never blocks a gate: Propel says so once and continues single-model.

Four Modes — selected automatically

Not every session needs the full pipeline, so Propel filters which skills and gates are active. You do not choose the mode. The agent reads your first message, picks one, and says which and why in a single line — then switches again on its own when the work crosses a boundary. Asking someone to classify their own problem before they have started working on it is exactly the kind of friction Propel exists to remove.

Four example first messages, each mapped to the mode Propel infers from it and the gates that then fire: a question about how code works selects Researcher; a request to port a paper section selects Engineer; a NaN loss with a traceback selects Debugger; launching a cluster sweep selects Trainer.
Figure 4. First message to inferred mode to active gates. /switch overrides it whenever you disagree.

Researcher

Gates 0 & 1

Literature, investigation, deep research. Picked when the question is about the problem space rather than the code.

Engineer

All Gates (0–4)

Full pipeline with every auditor. Picked when there is something to build — and when the request is ambiguous, since it filters nothing out by mistake.

Debugger

Gates 0, 1 & 4

Root-cause analysis with evidence-backed classification. Picked when there is a specific wrong behavior to explain.

Trainer

Gate 4 (runtime)

Launch, monitor, and fix CUDA/OOM/path errors. Picked when the code is settled and the problem is execution.

Crossing a boundary switches the mode rather than refusing the request. Ask a Trainer-mode session to change a loss function and it moves to Engineer — and tells you why: that change alters what the run measures, so it needs the design and audit path, not a runtime patch. The mode boundary exists to make sure the right gates fire, not to withhold capability. A narrower mode runs fewer gates; it never lowers the bar at the ones it runs.

What Runs Without You Asking

A framework that requires you to remember a command before investigating, or a flag before wanting a second opinion, has moved the cognitive load it was supposed to remove. In Propel, you describe the problem. Everything mechanical dispatches itself — and narrates what it did, so nothing happens invisibly.

TriggerWhat fires, automatically
Your first messageMode selection, announced in one line with its reason
A new taskGate 0 scoping questions, one at a time, then Q0 grounding
A question needing several files tracedSeveral investigator subagents in parallel, one question each
Any gate about to be presentedA Codex consult via codex-bridge, verified against the repo
An approved planimplementer then spec-reviewer, one component at a time
Any source editA hook names the required auditors — the dispatch is deterministic, not remembered
An approved training commandtrainer-operator, which keeps thousands of log lines out of your session
Three failed attemptsThe 3-strike break: stop, and ask what assumption all three share
Process is the agent's to decide. Research is yours. If the agent catches itself writing "you can run X if you want", it should just run X.

Twenty skills and thirteen subagents sit behind that table. You never name one — but if you want to see exactly what exists and what triggers it, the full references are in the documentation: all skills, by category, with triggers and the modes each is active in · all agents, split into read-only auditors and the workers that do the building.

See It In Action

Each mode shapes how the agent interacts with you. Watch simulated sessions for each mode.

Researcher Mode
Engineer Mode
Debugger Mode
Trainer Mode

When Would I Use This?

Three real-world starting points and how Propel guides you through each one.

1

“I have an idea and a reference codebase”

You know what you want to build. You have a reference implementation or paper to follow. You need to adapt it to your architecture and constraints.

Engineer Mode Gate 0 → Q0 → G1 → Q1 → G2 → G3 → G4
Gate 0
Scope it. Claude asks what exactly you're building, which parts of the reference to follow, and what to change. You provide the reference repo/paper.
Q0
Ground it. “Which file in the reference should I use as the starting point? What architecture pattern should I follow? What should I verify against?” — Claude asks for concrete anchors instead of guessing.
Investigation
Trace the reference. Claude reads the reference codebase, documents how it works in a scratch/ README, identifies what needs adapting vs. what can be reused directly.
Q1 → Design
Nail down details. Interface contracts, data formats, edge cases, config approach. Then a paper-to-code mapping with explicit component ordering and regression risk assessment.
Build → Audit
Implement with review. Each component gets a 3-stage review: implement → spec compliance → domain auditors (paper-alignment, silent-bug-detector, regression-guard). You approve at Gate 3 after each component.
What you get: Code that precisely matches your reference — not a plausible average. Every deviation is documented and human-approved.
2

“I inherited a codebase and need to understand it, then extend it”

Your teammate's project, a previous experiment, or an open-source repo. It generally works but you need to understand how before adding your features.

Researcher then Engineer G0 → G1 → /switch → G0 → Q0 → G1 → Q1 → G2 → G3
Researcher
Map the territory. Start in Researcher Mode. Claude traces the codebase end-to-end: entry points, data flow, key abstractions, configuration, and conventions. Everything goes into a scratch/ investigation README that persists across sessions.
Gate 1
Confirm understanding. Claude presents its findings: “Here's how the reward pipeline works, here are the 3 config files that control it, here's what surprised me.” You correct misunderstandings before anything gets built.
/switch engineer
Now build. Switch to Engineer Mode. The investigation README becomes the foundation. Q0 asks which existing patterns to follow. Q1 nails down how your new feature integrates with what's already there.
Design → Build
Extend with guardrails. Design respects existing conventions (from investigation). The regression-guard auditor verifies nothing breaks. Each new component is audited against the existing codebase, not just in isolation.
What you get: A documented understanding of the codebase that survives across sessions, plus new features that integrate cleanly without breaking existing functionality.
3

“Something is wrong and I need to find the real cause”

Training diverges, outputs look wrong, or a refactor broke something. You need to find the root cause — not just patch symptoms. Could be a code bug, a design flaw, or a config issue.

Debugger Mode Gate 0 → G1 → Classify → G4 → Fix
Gate 0
Characterize the symptom. Claude asks precise, disjunctive questions: “Is the loss exploding or collapsing? All configs or just one? After a specific commit or always?”
Investigation
Dispatch auditors. Based on the symptom, Claude dispatches relevant auditors — silent-bug-detector for numerical issues, paper-alignment-auditor for correctness, jax-logic-auditor for shape problems. Evidence is gathered, not guessed at.
Classify
Name the category. Every issue is classified as a code bug (concrete evidence: line numbers, values), design issue (requires literature backing), or config/environment problem (specific settings, version mismatches). The fix depends on the category.
Gate 4
Diagnose before fixing. Claude presents: Symptom → Root Cause → Why This Happens → Proposed Fix → Side Effects → What Won't Fix. You approve before any code changes. If it's a design issue, Claude searches the literature for validated alternatives and redirects to Engineer Mode for the redesign.
3-Strike
No infinite loops. If the same approach fails 3 times, Claude stops, re-examines assumptions, and asks whether the root cause hypothesis is wrong — instead of trying one more variation.
What you get: A proven root cause with evidence, not a symptom patch. Design issues get literature backing. Failed debugging approaches are documented so you don't repeat them.

Get Started

git clone https://github.com/KevinBian107/propel.git
cd propel

pip install -e .      # inside an activated conda env or virtualenv
propel                # setup: browser page locally, terminal prompts on a cluster

This is an editable install, so edits to the repo take effect immediately. Run it inside an activated conda env or virtualenv. Against a Homebrew or system Python, pip fails with externally-managed-environment and leaves you without the command.

Running propel with no arguments sets everything up: git, Claude Code and Codex, the two sign-ins, and Propel itself in your project. Claude Code and Codex come from the vendors' own native installers, so you need no Node, no npm and no sudo. There are two versions, and propel picks one by itself. On your laptop it opens a setup page in your browser, with one button per step. On a cluster (an HPC node, a container, a Jupyter terminal, anything over SSH) it runs the same steps as prompts in the terminal. Each sign-in there shows a code you approve from any device, so nothing needs forwarding.

Prefer the command line? cd into your project and run propel init. Want a specific version? propel launch (browser) or propel setup (terminal).

Then start claude and just describe what you're working on. Propel picks a mode, fires Gate 0, and tells you what it's doing. You never have to name a skill, an agent, or a model. Run /intro if you'd rather see the full tour first, and see the Getting Started guide for a complete walkthrough.

What You Bring to the Gates

Propel's constraints are necessary but not sufficient. Automating the mechanical half only pays off if the judgment half actually happens — and that half is yours. The framework guarantees the agent will stop and ask; the quality of the output depends entirely on what you bring when it does:

Your InputWhy It Matters
Research questionNot "implement X" but "test whether X improves Y under condition Z." The more specific, the less the agent guesses.
HypothesisWhat do you expect and why? This is what auditors verify against.
MethodWhich paper, which equations, which algorithmic choices. The agent cannot infer "use stop-gradient on the codebook as in Section 3.2."
Domain knowledgePitfalls not in any paper, configs that silently fail, things that only work in your setup.

Acknowledgments

Propel combines ideas from multiple sources: