Code2Act

Discrete Actions for Naturalistic Behavior in Embodied Biomechanical Animal Models

Kaiwen Bian 1,2 Charles Y. Zhang 3 Yuanjia Yang 1,2 Aidan Sirbu 4,5 Eric J. Leonardis 1,2 Blake A. Richards 4,5 Bence P. Ölveczky 3 Talmo D. Pereira 1,2
1 Salk Institute for Biological Studies 2 UC San Diego 3 Harvard University 4 Mila 5 McGill University
NeurIPS 2026 Workshop on Neuro-Symbolic Embodied Intelligence (NEmo) Oral
Paper (coming soon) Code Citation

Abstract

Embodied symbolic systems depend on a discrete vocabulary that compresses continuous motion into named actions, which make planning and annotation tractable. Many such vocabularies exist, from unsupervised segmentation to hand annotation, but they only describe motion, and making a body perform one of their symbols requires a hand-engineered controller or a policy trained for that symbol. We present Code2Act, which turns any discrete vocabulary into motor commands for a 38-actuator biomechanical rat. One recurrent controller reads only the symbol stream and proprioception, predicts the motor latent that a frozen pretrained imitation model would supply, and decodes through that model's inverse-dynamics module. A single training run then covers an entire vocabulary with no per-symbol reward and no per-symbol training. We drove one controller from three vocabularies of different provenance and size. An independent labeler can recover each steering symbol from the motion it produced, across arbitrary initial poses and external velocity impulses. Symbols are executed by the body and realized as distinct motion, and novel combinations of symbol pairs can produce within-distribution motion patterns.

The MLB controller following a chain of six symbols (walk, rear, walk, still, rear, turn) from a test-set start pose. The body's color is the commanded symbol.

Behavior has names. Bodies need commands.

Behavioral neuroscience and embodied AI both describe behavior with discrete vocabularies. Keypoint-MoSeq splits motion into syllables, pose classifiers assign discriminative classes, annotators write ethograms, and planners emit symbols. Over such a vocabulary, planning and annotation become tractable. But none of these symbols can be executed: they describe recorded motion, and making a body perform one usually takes a hand-engineered controller or a policy trained, with its own reward, for that symbol alone.

Code2Act closes that gap for the 38-actuator biomechanical rat of MIMIC-MJX. One recurrent controller reads only a stream of symbols and the body’s proprioception, and turns it into torques. Because the symbol stream is its only command input, a single training run covers an entire vocabulary, and the vocabulary can be swapped for any other labeling without changing the controller’s design.

On this page we walk through what the controller does, with the rollouts behind the paper’s figures rendered as video and several of the analyses made interactive.

A note on the examples. The video clips on this page are curated for the best visualization: from each set of rollouts we show an example chosen by a fixed rule, usually the clearest or best-recovering one, and each caption states that rule. The numbers are not curated. Tables, charts, round-trip bars and success rates use all rollouts, as in the paper.

How Code2Act works

Code2Act is built in three stages, each run as its own process. A pretrained imitation model supplies both a privileged teacher and an inverse-dynamics decoder. A steering strategy assigns every reference frame a discrete symbol. And a recurrent controller learns, by reinforcement learning on its own rollouts, to reproduce the motor intention the imitation model would have produced from the full reference, while seeing only the symbols.

interactive
KLaₜproprioception sᵖ (277-d)Reference motionmotion captureGoal sᵍ5-frame horizon · 640-dSteering strategyKPMS · KVAE · MLBsymbols cₜ … cₜ₊₄GRU stackhₜDistill head gμᵈ, σᵈEncoder qφμᵉ, σᵉ❄zInverse dynamics pθ[z, sᵖ] → aₜMuJoCo MJX rat38 actuators · 100 HzTracking rewardPPOinitialized from Stage 1
One controller, any vocabulary

A steering strategy turns the reference motion into a stream of discrete symbols. A recurrent controller reads only that stream and the body's proprioception, predicts the motor intention a pretrained imitation model would have produced, and decodes it to torques through that model's inverse-dynamics decoder.

The goal trajectory only reaches a frozen encoder that supplies a KL target during training. It never enters the action path.

symbol streamnow: Turn L
Figure 1. The pipeline, stage by stage.

Pick a tab to light up the modules and data paths active in each stage. The symbol box cycles through real bout sequences of each vocabulary. In the controller the goal trajectory reaches only the frozen encoder, which supplies a stop-gradient KL target and never the action path.

The key property comes from the pretrained latent. The imitation model’s intention encodes the change the body must apply to reach the next pose from its current state, so no single intention can serve a symbol across every body state. One symbol therefore maps to many intention trajectories, each selected by the body’s state at that moment. That is why a held symbol can reach the same motion from very different starting points, as the experiments below show.

Three vocabularies, one controller

We train the same controller design on three steering strategies (one training run each) that differ in where they come from and how large they are:

Below, each code is held constant on one body.

vocabulary
Figure 2. Every symbol, held.

One body holds each code for six seconds, drawn as in the paper’s codebook figures: the live body, a fading trail of its recent past, and its root path, colored by behavior (hue) and direction (shade). For each code we show the body whose motion the labeler reads as the code’s most common behavior, and among those the one with the median path length. Click a clip to enlarge it. Several directional variants share a base motion, and Turn L is executed as a tightly curving left walk.

Executed on the body, not every symbol is distinguishable. The nine MLB symbols steer to about five distinct behaviors: there are only about 100 labeled Walk R frames in the corpus, and a rear is the same movement whether it is labeled left or right. Keypoint-MoSeq is more redundant still, with usage concentrated on about 7 of its 100 syllables. Running a vocabulary on a body measures how many of its symbols correspond to distinct behaviors, a check the labels alone cannot provide.

Steer it yourself

The controller is not limited to holding one symbol. Chain symbols into a command stream and it moves between motor programs.

interactive
0.0 s · drag the timeline to scrub
controller hidden state
PC1 →

Faint clouds: GRU states while each symbol is held. The bright trace is this rollout, colored by the commanded symbol.

Figure 3. Interactive sequence viewer.

Each preset chains MLB symbols into a command stream for the paper’s MLB controller, running decoder-only from a test-set start pose. The command lane is the symbol stream; the readout lane is what the independent kinematic labeler reads from the motion. The right panel follows the controller’s GRU state on the first two principal components of its states while each symbol is held (faint clouds), so you can watch the state leave one symbol’s region and settle into the next. Drag the timeline to scrub. For each preset we rolled out 48 bodies from different start poses and show the one whose readout best matches the command; the agreement of the shown body and the median over all 48 are printed under the viewer.

Symbols steer realistic behavior

Checking individual symbols is only meaningful if the controller’s motion looks like a rat. We quantify realism with the kernel inception distance (KID) between controller rollouts and the expert MIMIC-MJX policy, computed in the feature space of a small VAE trained on expert motion. Each controller is driven by five code streams: its original codes, samples from the empirical transition matrix, two learned autoregressive priors (GPT and Mamba), and uniformly random codes.

Table 1. KID to the expert policy (lower is better), mean ± std over three feature-extractor seeds. The expert’s self-distance floor is about −0.1.
Steering strategyOriginalTransitionMambaGPTRandom
KPMS (K = 100)0.251 ± 0.0070.378 ± 0.0120.298 ± 0.0140.268 ± 0.0140.664 ± 0.024
KVAE (K = 25)0.247 ± 0.0310.266 ± 0.0220.292 ± 0.0260.244 ± 0.0150.998 ± 0.099
MLB (K = 9)0.162 ± 0.0270.198 ± 0.0360.246 ± 0.0200.205 ± 0.0370.729 ± 0.073

Every structured stream reaches a KID far below the random stream of the same strategy, with the original codes at or near the lowest and the learned priors close behind: realism tracks the temporal structure of the symbol stream. This is a property of the motion distribution, so we report it as a table rather than as a single clip.

Inside the controller

The body and the controller organize the same symbols differently. We compare per-symbol dynamics with dynamical similarity analysis (DSA), once on the body state and once on the GRU state. On the body, symbols separate cleanly: the three rears group with Immobile and the walks form their own block. Inside the controller the blocks blur, and the turns group with Walk R rather than with the other walks. This mirrors the data: in the corpus a turn is mostly a curving walk, and Walk R is the rarest symbol.

DSA distances between MLB symbols on the body state
(a) body state
DSA distances between MLB symbols on the GRU state
(b) controller (GRU) state
Walk LWalk RWalk SRear LRear RRear STurn LTurn RImmobile
Figure 4. Symbol dynamics on the body and in the controller.

Pairwise DSA angular distance between per-symbol operators (the paper’s Fig. 2a,b), each matrix reordered by clustering its own distances and drawn on its own color scale; the axis strips give each symbol’s category. Behavior labels play no part in the ordering.

The matrices summarize dynamics; the trajectories themselves are below. Each held symbol starts from a random pose, and the GRU state travels into a region of its own.

Figure 5. The hidden-state progression, in 3D.

GRU-state trajectories of the MLB controller holding each symbol, projected on the first three principal components. Press play, or drag the slider, to grow every trajectory from its start pose; drag the plot to rotate it. Click a symbol to hide it, or double-click to show it alone.

The same symbol, from anywhere

A symbolic layer needs each symbol to mean the same behavior whenever it is issued, from whatever state the body is in. A symbol that merely replayed a stored trajectory would work from one start pose. Because each symbol is defined outside the controller, re-labeling the motion it produces scores the controller against a criterion it was never trained on. We hold each symbol from 60 start poses and let the strategy’s own labeler read the motion back.

hold symbol
reads Walk L · symbol
reads Walk L · symbol
reads Walk L · symbol
reads Walk L · symbol

Four bodies, four random start poses, one held symbol: Walk L. Each tag is the readout of the same kinematic labeler that scores the round trip, computed on that body's settled motion.

symbol 48% · behavior 69%
Walk L67%33% → ImmobileWalk R73%23% → Rear RWalk S28% → ImmobileRear L62%32% → ImmobileRear R73%Rear STurn L98% → Walk LTurn R88%Immobile57%43% → Walk Lcommanded symbolsame behavior, other dir.different behavior

60 start poses per symbol, each strategy read back by its own labeler. Chance is 11% (symbol) and 25% (behavior).

Figure 6. Round trip: command a symbol, read it back.

Left: four of the 60 start poses for the selected MLB symbol (a seeded random draw), with the readout the round trip assigns to each. Right: each commanded symbol’s readouts split into the symbol itself, the same behavior in another direction, and a different behavior; the most common wrong destination is printed for large error segments. Switch strategies to compare KPMS and KVAE, which are read back by their own tokenizers.

For MLB the commanded symbol is recovered in 48% of instantiations and its behavior in 69%, against chance levels of 11% and 25%. Most of the shortfall is the directional collapse seen above: 8 of the 9 symbols place their most common readout inside the correct behavior. The exception, Turn L, reads as Walk L in 98% of instantiations, which makes it the most consistent symbol in the codebook, executed as a curving walk.

Kick it, and it comes back

The round trip only varies the start pose. We also perturb the body during execution: after it settles under a held symbol, a single-step velocity impulse of scale a hits it, and the controller, still reading only the symbol, has to bring it back.

held symbolimpulse a
unperturbed twinsame body, kick at 2 s

Both panels are the same body from the same start pose, holding Rear R. On the right a single-step velocity impulse hits at 2 s (the body flashes red), and the controller, still reading only the symbol, tries to bring it back to the motion of its unperturbed twin. For each impulse size we show the best-recovering of the 48 bodies kicked per symbol; the curves on the right average all 48, including those that do not recover.

displacement recovered after 5 s0%50%100%00.61.22.44impulse strengthpose deviation from unperturbed0.000.300.6000.61.22.44impulse strengthdistance between symbols
ImmobileRearWalkTurn
Figure 7. Kick and recovery.

The MLB controller. Left: one body and its unperturbed twin from the same start pose, holding Rear R or Walk L, kicked at the impulse size you pick. The kicked body flashes red at the impulse and fades back as it recovers. For each impulse size we show the best-recovering of 48 bodies: for the rear, among bodies that took at least a median-sized shove, the one whose posture ends closest to its twin’s; for the walk, among bodies walking steadily through their last three seconds, the one closest to its twin’s pace. Right: the share of the impulse-induced displacement that has decayed five seconds later, and the pose deviation from unperturbed rollouts against the unperturbed floor (dotted) and the distance between different symbols (dashed), averaged by behavior over all 48 bodies per symbol.

Five seconds after the impulse, 85% (MLB), 92% (KPMS) and 74% (KVAE) of the induced displacement has decayed on average, with little dependence on impulse size. Across a 6.7-fold range of impulse, MLB’s peak displacement grows 7.9-fold but its final displacement only 2.0-fold. The deviation from unperturbed motion grows with the impulse yet stays below the distance between distinct symbols for every behavior at a = 0.6. For MLB, rears cross that line at a = 1.2 and turns at a = 4.0, while walks and immobility never do; a rear sits near the limit of static stability. Recovery is partial and posture-dependent, but it holds across every tested symbol and strategy.

Transitions the corpus never contains

Can the controller execute symbol pairs it has never seen? For MLB, 30 ordered pairs have zero bout-level transitions anywhere in the reference corpus. We command each as two phases of 250 control steps, holding A and then B, from 60 initial states.

from (row) → to (column)
Walk LWalk RWalk SRear LRear RRear STurn LTurn RImmobileWalk LWalk RWalk SRear LRear RRear STurn LTurn RImmobile24313111415473134371312326204817532292642351921918432915256104
never in corpus (click) observed (number of bouts)
hold Rear Lthen Walk R

Rear L → Walk R never occurs in the training corpus. Across 50 start poses the labeler reads the held rear in 74% of bodies and the commanded walk after the switch in 98%.

Figure 8. Unobserved symbol pairs.

The matrix lists every ordered pair of MLB symbols. Colored cells are the 30 pairs that never occur in the corpus; click one to play its rollout. For each pair we show, among the 60 start poses, a body the labeler reads as the held symbol and then the commanded one with no unrequested behavior in between, choosing the one whose rear height, walking speed, turning or stillness is clearest. Grey cells occur in the corpus, with their bout counts. Below the video, the labeler’s reading of the held symbol and of the commanded one after the switch, across 50 diverse start poses.

Table 2. Composed MLB sequences, KID to the expert (lower is better), mean ± std over three feature-extractor seeds.
Code sequenceKID ↓
Corpus sequences0.204 ± 0.023
Unobserved pairs0.418 ± 0.045
Random codes0.735 ± 0.083

The resulting motion is less realistic than corpus-derived sequences but stays far from the random-code level, and the commanded behaviors remain recognizable pair by pair. Sequences with no support in the data still fall within the distribution of rat motion.

Discussion

Code2Act provides a symbolic layer that acts through physics: its symbols are simple to state, they drive the body, and they hold under perturbation. A practitioner with a segmentation, an ethogram or a set of annotated behaviors obtains an executor for all of it from a single training run, with no reward designed per behavior, no policy trained per symbol, and no controller written by hand. Swapping one vocabulary for another changes only the input stream, as we showed by training the same controller design on a Keypoint-MoSeq segmenter, a VAE codebook and a hand-written label set.

The resulting command space is finite, and the interface reports which of its symbols the body can actually realize as distinct behaviors. The retained symbols support counterfactuals: held, composed into unseen sequences, and checked directly on the body. For neuroethology, this offers a way to test a behavioral vocabulary by executing it. For symbolic control, it offers a motor layer that does not have to be built one behavior at a time.

The main prerequisite, and limitation, is a pretrained imitation model for the body of interest, which the interface reuses rather than replaces.

Citation

@inproceedings{bian2026code2act,
title = {Discrete Actions for Naturalistic Behavior in Embodied Biomechanical Animal Models},
author = {Bian, Kaiwen and Zhang, Charles Y. and Yang, Yuanjia and Sirbu, Aidan and
Leonardis, Eric J. and Richards, Blake A. and {\"O}lveczky, Bence P. and
Pereira, Talmo D.},
booktitle = {NeurIPS 2026 Workshop on Neuro-Symbolic Embodied Intelligence (NEmo)},
year = {2026}
}

Code2Act builds on MIMIC-MJX and track-mjx for the biomechanical rat and imitation pipeline, Keypoint-MoSeq for behavioral syllables, and MuJoCo and Brax for simulation and training.