TopoMimic

Shaped by What's Missing
Topological Invariance Simplification Discovers Behavioral Transitions

Kaiwen Bian 1,2 Eric J. Leonardis 2 Yuanjia Yang 1,2 Charles Zhang 3 Yusu Wang 1 Talmo D. Pereira 2
1 University of California San Diego 2 Salk Institute for Biological Studies 3 Harvard University
NeurIPS 2026 Workshop on Symmetry and Geometry in Neural Representations (NeurReps)

Abstract

Understanding how animals move between behaviors is a central question in neuroethology, yet most segmentation methods first partition time into discrete states, leaving transitions as the unmodeled residue between them. We take the opposite view: in a suitable feature space, sustained behaviors are dense regions of the trajectory's density landscape, and transitions are the sparse corridors and loops that connect them. Because this structure is a property of the landscape's shape rather than of any particular coordinates, it can be read off directly with topology and should persist across feature spaces. We build this landscape from frame-by-frame trajectories, extract its discrete-Morse graph, and keep two kinds of structure: paths, the corridors a trajectory follows between dense behavioral regions, and cycles, closed routes whose entry and exit take different ways. Clustering these structures yields transitions directly, with no upstream state partition. On a rat behavior dataset across three feature spaces, including the latent space of a neural-network controller that imitates the animal in a physics simulator, the approach recovers transitions at coverage comparable to state-space and topological-clustering baselines, while carrying substantially more information about the transition identity, forming clusters at least as tight as the baselines', and matching annotated transitions with less reliance on time-warping. More broadly, the results show that behavioral transitions are a topological feature of a representation: recoverable from the shape of its trajectories, and consistent across distinct feature spaces.

persistence δ = 0.023 loops recovered

The idea behind the method, on the road-network example it was built for. Points sampled around a hidden graph form a density terrain; discrete-Morse reconstruction recovers the graph as the terrain's ridges (red). Raising the persistence threshold δ cancels the least significant ridges first, until only the most persistent loops remain. In the paper the points are behavior frames and the ridges are behavioral transitions.

Behaviors are where a trajectory lingers. Transitions are what is missing.

How an animal moves from one behavior to the next, walking into a rear or a rear into a turn, is a central question in neuroethology. Yet the standard tools describe transitions only after a coarse-graining. An autoregressive HMM such as Keypoint-MoSeq assigns every frame a syllable and summarizes transitions as a matrix over syllables. A variational autoencoder partitions a latent embedding. Mapper builds a graph over overlapping density patches. In each case the states come first, and a transition is whatever lies between two of them, at whatever resolution the state assignment allows.

We take the opposite view. In a suitable feature space a sustained behavior is a dense region of the trajectory: the animal spends many frames there. A transition is the sparse corridor or loop that connects dense regions: the trajectory passes through it quickly and rarely. That structure belongs to the shape of the density landscape, a region of the space shaped by what is missing, so it can be read off directly with topology and should look the same in different coordinates.

Method overview: trajectory, skeleton, DM-cycle and DM-path, transition clusters
Figure 1. Method overview.

A trajectory is Takens-embedded and skeletonized into a discrete-Morse graph (Step 1), decomposed into DM-cycles (H₁ loops) and DM-paths (the split tree) (Step 2), and the pooled primitives are clustered under the 2-Wasserstein distance into behavioral transition clusters (Step 3). The rest of this page walks through each step with the data behind it.

This page walks through what we did, with the rollouts behind the paper’s figures rendered as video and the key constructions made interactive.

A rat, a physics simulator, and three ways to describe it

The trajectories come from MIMIC-MJX. Markerless 3D motion capture of a freely moving rat is fit by inverse kinematics to a biomechanical model, and a deep reinforcement learning policy is trained to imitate the resulting motion in the MuJoCo physics simulator. At each step an encoder maps a short window of the upcoming reference to a 16-dimensional latent intention, and a decoder combines it with proprioception to produce the actions that drive the body.

MIMIC-MJX pipeline: motion capture, inverse kinematics, imitation policy with encoder and decoder
Figure 2. Where the trajectories come from.

The MIMIC-MJX pipeline (Zhang et al., 2025). (A) Motion capture is fit to a biomechanical model and a policy learns to imitate it. (B) The encoder maps the reference window to the latent intention; the decoder turns intention and proprioception into action.

The same 842 clips (about 500 frames each at 30 Hz) are analyzed in three representations at increasing levels of abstraction:

Intention and joint angles are spaces a keypoint segmenter cannot offer, and they matter here for a geometric reason: a skeleton only forms where behaviorally similar frames recur near each other. Uncorrected drift pulls repeated instances of the same behavior apart, so we expect the cleanest structure in intention and qpos and the weakest in ekp.

The behaviors, and the transitions between them

To evaluate, and only to evaluate, every frame carries a behavior label (Walk, Immobile, Turn, Rear) and a turn direction (left, right, straight), produced by an open-source heuristic labeler and proofread by two annotators. Below is one instance of each of the twelve directed transitions, rendered from the policy rollouts in the simulator.

Figure 3. Rat behavior lexicon, in motion.

The video counterpart of the paper’s Figure 5. Each clip replays one representative ground-truth transition at half speed under the paper’s shared camera. The amber body is live; the six poses of the static figure are dropped into the scene, light to dark, as the body passes them, and the final held frame is the paper’s panel. The strip at the bottom is the ground-truth behavior over the window (tick marks are the six poses). Instances are chosen as in the paper: each side of the labeled boundary is scored against the kinematic signature of its own label (speed, yaw rate, torso height), and the instance whose weaker side scores best is shown. n is the number of instances of that transition in the dataset. Click a clip to enlarge it.

Purely rotational transitions (Turn → Immobile, Immobile → Turn) are the hardest to see from one viewpoint: the heading changes while the body stays in place.

How the method works

The pipeline has three steps. The first two run on each clip separately and turn it into a handful of topological primitives; the third pools the primitives from all clips and clusters them.

interactive
PER CLIPPOOLED OVER CLIPSBehavior clip≈500 frames, 30 HzFeature spaceintention · qpos · ekpTakens embeddingxₜ = [pₜ, pₜ₋τ, pₜ₋₂τ]Discrete-Morse Gk-NN ρ · sparse Rips · δSkeleton Gdensity ridgesDM-cyclesH₁ · min cycle basisDM-pathssplit-tree pairsW₂ distancesexact OT · L = 100MDS → K-meansK = 25Transition clusterse.g. Rear → Turn
Transitions read off the shape of a trajectory

Each recorded clip becomes a point cloud in a feature space. Its discrete-Morse skeleton traces the ridges of the cloud's density, and two kinds of structure are cut from it: DM-cycles, loops whose way out and way back differ, and DM-paths, corridors between dense behavioral regions.

Primitives from all clips are pooled and clustered under the 2-Wasserstein distance. Each cluster is a candidate behavioral transition. No state partition is fit anywhere, and ground-truth labels are used only to evaluate the clusters afterwards.

DM-cycle 1 (noise)DM-path 1 (noise)DM-cycle 2DM-path 2DM-cycle 3
clip 447, real time. Strip: ground-truth behavior. Rows below it: the 5 DM primitives the pipeline cuts from this clip in intention space, each drawn over the frames it covers (reds and purples = DM-cycles, green = DM-path).
ImmobileRearWalkTurn
Figure 4. The pipeline, step by step.

Pick a tab to light up the modules active in each step. The right panel shows each step on the animal itself, using one real clip (clip 447: Rear, Walk, Rear) and its stored intention-space outputs. Takens: the pose now with the poses τ and 2τ frames earlier, which together form one point of the cloud. DM skeleton: the body tinted by the k-NN inverse density of that point. Primitives: each stored DM-cycle and DM-path of the clip, replayed. W2 clustering: three real DM-cycles from the paper’s DMC-Int clustering (seed 0) with their exact W₂² distances.

Why a skeleton holds the transitions

The discrete-Morse graph reconstruction we build on was developed to recover road networks from noisy GPS samples, and that two-dimensional setting is the easiest place to see what it does. Points sampled around a hidden graph induce a density field. Viewed as a terrain, the hidden roads are its mountain ridges: formally the 1-stable manifolds, the integral lines that connect density peaks through saddles. Estimating them directly is fragile, so the construction works with discrete Morse theory on a simplicial complex over the points and uses persistent homology to cancel ridges that are not significant, under a threshold δ.

Below is that construction run for real (for a stage-by-stage account of the algorithm itself, see the PCD walkthrough). The point cloud and terrain are the road network of the paper’s Figure 3A; the red graph is the raw output of the same discrete-Morse backend we use on trajectories, recomputed at each δ.

interactive
persistence threshold δ = 0.5drag the terrain to rotate
0.020.050.10.20.350.50.81.2
271 graph edges
2 DM-cycles (β₁)
34 DM-paths
Figure 5. Ridges of a density terrain, recovered by discrete-Morse reconstruction.

Drag to rotate. The slider sets the persistence threshold δ; each position is a separate run of the backend (k = 5) on the same 796 points. Up to δ = 0.35 the recovered graph keeps three independent loops (the two diamonds plus one small noise loop) and every outward branch; at δ = 0.5 only the two diamonds remain as loops, at δ = 0.8 one, and at δ = 1.2 everything has been cancelled. The “DM-cycles + DM-paths” view decomposes each recovered graph exactly as Step 2 does on trajectories. Toggle “hidden graph” to compare with the roads the points were sampled from.

On a trajectory the same logic applies, with behaviors in place of junctions. Frames of a sustained behavior pile up into a density peak; the frames where the body reconfigures from one behavior to the next are few and spread out, so they form a low-density ridge between peaks. The skeleton is that set of ridges, and Step 2 cuts it into two kinds of piece.

animation
density fbacsaddle 2saddle 1global min

A DM skeleton: a split tree with one loop attached. Height is density f; peaks are dense behavioral regions.

  • DM-path · peak c dies at saddle 2
  • DM-path · peak a dies at saddle 1
  • DM-path · essential: global max to global min
  • DM-cycle · the loop (minimum cycle basis)
Figure 6. The two DM primitives on one skeleton.

The synthetic skeleton of the paper’s Figure 3B (a split tree plus one loop), decomposed step by step. The DM-cycle is the H₁ loop: a recurrence where the trajectory leaves a region and returns by a different route. Because every point of a Takens cloud is a short window of motion, that recurrence is one of dynamics, so a gait closes a loop without ever passing through a common pose. DM-paths are the persistence pairs of the remaining split tree, read as transition corridors. This is standard 0-dimensional persistence on a tree; what is ours is the reading of each pair as a corridor.

What the discovered clusters look like

Pooled over all clips, the primitives are compared as distributions under the 2-Wasserstein distance and clustered (Step 3). The baselines assign every frame a discrete label, a syllable or a Mapper node, so their transition units are ordered pairs of consecutive sustained labels. Below are clusters from the four variants of the paper’s main comparison. Within a cluster, the two members should trace the same gross movement.

method
Immobile → Turn Leftcluster 1 · 64 members
clip 117 · 291–358 · 0.5× speed
clip 41 · 241–397 · 1× speed
Immobile → Rear Straightcluster 2 · 31 members
clip 1 · 52–196 · 1× speed
clip 1 · 352–463 · 0.5× speed
Rear → Turn Leftcluster 10 · 23 members
clip 11 · 245–479 · 1× speed
clip 321 · 71–280 · 1× speed
Rear → Immobile Leftcluster 20 · 32 members
clip 58 · 126–346 · 1× speed
clip 70 · 84–244 · 1× speed
Immobile → Rear Leftcluster 21 · 42 members
clip 38 · 235–484 · 1.5× speed
clip 123 · 251–405 · 1× speed
Figure 7. Discovered transition clusters, in motion.

The video counterpart of the paper’s Figures 10 to 13 (seed 0). The title of each cluster is its majority-vote ground-truth label; each figure shows up to six uniquely labeled transition clusters and two members per cluster. Members are chosen as in the paper: the first 14 members of a cluster are scored on whether both labeled behaviors appear in order and on how few frames belong to neither, with a penalty for members too spread out to fit the shared camera, and the best two are shown. Each video plays a member from its first to its last frame (for DM primitives, through the end of the last Takens window), with the paper’s six poses dropped as the amber body passes them; long members are played faster, at the speed noted on the clip. *Ours.

The baselines’ units are often long: a Keypoint-MoSeq pair can span most of a 500-frame clip, because a transition is the boundary between two sustained syllables. A DM primitive is a few seconds of behavior (on average 62 to 191 frames, depending on primitive and space), cut where the skeleton says the trajectory changes.

How much transition information survives the coarse-graining

Every method turns a clip into a per-frame label T: a primitive index for ours, a pair type for the baselines. We ask how much that label tells about the directed transition class Y each frame belongs to, by the mutual information I(T;Y), and place every configuration on the Information Bottleneck plane against its capacity H(T).

y axis
0123450.000.250.500.751.00H(T): capacity of the coarse-graining (bits)I(T;Y): transition information (bits)up and left dominates
variantI(T;Y)AMI
DMP-Qpos*0.8950.230
DMC-Intention*0.8730.212
KPMS-KP0.5250.159
DMP-Intention*0.5100.172
DMC-Qpos*0.4740.166
DMP-Ekp*0.3630.163
Mapper-Qpos0.3220.121
KPMS-Int0.2860.108
Mapper-EKP0.2760.110
DMC-Ekp*0.2340.123
Mapper-Int0.1550.080

Baselines: mean over 3 seeds. DM decomposition is deterministic. H(Y) = 2.67 bits bounds I(T;Y).

Figure 8. The Information Bottleneck plane.

The paper’s Figure 2 and Table 8, interactive. One dot per (method, seed); hover for values, click a family to hide it, switch the y axis to the cardinality-corrected AMI. DM decomposition is deterministic, so DM variants have one dot each.

DM cycles and paths in intention and qpos reach the top of the plane. The qpos DM-path carries the most transition information of any method, 0.90 bits, against 0.53 for the strongest baseline, Keypoint-MoSeq on keypoints, with intention DM-cycles close behind at 0.87. Correcting for label cardinality the lead is AMI 0.23 against 0.16; Mapper and Keypoint-MoSeq on intention stay at AMI ≤ 0.12. The ordering within DM, intention ≈ qpos > ekp, is the one the drift argument predicts.

Alignment, coherence and warping

Do the discovered clusters match the annotated transitions frame by frame? Each family is represented by its highest-information variant. Both DM variants align closest to the annotated boundaries on average, 0.04 to 0.06 below Keypoint-MoSeq and 0.01 to 0.03 below Mapper in mean DTW, but per row the picture is less one-sided: the DM columns win 6 of the 11 DTW rows. The warping decomposition separates the cases. The DM columns have the lower non-diagonal fraction on 9 of 11 rows and the higher sync ratio on 9 of 11; Keypoint-MoSeq leads on neither. Where Keypoint-MoSeq approaches our DTW, it does so by stretching a primitive onto the annotated window.

metric (lower is better)

Frame-level alignment to the annotated window under a monotone time correspondence.

TransitionDMC-Int*DMP-Qpos*Mapper-QposKPMS-KP
Immobile → Rear [L]—0.854—0.831
Rear → Immobile [L]—0.855†0.8790.855
Immobile → Rear [R]—0.880†—0.848†
Immobile → Rear [S]0.9340.8920.8950.967
Rear → Immobile [S]0.863†—0.8560.877
Immobile → Turn [L]0.8090.7880.8010.770
Turn → Immobile [L]0.7400.7210.7900.849†
Immobile → Turn [R]0.7590.756†0.715†0.747
Immobile → Walk [L]0.759†———
Rear → Turn [L]0.8260.7780.8180.918
Turn → Walk [L]0.643———

One variant per family, chosen by transition information. Bold = best in the row among populated cells. — = not discovered. † = one seed only. Direction tags: [L] left, [R] right, [S] straight. *Ours.

Figure 9. Alignment by directed transition.

Paper Table 1. All metrics are computed in the same root-stripped joint-angle space. Per-variant grids for all twelve variants are in the paper’s appendix.

On KID, which ignores frame order and cannot warp, intention DM-cycles trail Keypoint-MoSeq by 13%, and we claim no distributional advantage. KID rewards breadth: a Keypoint-MoSeq cluster holds a median of about 14,000 frames against 3,800 for qpos DM-paths, sampling the window more fully while being markedly less coherent. Coverage is comparable: intention DM-cycles reach 10 of the 16 discovered directed transitions, against 11 for Mapper on intention and 9 for Keypoint-MoSeq on keypoints. The families are not nested. Four transitions, among them Walk → Turn [L] and Walk → Immobile [S], are found only by DM primitives, and Rear → Immobile [R] only by Mapper.

Three controls rule out simpler explanations. None of 110 baseline variants, sweeping Keypoint-MoSeq stickiness and state count and Mapper lens and interval count, improves on the tuned baselines. A full-clip control that skips the decomposition does worse than both DM variants in alignment (0.888) and coherence (0.797), so the decomposition, not the Wasserstein clustering, recovers the transitions. And DTW varies by at most 0.03 across the k-NN parameter in every space.

Citation

@inproceedings{bian2026shaped,
title = {Shaped by What's Missing: Topological Invariance Simplification
Discovers Behavioral Transitions},
author = {Bian, Kaiwen and Leonardis, Eric J. and Yang, Yuanjia and
Zhang, Charles and Wang, Yusu and Pereira, Talmo D.},
booktitle = {Proceedings of the NeurIPS Workshop on Symmetry and Geometry
in Neural Representations (NeurReps)},
series = {Proceedings of Machine Learning Research},
year = {2026}
}

The trajectories come from MIMIC-MJX. The discrete-Morse reconstruction uses PCD-Graph-Recon-DM (Magee and Wang, 2022). Baselines use Keypoint-MoSeq and tda-mapper.