
Perception · Kinematics
Animal behaviour analysis
A 48-hour scientific prototype that turns behavioural video into pose-derived movement features and a first supervised classification baseline.
CENTURI Living Systems
- Duration
- 2 days
- Scope
- 39-keypoint tracking
- Contribution
- Tracking & modeling
From manual behavioural annotation to a traceable pose-and-motion pipeline for studying mouse-pup behaviour.
The challenge
Behaviour had to become measurable before it could become comparable.
AMBIA — Automated Mouse Behavior Imaging Annotation — was a 48-hour prototype developed during the CENTURI Living Systems hackathon for quantitative biology. The scientific setting couples behavioural observation with in vivo calcium imaging: neural signals are only interpretable when the animal’s actions can be timed and labelled with comparable care.
Manual annotation of mouse-pup behaviour from 2D video is slow, labour-intensive, and difficult to keep consistent across sessions. Camera placement, lighting, subject size, and developmental stage all vary, so labels written by hand do not transfer cleanly from one recording to the next.
The prototype did not claim to replace human annotation. It asked whether a traceable computational path — from heterogeneous video and interval labels to pose-derived motion features and a first supervised baseline — could make behaviour more inspectable under those constraints.
Working with a variable corpus
Heterogeneity was part of the problem, not an afterthought.
The working corpus comprised 36 sessions: 14 annotated and 22 unannotated. Recordings differed in animal age, camera configuration, spatial resolution, frame rate, and genetic lineage. That variability is exactly what limits naive transfer of a behaviour model trained on one acquisition setup.
The fourteen annotated sessions were standardised into a shared label space. Together they account for 2,832 standardised behavioural intervals after upgrade. That inventory documents the annotated material available to the prototype; it is not a multi-session validation claim.
Making behaviour computable
Labels had to become frame-aligned before pose could become features.
Observed annotation vocabularies were heterogeneous. Across the annotated material, 24 observed labels were mapped onto 10 canonical classes, so intervals from different sessions could share a common taxonomy before any modelling step.
The computational chain then follows a deliberate sequence:
Each stage produces an intermediate artefact that can be inspected independently. On the full annotated corpus, standardisation reaches frame-ready labels. The pose → feature → classification branch is demonstrated on session em001/14, the only session for which SuperAnimal pose output was available in the surviving project artefacts.

Pixel-level motion energy provides a first scalar reading of activity in the frame. Pose-derived features later offer a structured alternative once keypoints are available.
From pose to motion features
Quality selection came before feature engineering.
Pose coordinates for em001/14 come from a SuperAnimal quadruped model retained at 39 keypoints (position and likelihood per part). The team did not train DeepLabCut for this prototype; the available SuperAnimal pose product was consumed as an upstream tracking output.
Not every keypoint is equally reliable under a side or overhead view of a pup. A quality screen kept 9 anatomical points on likelihood, coverage, and jump-rate criteria, and set aside the rest for review or discard.
From those retained parts, the pipeline built 94 kinematic features per frame — local displacements, speeds, accelerations, rolling summaries, and a small set of global body-motion aggregates.
Verified coordinate rules:
- likelihood below 0.50 treated as missing;
- linear interpolation of gaps up to 5 consecutive frames;
- rolling statistics over a 30-frame window.
Longer gaps remain missing rather than being filled with implausible motion.
What the baseline showed
A first classifier, on one session, under a temporal split.
On em001/14, a Random Forest baseline was trained on the pose-derived supervised dataset. Evaluation used a temporal 80/20 train–test split so that later frames were held out rather than randomly interleaved. The reported classes were Resting, Complexing, and Mouthing.
On the held-out temporal partition the baseline reached:
- accuracy 0.713
- macro F1 0.291
- weighted F1 0.617
Per-class F1 scores were sharply asymmetric:
- Resting — 0.833
- Complexing — 0.040
- Mouthing — 0.000


These figures document one reconstructable baseline run. They do not stand in for a multi-session benchmark.
Reading beyond accuracy
A high overall score can still miss the behaviours that matter least in the majority class.
Accuracy near 0.71 is dominated by Resting, which occupies most of the labelled frames in the test partition. Complexing is barely recovered; Mouthing is not recognised at all in this evaluation. Macro F1 (0.291) makes that asymmetry explicit in a way that accuracy and weighted F1 do not.
The surviving artefacts do not isolate a single cause. Class imbalance, temporal structure, pose noise after quality filtering, and the limits of a single-session feature space can all contribute. The honest reading is diagnostic: the baseline privileges the majority state and fails on rarer behaviours under this protocol.
What we learned
Classification was the probe. Representation was the real problem.
The workable contribution of the prototype is not a claim that mouse-pup behaviour is solved. It is that annotations can be standardised, that pose quality can be screened before modelling, and that kinematic features can be evaluated under an explicit temporal protocol — while still exposing how quickly performance collapses on rare classes.
The central difficulty was not only to classify a movement, but to produce a behavioural representation that remains stable under heterogeneous acquisitions and imbalanced labels.
Where we would go next
A second iteration would treat evaluation design as carefully as feature design.
Validate by session or animal
Hold out entire sessions or animals rather than only later frames within one recording, so that transfer across acquisitions can be measured directly.
Rebalance and report by class
Pair majority-aware training with class-wise metrics as primary readouts, so rare behaviours cannot disappear behind overall accuracy.
Protocolise multi-camera setups
Make camera count, viewpoint, and calibration part of the experimental design whenever dual-camera sessions are used.
Stratify by age and lineage
Evaluate separately across developmental ages and genetic backgrounds instead of pooling them into a single score.
These steps are proposed extensions. They were not completed within the 48-hour prototype.
Techniques & technologies
Techniques
- Annotation standardization
- Frame-level labelling
- Pose-quality assessment
- Kinematic feature engineering
- Temporal evaluation
- Class-imbalance diagnosis
Technologies
- Python
- NumPy
- Pandas
- OpenCV
- DeepLabCut
- scikit-learn