Projects
Macro detail of cinema camera follow-focus gear under warm and cool studio lighting

Perception · Kinematics

Animal behaviour analysis

A 48-hour scientific prototype that turns behavioural video into pose-derived movement features and a first supervised classification baseline.

CENTURI Living Systems

Duration
2 days
Scope
39-keypoint tracking
Contribution
Tracking & modeling

From manual behavioural annotation to a traceable pose-and-motion pipeline for studying mouse-pup behaviour.

The challenge

Behaviour had to become measurable before it could become comparable.

AMBIA — Automated Mouse Behavior Imaging Annotation — was a 48-hour prototype developed during the CENTURI Living Systems hackathon for quantitative biology. The scientific setting couples behavioural observation with in vivo calcium imaging: neural signals are only interpretable when the animal’s actions can be timed and labelled with comparable care.

Manual annotation of mouse-pup behaviour from 2D video is slow, labour-intensive, and difficult to keep consistent across sessions. Camera placement, lighting, subject size, and developmental stage all vary, so labels written by hand do not transfer cleanly from one recording to the next.

The prototype did not claim to replace human annotation. It asked whether a traceable computational path — from heterogeneous video and interval labels to pose-derived motion features and a first supervised baseline — could make behaviour more inspectable under those constraints.

Thirty-second excerpt from a longer behavioural recording session — browser-ready H.264 conversion of the source AVI.

Working with a variable corpus

Heterogeneity was part of the problem, not an afterthought.

The working corpus comprised 36 sessions: 14 annotated and 22 unannotated. Recordings differed in animal age, camera configuration, spatial resolution, frame rate, and genetic lineage. That variability is exactly what limits naive transfer of a behaviour model trained on one acquisition setup.

The fourteen annotated sessions were standardised into a shared label space. Together they account for 2,832 standardised behavioural intervals after upgrade. That inventory documents the annotated material available to the prototype; it is not a multi-session validation claim.


Making behaviour computable

Labels had to become frame-aligned before pose could become features.

Observed annotation vocabularies were heterogeneous. Across the annotated material, 24 observed labels were mapped onto 10 canonical classes, so intervals from different sessions could share a common taxonomy before any modelling step.

Annotation interface used to inspect behavioural sequences and register time intervals.

The computational chain then follows a deliberate sequence:

Annotation-to-baseline pipeline from labels through pose features to classifier

Each stage produces an intermediate artefact that can be inspected independently. On the full annotated corpus, standardisation reaches frame-ready labels. The pose → feature → classification branch is demonstrated on session em001/14, the only session for which SuperAnimal pose output was available in the surviving project artefacts.

Motion-energy time series plotted over successive video frames
Motion-energy series over 24,764 frames — a pixel-based activity trace aligned with the recording timeline, not a classification result.

Pixel-level motion energy provides a first scalar reading of activity in the frame. Pose-derived features later offer a structured alternative once keypoints are available.

From pose to motion features

Quality selection came before feature engineering.

Pose coordinates for em001/14 come from a SuperAnimal quadruped model retained at 39 keypoints (position and likelihood per part). The team did not train DeepLabCut for this prototype; the available SuperAnimal pose product was consumed as an upstream tracking output.

Pose-tracking visualisation used to inspect anatomical motion before feature engineering.

Not every keypoint is equally reliable under a side or overhead view of a pup. A quality screen kept 9 anatomical points on likelihood, coverage, and jump-rate criteria, and set aside the rest for review or discard.

From those retained parts, the pipeline built 94 kinematic features per frame — local displacements, speeds, accelerations, rolling summaries, and a small set of global body-motion aggregates.

Verified coordinate rules:

  • likelihood below 0.50 treated as missing;
  • linear interpolation of gaps up to 5 consecutive frames;
  • rolling statistics over a 30-frame window.

Longer gaps remain missing rather than being filled with implausible motion.


What the baseline showed

A first classifier, on one session, under a temporal split.

On em001/14, a Random Forest baseline was trained on the pose-derived supervised dataset. Evaluation used a temporal 80/20 train–test split so that later frames were held out rather than randomly interleaved. The reported classes were Resting, Complexing, and Mouthing.

On the held-out temporal partition the baseline reached:

  • accuracy 0.713
  • macro F1 0.291
  • weighted F1 0.617

Per-class F1 scores were sharply asymmetric:

  • Resting — 0.833
  • Complexing — 0.040
  • Mouthing — 0.000
Confusion matrix for the Random Forest baseline on session em001/14
Confusion matrix for the Random Forest baseline — session em001/14 only.
Top twenty feature importances from the Random Forest baseline on session em001/14
Top-20 feature importances from the Random Forest baseline — session em001/14 only.

These figures document one reconstructable baseline run. They do not stand in for a multi-session benchmark.

Reading beyond accuracy

A high overall score can still miss the behaviours that matter least in the majority class.

Accuracy near 0.71 is dominated by Resting, which occupies most of the labelled frames in the test partition. Complexing is barely recovered; Mouthing is not recognised at all in this evaluation. Macro F1 (0.291) makes that asymmetry explicit in a way that accuracy and weighted F1 do not.

The surviving artefacts do not isolate a single cause. Class imbalance, temporal structure, pose noise after quality filtering, and the limits of a single-session feature space can all contribute. The honest reading is diagnostic: the baseline privileges the majority state and fails on rarer behaviours under this protocol.


What we learned

Classification was the probe. Representation was the real problem.

The workable contribution of the prototype is not a claim that mouse-pup behaviour is solved. It is that annotations can be standardised, that pose quality can be screened before modelling, and that kinematic features can be evaluated under an explicit temporal protocol — while still exposing how quickly performance collapses on rare classes.

The central difficulty was not only to classify a movement, but to produce a behavioural representation that remains stable under heterogeneous acquisitions and imbalanced labels.

Where we would go next

A second iteration would treat evaluation design as carefully as feature design.

Validate by session or animal

Hold out entire sessions or animals rather than only later frames within one recording, so that transfer across acquisitions can be measured directly.

Rebalance and report by class

Pair majority-aware training with class-wise metrics as primary readouts, so rare behaviours cannot disappear behind overall accuracy.

Protocolise multi-camera setups

Make camera count, viewpoint, and calibration part of the experimental design whenever dual-camera sessions are used.

Stratify by age and lineage

Evaluate separately across developmental ages and genetic backgrounds instead of pooling them into a single score.

These steps are proposed extensions. They were not completed within the 48-hour prototype.


Techniques & technologies

Techniques

  • Annotation standardization
  • Frame-level labelling
  • Pose-quality assessment
  • Kinematic feature engineering
  • Temporal evaluation
  • Class-imbalance diagnosis

Technologies


Resources