Tracer Developers Quickstart Results Whitepaper GitHub

Punches in.
Scored events out.

Two independent sensing tiers — wrist IMU and phone camera — into one on-device cascade. Every constant and number on this page is taken from the code or regenerated from data in the repository, never from memory.

Wrist, 3-class
0.693
all 238 captured punches, held out by stance
Wrist, straight vs arc
0.887
hook and uppercut merged
Camera, mean-class recall
0.614
BoxingVI, held out by subject

The wrist tier takes two Seeed XIAO nRF52840 Sense boards over BLE into a peak detector and a feature classifier fitted on our own straps. The camera tier takes the phone camera into RTMPose keypoints, a wrist-velocity spotter, and an ST-GCN++ trunk. Neither requires the other.

They are built together because they fail differently. A single camera sees trajectory but loses depth, so it cannot separate vertical drive from horizontal arc. A wrist sensor measures that drive directly but has no idea where the hand is in space. Both tiers reach their limit on the same boundary — hook against uppercut — which is the argument for fusing them.

Quickstart

Run the tests

# Python: converters, spotter parity, CoreML parity
.venv/bin/python -m pytest tests/ -q

# Swift: spotter, calibration solver, gate, round clock
cd ios && xcodebuild test -scheme FierceReflex \
  -destination 'platform=iOS Simulator,name=iPhone 17 Pro'

Fit the wrist classifier from captured clips

.venv/bin/python scripts/fit_own_rig_shape.py
# prints the confusion matrix, writes
# ios/FierceReflex/OwnRigShapeClassifier.swift

Build the camera training set and train

.venv/bin/python scripts/convert/boxingvi_public_to_pkl.py
./scripts/train_pose.sh        # picks up boxingvi_pub_3c_* if present

Install to a device

cd ios && xcodegen generate
xcodebuild -scheme FierceReflex -destination 'id=<UDID>' \
  -allowProvisioningUpdates build
xcrun devicectl device install app --device <UDID> <built>.app

Repo layout

PathWhat lives there
ios/FierceReflex/The app. Spotters, classifiers, calibration, capture, review UI.
inference/Python reference implementations the Swift mirrors.
training/ST-GCN++ and IMU model definitions and training loops.
scripts/convert/Dataset converters. Each is the only path to its PKL.
posedata/Corpora. imu_own_3c/ is ours; pkl/ is converted BoxingVI.
firmware/xiao_imu/340-line BLE streamer. No model runs on the board.
docs/whitepaper/Typst source and figures for the technical paper.

The cascade

Both tiers share a shape: detect → window → classify → gate → resolve hand. Detection decides when a punch happened, windowing cuts a fixed clip around the peak, the classifier names the shape, the gate decides whether the answer is trustworthy, and the hand is resolved without a model at all — the sensor reports which wrist fired, and stance is a setting, so jab and cross separate from straight by definition.

Design rule

Nothing downstream of the gate ever sees raw signal. Coaching cues and the round-break agent consume structured scores only, so any spoken cue traces back to the score that triggered it.

Wrist tier

Firmware streams the onboard LSM6DS3TR-C at a declared 104 Hz in batches of four over BLE. All detection, calibration and classification happen on the phone; the board carries no model.

Detection

IMUPunchSpotter arms peak search at 3.0 g and emits only when the confirmed peak reaches 4.0 g. The two are separate on purpose: arming low lets search begin on the rise so the reported crest is the real one, while emitting high keeps non-punches out.

ThresholdNon-punch rejectedReal punches lost
3.0 g0.0%0.0%
4.0 g shipped86.1%5.9%
5.5 g100.0%12.6%

A Schmitt-style re-arm at rearmRatio 0.5 stops a punch double-firing on its return stroke. Measured on the public clips, re-arm took single-detection rate from 70.3% to 90.3% overall and from 51.3% to 89.9% on uppercuts, which were worst because a drive up and a return down are both large.

Measurement trap

The original 3.0 g guard rejected zero non-punch clips, and could not have rejected any: every clip in the set was recorded by that detector at 3.0 g, so by construction nothing in it could fail its own test. An offline dataset cannot contain the false positives of the gate that collected it. Found by plotting the distribution, not by validation accuracy.

Features

Seven features, and every one is invariant under rotation of the sensor frame: peakA, meanA, sdA, peakW, meanW, wOverA, peakPos. Verified numerically — 274 clips against 25 random rotations, zero predicted labels changed.

Eight axis-reading features were removed to get there. They tied the model to how each board sits, and a model fitted before calibration existed and then fed rotated input returned hook for 62% of all punches while validating at 0.74. A rotated input is not noise: the model does not report that it has not seen something, it extrapolates confidently into whichever class holds the largest region of feature space.

Why sixty parameters

ModelParametersAccuracy
Feature classifier shipped600.757
1D CNN~194k0.458
Kautz-style CNN~190k0.442
InceptionTime~400k0.383
DeepConvLSTM~300k0.350

Every network reaches 1.000 training accuracy within three epochs. Capacity has to follow the labelled corpus, not precede it.

Calibration

Six static holds — logo up/down, USB up/down, left/right edge up — give three opposed pairs. Each canonical axis is half the difference of its pair, which cancels any constant accelerometer bias that a single-sided reading would bake into the rotation. Gram-Schmidt then forces an exact orthonormal basis.

This is deliberately not a general SVD/Kabsch fit: the six poses are already an orthogonal triad up to noise, and a hand-rolled 3×3 SVD is a well-known way to get this subtly wrong. A reflection (determinant −1) is rejected outright, because no mounting can mirror a board.

Calibration is applied in IMUSession.ingest, before the spotter, so one frame holds for the detector, the metrics, the clips written to disk and the classifier alike. Because the shipped features are rotation-invariant, recalibrating can no longer affect classification — it still matters for metrics and for left/right comparability.

Camera tier

RTMPose-s at 192×256 produces COCO-17 keypoints. WristVelocitySpotter triggers on shoulder-width-normalised, body-relative wrist speed with a two-sided local maximum test, a confidence gate at KEYPOINT_VALID_MIN 0.3 on both wrist and hip, and the same 0.5 re-arm ratio as the wrist tier.

The window is 48 frames centred on the peak — 24 before and 24 after. Training clips were stretched to that length by mmaction's UniformSampleFrames, so the runtime uses the evaluation branch of the same sampler (bin centres) rather than feeding 48 real frames to a model trained on roughly 0.4 s of stretched motion.

Prediction gate

The gate turns a softmax into one of three honest answers: a single class, a merged pair, or an abstention.

ParameterValueMeaning
pairConf0.65Mass a pair needs before a merged label is honest
arcMargin0.30Margin hook needs over uppercut to be called alone
straightMargin0.22Margin for the straight family

HARD_PAIRS contains only hook/uppercut. When the two are close but their combined mass is high, the gate emits hook/uppercut rather than guessing — not confident, but not uninformative.

Measured

Wrist results are held out by stance: trained on orthodox, tested on southpaw, and the reverse. With one athlete a random per-clip split would shuffle the same repetitions across train and test and report a number that means nothing.

EvaluationnAccuracy
Wrist, 3-class, all captured punches2380.693
Wrist, 3-class, above the 4.0 g threshold2240.754
Wrist, straight vs arc2380.887
Camera, top-1 (held out by subject)11200.663
Camera, mean-class recall11200.614

The 4.0 g row is what the live system does, since the spotter never emits a weaker window. The 238-clip row includes punches the product would decline to classify, and is reported first because it describes the wider set of inputs.

Whitepaper

A ten-page technical paper covering the same system for a non-engineering reader: the sensing requirement measured rather than quoted, the custom glove board specification, the rubric scoring design, the annotation plan, and manufacturing cost.

It is built from this repository, not written alongside it. Every figure is regenerated from data on disk by scripts/whitepaper_figures.py and scripts/whitepaper_diagrams.py, so a retrain regenerates the paper rather than invalidating it. The one exception is the camera confusion matrix, quoted from MODELS.md and marked as such in the figure source.

Known limits

  • One athlete. 274 clips. Cross-subject accuracy is not yet measured, and wrist data has already proven able to fail completely between mountings — models fitted on fist-mounted public data returned hook for essentially every input on straps.
  • Hook against uppercut is the whole residual error in both tiers, and it is a sensor-placement limit. Published ablation attributes about 8 points on that exact boundary to trunk rotation, which the wrists cannot observe.
  • Confidence is not well calibrated. Wrong predictions carry median confidence 0.668 against 0.866 for correct ones. Abstaining below 0.7 reaches 0.795 while discarding a third of all punches, so no cheap threshold helps.
  • The radio delivers 33 Hz against a declared 104 Hz, unimodally — 84.6% of intervals fall between 25 and 35 ms. Consistent between capture and live, so nothing is mismatched by it, but it bounds how finely punch shape can be resolved. Addressable in BLE batching with no new data.

Corpora

CorpusClipsSplit
posedata/imu_own_3c/ — our straps274by stance (154 orthodox / 120 southpaw)
posedata/pkl/boxingvi_pub_3c_*5,451by subject, V4+V5 held out (4,293 / 1,158)
On the public release

The BoxingVI README advertises 6,915 clips from 18 athletes across 20 videos. The distributed archive contains V1–V10, 5,570 annotated rows, ten subjects. The release also carries two incompatible layouts — nine videos pre-segmented where row i is clip i, one continuous where the start/end columns are real indices. Our converter detects the layout per video, finds the class column by content rather than position, and skips any video whose clip count disagrees with its label count.

Collecting

The in-app collection section drives a self-ticking checklist: six real punches per stance (jab, cross, lead/rear hook, lead/rear uppercut), footwork negatives, and the six-pose calibration. Each clip is written with its full context — mount, stance, surface, intensity, hand role — because a training run that cannot filter by those cannot answer the questions the protocol was designed to answer.

.venv/bin/python scripts/convert/ios_punches_to_imu_clips.py \
  --input punches_*.jsonl --labels labels_*.jsonl --subject V11

A measured rule: accuracy saturates at roughly 85 clips per stance and is flat thereafter. More punches from the same athlete add nothing; each new athlete opens ground none of the existing data covers.

Review loop

While recording, the live feed becomes a review queue. Each punch shows what the model called it, with every class beside it at equal weight and nothing pre-selected.

Why a prediction is not a label

Fitting a model to its own output teaches it to repeat its errors, and the errors sit exactly where it is confident and wrong. The only new information in the loop is a human saying what the punch was.

  • Confirmed — cheap to give, and the confirmed/corrected ratio is live accuracy, with no held-out split to argue about.
  • Corrected — rarer, and worth more per clip than anything else collectable, because it falls exactly on the misplaced boundary.
  • Not a punch — the only measure of the detector's live precision, which no offline set can supply.

Verdicts append to labels_<session>.jsonl; last line per clip wins. Re-labelling appends rather than edits, so a clip relabelled three times is visibly ambiguous. An unreviewed punch is stored as neither and is never silently promoted.

Training

# camera trunk — prefers boxingvi_pub_3c_* when present
./scripts/train_pose.sh

# wrist classifier — deterministic, regenerates the Swift file
.venv/bin/python scripts/fit_own_rig_shape.py --gate 4.0

The wrist fit is deterministic (zero init, full batch, no shuffling), so the generated Swift is byte-identical across runs and a diff means the data changed.

CoreML export

FLOAT16 is the silent default

coremltools defaults to FLOAT16 and will quietly halve your precision. Pin compute_precision=ct.precision.FLOAT32. On the IMU model this moved parity error from 1.02e−2 to 2e−6.

The export script takes --classes and refuses on a mismatch, after an earlier version sliced a hardcoded 4-class list to length and silently wrote ['jab','cross','hook'] onto a 3-class trunk.

Tests

Eight Swift test files and eight Python ones. The ones worth knowing about: spotter parity between the Swift and Python implementations, the calibration solver against known rotations, CoreML parity against the PyTorch source, and gate behaviour on merged pairs.

Constants

ConstantValueWhere
minPunchG3.0 gIMU peak search arms
minEmitG4.0 gIMU candidate emitted
rearmRatio0.5Both spotters
refractory250 msIMU; a 1–2 lands 300–400 ms apart
clipDuration48/52 s923 ms, the training clip length
leadFraction0.45Fraction of the clip before the peak
KEYPOINT_VALID_MIN0.3Pose confidence gate
refractoryFrames15Pose, ~0.5 s at 30 fps
CLIP_LEN48Pose window, peak-centred

Scripts

ScriptPurpose
scripts/fit_own_rig_shape.pyFit the wrist classifier, regenerate the Swift weights
scripts/convert/boxingvi_public_to_pkl.pyAuthors' release → training PKLs
scripts/convert/ios_punches_to_imu_clips.pyCaptured sessions + review labels → clips
scripts/convert/collapse_4c_to_3c.pyCollapse jab+cross into straight
scripts/train_pose.shTrain the camera trunk
scripts/flash_imu.sh./scripts/flash_imu.sh left|right
scripts/whitepaper_figures.pyRegenerate every whitepaper figure from data
scripts/whitepaper_diagrams.pySchematic figures for the whitepaper

Phase-2 archive

Several documents in docs/ describe a different architecture — YOLO detection, OC-SORT tracking, PoseC3D — explored before the current cascade. They are deliberately retained as the design paper trail for a multi-fighter phase and are not a description of what ships today.

Affected: docs/ARCHITECTURE.md, docs/TECHNICAL_OVERVIEW.md, docs/API_REFERENCE.md, docs/PROJECT_STATUS.md. For current behaviour, this site and MODELS.md are authoritative.

Tracer — developer documentation. Numbers regenerate from scripts/whitepaper_figures.py and scripts/fit_own_rig_shape.py.