Punches in.
Scored events out.
Two independent sensing tiers — wrist IMU and phone camera — into one on-device cascade. Every constant and number on this page is taken from the code or regenerated from data in the repository, never from memory.
The wrist tier takes two Seeed XIAO nRF52840 Sense boards over BLE into a peak detector and a feature classifier fitted on our own straps. The camera tier takes the phone camera into RTMPose keypoints, a wrist-velocity spotter, and an ST-GCN++ trunk. Neither requires the other.
They are built together because they fail differently. A single camera sees trajectory but loses depth, so it cannot separate vertical drive from horizontal arc. A wrist sensor measures that drive directly but has no idea where the hand is in space. Both tiers reach their limit on the same boundary — hook against uppercut — which is the argument for fusing them.
Quickstart
Run the tests
# Python: converters, spotter parity, CoreML parity
.venv/bin/python -m pytest tests/ -q
# Swift: spotter, calibration solver, gate, round clock
cd ios && xcodebuild test -scheme FierceReflex \
-destination 'platform=iOS Simulator,name=iPhone 17 Pro'
Fit the wrist classifier from captured clips
.venv/bin/python scripts/fit_own_rig_shape.py
# prints the confusion matrix, writes
# ios/FierceReflex/OwnRigShapeClassifier.swift
Build the camera training set and train
.venv/bin/python scripts/convert/boxingvi_public_to_pkl.py
./scripts/train_pose.sh # picks up boxingvi_pub_3c_* if present
Install to a device
cd ios && xcodegen generate
xcodebuild -scheme FierceReflex -destination 'id=<UDID>' \
-allowProvisioningUpdates build
xcrun devicectl device install app --device <UDID> <built>.app
Repo layout
| Path | What lives there |
|---|---|
ios/FierceReflex/ | The app. Spotters, classifiers, calibration, capture, review UI. |
inference/ | Python reference implementations the Swift mirrors. |
training/ | ST-GCN++ and IMU model definitions and training loops. |
scripts/convert/ | Dataset converters. Each is the only path to its PKL. |
posedata/ | Corpora. imu_own_3c/ is ours; pkl/ is converted BoxingVI. |
firmware/xiao_imu/ | 340-line BLE streamer. No model runs on the board. |
docs/whitepaper/ | Typst source and figures for the technical paper. |
The cascade
Both tiers share a shape: detect → window → classify → gate
→ resolve hand. Detection decides when a punch happened, windowing cuts
a fixed clip around the peak, the classifier names the shape, the gate decides whether
the answer is trustworthy, and the hand is resolved without a model at all — the
sensor reports which wrist fired, and stance is a setting, so jab and cross separate
from straight by definition.
Nothing downstream of the gate ever sees raw signal. Coaching cues and the round-break agent consume structured scores only, so any spoken cue traces back to the score that triggered it.
Wrist tier
Firmware streams the onboard LSM6DS3TR-C at a declared 104 Hz in batches of four over BLE. All detection, calibration and classification happen on the phone; the board carries no model.
Detection
IMUPunchSpotter arms peak search at 3.0 g and emits only
when the confirmed peak reaches 4.0 g. The two are separate on purpose:
arming low lets search begin on the rise so the reported crest is the real one, while
emitting high keeps non-punches out.
| Threshold | Non-punch rejected | Real punches lost |
|---|---|---|
| 3.0 g | 0.0% | 0.0% |
| 4.0 g shipped | 86.1% | 5.9% |
| 5.5 g | 100.0% | 12.6% |
A Schmitt-style re-arm at rearmRatio 0.5 stops a punch double-firing on
its return stroke. Measured on the public clips, re-arm took single-detection rate from
70.3% to 90.3% overall and from 51.3% to 89.9% on uppercuts, which were worst because a
drive up and a return down are both large.
The original 3.0 g guard rejected zero non-punch clips, and could not have rejected any: every clip in the set was recorded by that detector at 3.0 g, so by construction nothing in it could fail its own test. An offline dataset cannot contain the false positives of the gate that collected it. Found by plotting the distribution, not by validation accuracy.
Features
Seven features, and every one is invariant under rotation of the sensor
frame: peakA, meanA, sdA,
peakW, meanW, wOverA, peakPos.
Verified numerically — 274 clips against 25 random rotations, zero predicted
labels changed.
Eight axis-reading features were removed to get there. They tied the model to how
each board sits, and a model fitted before calibration existed and then fed rotated
input returned hook for 62% of all punches while validating at 0.74. A
rotated input is not noise: the model does not report that it has not seen something,
it extrapolates confidently into whichever class holds the largest region of feature
space.
Why sixty parameters
| Model | Parameters | Accuracy |
|---|---|---|
| Feature classifier shipped | 60 | 0.757 |
| 1D CNN | ~194k | 0.458 |
| Kautz-style CNN | ~190k | 0.442 |
| InceptionTime | ~400k | 0.383 |
| DeepConvLSTM | ~300k | 0.350 |
Every network reaches 1.000 training accuracy within three epochs. Capacity has to follow the labelled corpus, not precede it.
Calibration
Six static holds — logo up/down, USB up/down, left/right edge up — give three opposed pairs. Each canonical axis is half the difference of its pair, which cancels any constant accelerometer bias that a single-sided reading would bake into the rotation. Gram-Schmidt then forces an exact orthonormal basis.
This is deliberately not a general SVD/Kabsch fit: the six poses are already an orthogonal triad up to noise, and a hand-rolled 3×3 SVD is a well-known way to get this subtly wrong. A reflection (determinant −1) is rejected outright, because no mounting can mirror a board.
Calibration is applied in IMUSession.ingest, before the spotter, so one
frame holds for the detector, the metrics, the clips written to disk and the classifier
alike. Because the shipped features are rotation-invariant, recalibrating can no longer
affect classification — it still matters for metrics and for left/right
comparability.
Camera tier
RTMPose-s at 192×256 produces COCO-17 keypoints. WristVelocitySpotter
triggers on shoulder-width-normalised, body-relative wrist speed with a two-sided local
maximum test, a confidence gate at KEYPOINT_VALID_MIN 0.3 on both wrist and
hip, and the same 0.5 re-arm ratio as the wrist tier.
The window is 48 frames centred on the peak — 24 before and 24
after. Training clips were stretched to that length by mmaction's
UniformSampleFrames, so the runtime uses the evaluation branch of the same
sampler (bin centres) rather than feeding 48 real frames to a model trained on roughly
0.4 s of stretched motion.
Prediction gate
The gate turns a softmax into one of three honest answers: a single class, a merged pair, or an abstention.
| Parameter | Value | Meaning |
|---|---|---|
pairConf | 0.65 | Mass a pair needs before a merged label is honest |
arcMargin | 0.30 | Margin hook needs over uppercut to be called alone |
straightMargin | 0.22 | Margin for the straight family |
HARD_PAIRS contains only hook/uppercut. When the two are
close but their combined mass is high, the gate emits hook/uppercut rather
than guessing — not confident, but not uninformative.
Measured
Wrist results are held out by stance: trained on orthodox, tested on southpaw, and the reverse. With one athlete a random per-clip split would shuffle the same repetitions across train and test and report a number that means nothing.
| Evaluation | n | Accuracy |
|---|---|---|
| Wrist, 3-class, all captured punches | 238 | 0.693 |
| Wrist, 3-class, above the 4.0 g threshold | 224 | 0.754 |
| Wrist, straight vs arc | 238 | 0.887 |
| Camera, top-1 (held out by subject) | 1120 | 0.663 |
| Camera, mean-class recall | 1120 | 0.614 |
The 4.0 g row is what the live system does, since the spotter never emits a weaker window. The 238-clip row includes punches the product would decline to classify, and is reported first because it describes the wider set of inputs.
Whitepaper
A ten-page technical paper covering the same system for a non-engineering reader: the sensing requirement measured rather than quoted, the custom glove board specification, the rubric scoring design, the annotation plan, and manufacturing cost.
It is built from this repository, not written alongside it. Every figure is
regenerated from data on disk by scripts/whitepaper_figures.py and
scripts/whitepaper_diagrams.py, so a retrain regenerates the paper
rather than invalidating it. The one exception is the camera confusion matrix,
quoted from MODELS.md and marked as such in the figure source.
Known limits
- One athlete. 274 clips. Cross-subject accuracy is not yet
measured, and wrist data has already proven able to fail completely between mountings
— models fitted on fist-mounted public data returned
hookfor essentially every input on straps. - Hook against uppercut is the whole residual error in both tiers, and it is a sensor-placement limit. Published ablation attributes about 8 points on that exact boundary to trunk rotation, which the wrists cannot observe.
- Confidence is not well calibrated. Wrong predictions carry median confidence 0.668 against 0.866 for correct ones. Abstaining below 0.7 reaches 0.795 while discarding a third of all punches, so no cheap threshold helps.
- The radio delivers 33 Hz against a declared 104 Hz, unimodally — 84.6% of intervals fall between 25 and 35 ms. Consistent between capture and live, so nothing is mismatched by it, but it bounds how finely punch shape can be resolved. Addressable in BLE batching with no new data.
Corpora
| Corpus | Clips | Split |
|---|---|---|
posedata/imu_own_3c/ — our straps | 274 | by stance (154 orthodox / 120 southpaw) |
posedata/pkl/boxingvi_pub_3c_* | 5,451 | by subject, V4+V5 held out (4,293 / 1,158) |
The BoxingVI README advertises 6,915 clips from 18 athletes across 20 videos. The distributed archive contains V1–V10, 5,570 annotated rows, ten subjects. The release also carries two incompatible layouts — nine videos pre-segmented where row i is clip i, one continuous where the start/end columns are real indices. Our converter detects the layout per video, finds the class column by content rather than position, and skips any video whose clip count disagrees with its label count.
Collecting
The in-app collection section drives a self-ticking checklist: six real punches per stance (jab, cross, lead/rear hook, lead/rear uppercut), footwork negatives, and the six-pose calibration. Each clip is written with its full context — mount, stance, surface, intensity, hand role — because a training run that cannot filter by those cannot answer the questions the protocol was designed to answer.
.venv/bin/python scripts/convert/ios_punches_to_imu_clips.py \
--input punches_*.jsonl --labels labels_*.jsonl --subject V11
A measured rule: accuracy saturates at roughly 85 clips per stance and is flat thereafter. More punches from the same athlete add nothing; each new athlete opens ground none of the existing data covers.
Review loop
While recording, the live feed becomes a review queue. Each punch shows what the model called it, with every class beside it at equal weight and nothing pre-selected.
Fitting a model to its own output teaches it to repeat its errors, and the errors sit exactly where it is confident and wrong. The only new information in the loop is a human saying what the punch was.
- Confirmed — cheap to give, and the confirmed/corrected ratio is live accuracy, with no held-out split to argue about.
- Corrected — rarer, and worth more per clip than anything else collectable, because it falls exactly on the misplaced boundary.
- Not a punch — the only measure of the detector's live precision, which no offline set can supply.
Verdicts append to labels_<session>.jsonl; last line per clip wins.
Re-labelling appends rather than edits, so a clip relabelled three times is visibly
ambiguous. An unreviewed punch is stored as neither and is never silently
promoted.
Training
# camera trunk — prefers boxingvi_pub_3c_* when present
./scripts/train_pose.sh
# wrist classifier — deterministic, regenerates the Swift file
.venv/bin/python scripts/fit_own_rig_shape.py --gate 4.0
The wrist fit is deterministic (zero init, full batch, no shuffling), so the generated Swift is byte-identical across runs and a diff means the data changed.
CoreML export
coremltools defaults to FLOAT16 and will quietly halve your precision. Pin
compute_precision=ct.precision.FLOAT32. On the IMU model this moved
parity error from 1.02e−2 to 2e−6.
The export script takes --classes and refuses on a mismatch, after an
earlier version sliced a hardcoded 4-class list to length and silently wrote
['jab','cross','hook'] onto a 3-class trunk.
Tests
Eight Swift test files and eight Python ones. The ones worth knowing about: spotter parity between the Swift and Python implementations, the calibration solver against known rotations, CoreML parity against the PyTorch source, and gate behaviour on merged pairs.
Constants
| Constant | Value | Where |
|---|---|---|
minPunchG | 3.0 g | IMU peak search arms |
minEmitG | 4.0 g | IMU candidate emitted |
rearmRatio | 0.5 | Both spotters |
refractory | 250 ms | IMU; a 1–2 lands 300–400 ms apart |
clipDuration | 48/52 s | 923 ms, the training clip length |
leadFraction | 0.45 | Fraction of the clip before the peak |
KEYPOINT_VALID_MIN | 0.3 | Pose confidence gate |
refractoryFrames | 15 | Pose, ~0.5 s at 30 fps |
CLIP_LEN | 48 | Pose window, peak-centred |
Scripts
| Script | Purpose |
|---|---|
scripts/fit_own_rig_shape.py | Fit the wrist classifier, regenerate the Swift weights |
scripts/convert/boxingvi_public_to_pkl.py | Authors' release → training PKLs |
scripts/convert/ios_punches_to_imu_clips.py | Captured sessions + review labels → clips |
scripts/convert/collapse_4c_to_3c.py | Collapse jab+cross into straight |
scripts/train_pose.sh | Train the camera trunk |
scripts/flash_imu.sh | ./scripts/flash_imu.sh left|right |
scripts/whitepaper_figures.py | Regenerate every whitepaper figure from data |
scripts/whitepaper_diagrams.py | Schematic figures for the whitepaper |
Phase-2 archive
Several documents in docs/ describe a different architecture — YOLO
detection, OC-SORT tracking, PoseC3D — explored before the current cascade. They
are deliberately retained as the design paper trail for a multi-fighter phase and are
not a description of what ships today.
Affected: docs/ARCHITECTURE.md, docs/TECHNICAL_OVERVIEW.md,
docs/API_REFERENCE.md, docs/PROJECT_STATUS.md. For current
behaviour, this site and MODELS.md are authoritative.