Teaching AI to Warm Me Up: The Body the Model Can’t See
The starting point
This assignment asked us to build a p5.js experiment driven by an ml5.js body model. Rather than make abstract interactive art, I wanted to build something I’d actually use: a warmup coach. Stand in front of the camera, do your warm-up moves, and the system recognizes what you’re doing and tracks your progress. As someone who trains regularly, I was curious whether AI can really “see” a proper warmup.
Before building, I read three pieces that reshaped how I approached the project.
What the readings gave me
Maya Man’s Mixing Movement and Machine reframed machine learning for me: not as a replacement for creativity, but as a new material. The interesting work is finding the fit between what the model does well and your creative intent, letting the body itself become the interface. That set my direction — instead of asking the model to make judgments it’s bad at, use its most reliable ability (joint tracking) to drive something clear.
Philipp Schmitt’s Humans of AI reminded me that behind every seemingly objective model are human choices — who labeled the data, who was filmed into the training set. The model is not a neutral eye.
And Ellen Nickles’ Model and Data Provenance Project gave me a concrete framework: a model needs a “biography” — what data trained it, who made it, what its limits are. When we call a pre-trained model, we’re accepting someone else’s definition of what the world looks like and how a body should move.
What I built
I used ml5.js’s BlazePose model to track body keypoints, then inferred four basic warm-up exercises from their relative positions:
- Shoulder circles — wrists raised above the shoulders
- Leg kicks — knee lifted above the hip
- Hip rotations — the width difference between hips and shoulders
- Arm circles — how far the wrists travel from the shoulders
Visually, I deliberately hid the raw camera feed, keeping only a glowing skeleton on a dark background. Partly it looks cleaner and records better — but it’s also an intentional statement: I wanted the viewer to see only what the model sees, a set of coordinates and lines, not my face. When a move is detected, its card lights up green and the progress bar advances.
What the model sees — and what it doesn’t
Once it was actually running, the most interesting part wasn’t what it could detect, but what it couldn’t:
It’s accurate about position — where my limbs are, relative joint heights, the confidence of each point. But it’s completely blind to:
- Movement quality — it knows my knee went up, not whether I did it correctly or with control
- Injury risk — if my form is dangerous (knees caving in), the model doesn’t care
- Individual range of motion — the same coordinate might be my limit and only halfway for someone more flexible
In other words, it knows where my body is, but not whether I’m doing it right. And a warm-up is all about the latter.
Thinking about provenance
After reading Nickles’ project, I started questioning the data behind BlazePose. It comes from Google, trained on large amounts of human-body video — but who is in that video?
I suspect its diversity has limits: likely skewed toward young, “standard” bodies, and the very definition of a “normal pose” is already framed by the training data. That means: if my warm-up style departs from what it has seen, it struggles to read me. How would it recognize a wheelchair user’s upper-body warmup? Stretching styles from different bodies, different cultures?
Questions I still have
- What is the actual demographic makeup of the training video — age, body type, ability, region?
- What assumptions about “a normal pose” are baked into the model?
- If I could add to this model’s “biography,” what I’d most want isn’t where it succeeds, but where it fails — failure cases reveal a model’s boundaries far better than its successes.
How provenance changed my process
I used to evaluate a model with one question: “Does it work?”
Now I add a second: “Who does it work for, and who does it fail?”
That shift matters. Building this coach, I realized I can’t assume the model is fair to everyone. If I were to turn it into a real product, I’d need to know who it works for, who it misreads, and design fallbacks for those edge cases. Understanding provenance isn’t academic fussiness — it’s the precondition for building responsibly.
Closing
AI isn’t an “objective eye.” It’s a mirror of a particular dataset. What it can see reflects whose body, whose way of moving. A warmup coach that actually works starts from knowing who the model serves and who it fails — and only then using it responsibly.
sketch:https://editor.p5js.org/ervin021013/sketches/Z1bESrIUE
Leave a comment