For this experiment I wanted to treat the face as an input device rather than a subject. FaceMesh (via ml5.js) hands you around 468 landmarks on every frame — a dense, live point cloud of your own face. The obvious thing to do is draw dots on top of the video. The more interesting thing, I decided, was to take those numbers and push them into a completely different domain than the face itself: sound, geometry, text. The surprise comes from the translation, not the tracking.
The trick that made all three feel solid was normalizing every measurement by the distance between my eyes. Raw pixel distances change the moment you lean toward the camera, but “mouth gap ÷ eye distance” stays stable whether I’m close or far. Once I had a handful of scale-invariant expression values — mouth openness, smile width, eyebrow raise, head tilt — each sketch was really just a mapping decision.
1. Facesynth — a silent face that sings
[GIF: opening and closing my mouth to play notes]
The first sketch turns my face into an instrument. Mouth openness sets the pitch, and I quantized it to a pentatonic scale so it always lands on something musical instead of a siren. Raising my eyebrows jumps up an octave, smiling adds a shimmering second voice, and tilting my head pans the sound left and right. Every note I “sing” leaves a glowing orb that drifts upward, so the melody becomes something you can see accumulating on screen.
What surprised me was how performable it became once the pitch was quantized. An unquantized version felt like a theremin I was fighting; snapping to a scale turned it into something I actually wanted to play with for ten minutes.
2. Shatter — a portrait that flies apart
[GIF: mouth open → shards explode → mouth closed → shards reassemble]
Here I used faceMesh.getTriangles(), which returns the real mesh triangulation, and filled each triangle with the color sampled from the webcam underneath it — a live, low-poly, stained-glass version of my face. Then I gave every shard its own physics: opening my mouth kicks the triangles outward with real momentum and a bit of gravity, so my face explodes; closing my mouth springs them all back home. Tilting my head spins the whole mosaic.
This is the one that got an actual reaction out of people. There’s something uncanny about watching your own face hold together as a thin membrane of glass and then blow apart because you opened your mouth.
3. Facetype — a portrait made of words
[GIF: face emerging out of streaming text]
The last one rebuilds my face out of flowing text. It samples the brightness of the video on a grid and stamps the next letter of a sentence at a size based on how bright that spot is — bright areas get big letters, shadows drop out into negative space. The phrase streams continuously through the image, I rendered it as a two-tone duotone so it reads like a designed poster, and opening my mouth surges the type bigger. The sentence is a variable at the top of the file, so the portrait literally writes whatever I want it to say.
Process — and arguing with an AI agent
The assignment invited using an AI coding agent to spin up multiple iterations, so I leaned into that and treated it as a collaborator I had to direct. That part was its own lesson: the agent was fast at generating variations, but the interesting work was still mine — deciding which expression should drive which parameter, and why. A good mapping is a design decision, not a code decision.
It wasn’t frictionless. One version threw TypeError: scale is not a function because I’d named my array of note frequencies scale, which quietly shadowed p5’s built-in scale() that I was using to mirror the video. A good reminder that the webcam feeling like magic doesn’t exempt you from boring namespace bugs.
Reflection
The reference points I keep coming back to for this kind of work are Zach Lieberman’s daily face sketches and Kyle McDonald’s older FaceOSC experiments — the lineage of treating the face as a live controller rather than a photograph. What this exercise clarified for me is that the camera is just a sensor; the design happens in the mapping. The same 468 points became an instrument, a sculpture, and a sentence depending only on where I decided to send them. That gap — between raw model output and a felt, embodied interaction — is where the interesting choices live.
Leave a comment