OVERHEAD · FRONT + 45° · SYNC +20 MS
READINESS
TAKES
| swing_07 | 240 fps | solved |
| swing_08 | 240 fps | solved |
| swing_09 | 120 fps | 4D · 2 h |
A metric 3D skeleton from one iPhone. No markers, no suit, no lab.
For thirty years the answer to "how does this body move" was an optical lab: a dozen cameras on a truss, a suit of reflective markers, an afternoon of calibration and a technician to clean the data. It works. It also costs more than a car, never leaves the building, and cannot be pointed at a batter in a cage or a patient in a hallway.
Annual growth of markerless capture through 2030, well ahead of the whole market.
What an optical lab costs, from an entry rig to an in-stadium install.
Joint-angle error two iPhones reach against a marker lab. Multi-camera markerless averages 2.3° in gait.
Hours of cleanup per hour of optical capture. Calibration alone runs one to two hours.
Three capture modes on one screen: a live 3D preview, high-speed capture at 30, 60, 120 or 240 fps, and Track, which records the camera's own motion so a handheld shot can be stabilized later. A wrist remote starts and stops the take from the batter's box.
Every take opens with 1.2 seconds of depth. From it Arikex solves bone lengths, stature, distance and a floor plane, then rejects anything that is not a body, like the ghost a mirror returns. High-speed frames inherit the scale: physics does not care about frame rate.
A 26-keypoint model runs on the phone's neural engine, then a whole-take optimizer lifts it into meters: bone-length constancy, a rigid torso, temporal smoothness, ground contact, with LiDAR arbitrating the folds a single camera cannot see. Two phones triangulate instead.
Joint angles, hip-shoulder separation, trunk lean and their velocities, timed to the frame. A Deep Analysis pass hands only those numbers to a language model and returns a report where every timestamp scrubs the footage. Video never leaves the phone.
Every take is a folder of plain files: video, depth, keypoints, calibration, camera track. The same analysis core runs on the phone and on current-generation Apple silicon, so a phone can do everything alone or just shoot and send. When a session needs a dynamic Gaussian splat, the workstation hands one command to a CUDA box and brings the frames back.
The capture core is the same everywhere: a metric, anatomy-first skeleton with per-segment poses, velocities and confidence. What changes is what you do with it.
Hip-shoulder separation, trunk lean, elbow and knee angles and their angular velocities, timed to the frame at 240 fps. A Deep Analysis pass turns the numbers into cues a coach would give, with every timestamp scrubbing the footage. In-stadium systems do this for $500,000 and a camera crew. A cage needs a tripod.
A metric skeleton with a real floor gives knee and hip flexion through the gait cycle, stride length in centimeters, cadence and left-right asymmetry, from one phone on a stand. Two phones at front and 45° triangulate for the numbers a clinic can act on. Video stays on the device.
Every solved take exports BVH: a pelvis-rooted hierarchy with calibrated bone offsets, root translation and per-joint rotations, plus USDZ of the posed rig. Drop it on any standard humanoid rig in your engine or DCC. Fingers come from the hand model, feet carry real sole orientation.
Stage places a take on a LiDAR-scaled floor with gray props and a camera with a real focal length, then renders a gray-box pass at 24 fps with depth, matte and camera JSON. Sent to a current video generation model, the same framing and the same move come back photoreal. Real faces are never uploaded, the mannequin is.
Track mode records the camera's own path with the performance, so a handheld take becomes both an actor and a camera move. Stabilize it to lock the actor to the floor, or import the move as the Stage camera. Export USDZ, BVH and per-frame camera JSON to your DCC, or skip it entirely.
A single camera cannot tell which way a limb folds in depth. Arikex does not pretend otherwise: LiDAR arbitrates the folds it can measure, the solver enforces what a body can do, and every report carries the take's own self-check, reconstructed stature against the LiDAR measurement. The path to lab-grade numbers is a second phone, and the workstation already walks it.
| Method | Cameras | Joint-angle error |
|---|---|---|
| Optical markers (Vicon, OptiTrack) | 8 to 40 | ~1 mm position |
| Markerless multi-camera, gait meta-analysis | 8+ | 2.3° mean |
| Theia3D, 16 validation studies | 8+ | < 6° RMSE |
| Two iPhones (OpenCap, marker-validated) | 2 | 4.5° MAE |
| One iPhone (OpenCap Monocular, 2026) | 1 | 4.8° MAE |
| Inertial suit (Xsens) | 0 | 6 to 18° RMSE |
| Arikex, one phone with LiDAR scale | 1 | validation shoot planned |
| Arikex, two phones front + 45° | 2 | OpenCap-class pipeline |
Arikex is in private testing with athletes, coaches and clinicians now. If you run a facility, a clinic, a studio or a pipeline that needs motion, we want to point it at yours.
Request a capture session