Arikex

Motion capture.Now in your pocket.

A metric 3D skeleton from one iPhone. No markers, no suit, no lab.

ANATOMICAL BONE MODEL
16 rigid segments · the app's own rig
BASEBALL · HIT
24 fps · Mixamo on the Arikex rig
240 fps capture · 1.2 s LiDAR calibration · 26 keypoints · metric 3D on the phone
Drag to orbit
iPhonecaptures, calibrates and solves the skeleton itself
Macingests every phone, triangulates, stages the shot
NVIDIAtrains the athlete as a 4D splat for bullet time
Why now

Motion capture has been a room.

For thirty years the answer to "how does this body move" was an optical lab: a dozen cameras on a truss, a suit of reflective markers, an afternoon of calibration and a technician to clean the data. It works. It also costs more than a car, never leaves the building, and cannot be pointed at a batter in a cage or a patient in a hallway.

0%

Annual growth of markerless capture through 2030, well ahead of the whole market.

Grand View Research
$25kto$500k

What an optical lab costs, from an entry rig to an in-stadium install.

MoCap Online · KinaTrax
0°

Joint-angle error two iPhones reach against a marker lab. Multi-camera markerless averages 2.3° in gait.

Uhlrich et al. · 2024 meta-analysis
2to8×

Hours of cleanup per hour of optical capture. Calibration alone runs one to two hours.

Studio setup guides, 2025

The optical lab

  • 8 to 40cameras on a truss in a dedicated 3 × 3 × 2.5 m volume
  • 1 to 2 hof wand calibration before the first take, then a marker suit on every subject
  • ~1 mmmarker precision, the bar nobody argues with
  • $500 to $5ka studio day plus a technician to clean the trajectories
  • Next weekthe athlete goes home, the numbers arrive later

Arikex

  • 1 to 4iPhones on tripods or in your hand, wherever the athlete already is
  • 1.2 sLiDAR pre-roll on every take gives the solve true bone lengths, stature and the floor
  • 240 fpscapture with a metric 3D skeleton solved after the take, live preview at 60
  • 0 suitsa 26-keypoint body-and-feet model runs on the phone, hands from Apple Vision
  • Minutesfrom swing to joint angles, velocities and a coaching read, on the device that shot it
How it works

Four steps. One take.

1 of 4 · Capture

Point the phone. Tap the watch.

Three capture modes on one screen: a live ARKit preview, Vision at 30, 60, 120 or 240 fps, and Track, which records the camera's own motion so a handheld shot can be stabilized later. A wrist remote starts and stops the take from the batter's box.

1080p · 240 fpsLiDAR pre-rollApple Watch remote
2 of 4 · Calibrate

LiDAR measures the body before it moves.

Every take opens with 1.2 seconds of depth. From it Arikex solves bone lengths, stature, distance and a floor plane, then rejects anything that is not a body, like the ghost a mirror returns. High-speed frames inherit the scale: physics does not care about frame rate.

Bone lengths in cmFloor planeAnatomy check
3 of 4 · Solve

2D keypoints become a metric skeleton.

A 26-keypoint model runs on the phone's neural engine, then a whole-take optimizer lifts it into meters: bone-length constancy, a rigid torso, temporal smoothness, ground contact, with LiDAR arbitrating the folds a single camera cannot see. Two phones triangulate instead.

RTMPose on deviceSequence refiner2-camera triangulation
4 of 4 · Understand

Angles, velocities, and a coach who read them.

Joint angles, hip-shoulder separation, trunk lean and their velocities, timed to the frame. A Deep Analysis pass hands only those numbers to a language model and returns a report where every timestamp scrubs the footage. Video never leaves the phone.

metrics.jsonBVH · USDZClaude or Gemini
REC 00:01.20
VISION · 240
2D · 26 kpt
SOLO
One capture core, three machines

The phone records. The Mac reasons. The GPU dreams in 4D.

Every take is a folder of plain files: video, depth, keypoints, calibration, camera track. The same analysis code compiles for iOS and macOS, so a phone can do everything alone or just shoot and send. When a session needs a dynamic Gaussian splat, the Mac hands one command to an NVIDIA box and brings the frames back.

iPhone

Capture · Solve
  • On-device 2D pose. A 26-keypoint body-and-feet model through the neural engine, hands from Apple Vision.
  • Metric 3D on the phone. LiDAR calibration plus a sequence solver. Bone model, 3D viewer, BVH, metrics, AI report.
  • Two phones, one button. A director phone syncs a second over peer-to-peer Wi-Fi and fires both takes together.
  • Apple Watch remote. Set the tripod, get in the box, tap the wrist.

Mac

Ingest · Triangulate · Stage
  • Same engines, bigger machine. Vision, RTMPose, Metric 3D and Isolate run on the Mac through the phone's own code, so a phone can just record and send.
  • Multi-camera truth. Phones find the Mac over Bonjour and deliver take packages; two views triangulate a real 3D skeleton through Pose2Sim and OpenSim.
  • Stage. Put a take on a floor, add gray props and a camera with a real lens, render a gray-box reference for a video model.

NVIDIA

4D splats · Bullet time
  • One command over Tailscale. The Mac converts a session into a per-frame COLMAP layout, syncs it to WSL2 on the PC, trains, and pulls one PLY per frame back.
  • The subject as a 4D splat. Four phones 45° apart give a 135° arc you can freeze, scrub and orbit. Bullet time for a swing, for coaching, never for generation.
  • Nothing else crosses. Every other stage is Mac-native, including a static splat of one frozen instant.
Built to adapt

One skeleton. Five rooms it walks into.

The capture core is the same everywhere: a metric, anatomy-first skeleton with per-segment poses, velocities and confidence. What changes is what you do with it.

A pitching lab in a batting cage.

Hip-shoulder separation, trunk lean, elbow and knee angles and their angular velocities, timed to the frame at 240 fps. A Deep Analysis pass turns the numbers into cues a coach would give, with every timestamp scrubbing the footage. In-stadium systems do this for $500,000 and a camera crew. A cage needs a tripod.

42°Peak separation
240Frames per second
1.08 sTime of peak

Gait analysis in the hallway, not the lab.

A metric skeleton with a real floor gives knee and hip flexion through the gait cycle, stride length in centimeters, cadence and left-right asymmetry, from one phone on a stand. Two phones at front and 45° triangulate for the numbers a clinic can act on. Video stays on the device.

14°Left / right asymmetry
1.24 mStride
108Steps per minute

From take to rig without a suit.

Every solved take exports BVH: a pelvis-rooted hierarchy with calibrated bone offsets, root translation and per-joint rotations, plus USDZ of the posed rig. Drop it on a Rigify or Mixamo character in Blender, Unity or Unreal. Fingers come from Apple's hand tracker, feet carry real sole orientation.

25Joints in BVH
60 fpsRotations
0SMPL

The reference a video model deserves.

Stage places a take on a LiDAR-scaled floor with gray props and a camera with a real focal length, then renders a gray-box pass at 24 fps with depth, matte and camera JSON. Sent to Seedance 2.5 or Runway Act-Two, the same framing and the same move come back photoreal. Real faces are never uploaded, the mannequin is.

9 sShot at 720p
405Credits, proven
<3 minReturn

Block the shot where you shot it.

Track mode records the camera's own path with the performance, so a handheld take becomes both an actor and a camera move. Stabilize it to lock the actor to the floor, or import the move as the Stage camera. Export USDZ, BVH and per-frame camera JSON to Blender, or skip Blender entirely.

3.84 mCamera travel
115°Sweep
36 mmSensor reference
Accuracy, honestly

We say what a camera can and cannot see.

A single camera cannot tell which way a limb folds in depth. Arikex does not pretend otherwise: LiDAR arbitrates the folds it can measure, the solver enforces what a body can do, and every report carries the take's own self-check, reconstructed stature against the LiDAR measurement. The path to lab-grade numbers is a second phone, and the Mac already walks it.

MethodCamerasJoint-angle error
Optical markers (Vicon, OptiTrack)8 to 40~1 mm position
Markerless multi-camera, gait meta-analysis8+2.3° mean
Theia3D, 16 validation studies8+< 6° RMSE
Two iPhones (OpenCap, marker-validated)24.5° MAE
One iPhone (OpenCap Monocular, 2026)14.8° MAE
Inertial suit (Xsens)06 to 18° RMSE
Arikex, one phone with LiDAR scale1validation shoot planned
Arikex, two phones front + 45°2OpenCap-class pipeline
Sources: Uhlrich et al. 2024 · PMC 2024 meta-analysis · AI in Medicine 2025 · OpenCap Monocular 2026. Arikex figures will be published after a marker-lab shoot.
Arikex mark

Point a phone at it.

Arikex is in private testing with athletes, coaches and clinicians now. If you run a facility, a clinic, a studio or a pipeline that needs motion, we want to point it at yours.

Request a capture session
Arikex · 2026Apache and BSD stack · no SMPL · video stays on the deviceiPhone 16 Pro · macOS 14 · RTX 4090