Skip to content

Performance ​

Both platforms implement all of this.

Profiles ​

tsx
<PoseCamera profile="auto" />   // default

Every profile is one row of the same model. It sets a ceiling, how much of the time inference may run (its duty), how it idles, and when heat counts:

ProfileCeilingDuty, nominal / fairIdle after 2 s / 20 sHeat acts at
auto (default)the camera's rate85% / 70%12 / 5 fpsserious, critical
qualitythe camera's rate95% / 85%15 / 8 fpsserious, critical
balanced24 fps70% / 60%12 / 5 fpsserious, critical
efficient15 fps50% / 40%8 / 3 fpsfair (×0.75), serious, critical
unrestrictedthe camera's rate100%offcritical only

Every profile measures the device and budgets against what its inference costs; they differ only in how much of that they spend. auto runs at the camera's rate whenever the device can do so with 15% to spare, which is 30 fps on any recent phone.

Inference never runs faster than the camera. The camera is pinned to 30 fps where it allows, so auto-exposure cannot halve the rate in a dim room, and nothing above 30 is offered: a phone asked for 60 ran warm within minutes for a skeleton that looked identical at half that.

One axis stops lower than a spec sheet would suggest, and it is deliberate. Analysis tops out at 480p. MediaPipe resizes whatever it is handed to 256 by 256 before the landmark model sees it, so a 720p analysis buffer is close to a megapixel captured, converted and copied every frame in order to be discarded inside the graph. A distant subject is the one case a larger buffer helps, and analysisResolution is there to ask for it.

The rate ​

Every inference reports what it cost, dispatch to result. The median of that cost, p50, sets the rate at which inference is busy a given share of the time, and the profile's duty picks the share:

text
capacity(duty) = duty × 1000 ÷ p50

nominal   min(camera, capacity(duty nominal))
fair      min(camera, capacity(duty fair))
serious   min(camera ÷ 2, capacity(50%))
critical  detection paused, preview kept

Why not 100%? One frame is in flight and one waits: in MediaPipe's live mode on iOS, in the camera's latest-frame slot on Android. At 100% every frame waits for the one before it, which adds a frame of latency and runs the GPU without a breath; the spare 15% keeps that queue empty and the device cool, for at most a few frames a second less on a slow phone. What that gives under auto:

p50NominalFairSerious
16 ms303015
20 ms303015
25 ms302815
30 ms282315
40 ms211712
60 ms14118

An iPhone 15 measures 16 to 18 ms for the full model on its GPU, the first row: the camera's 30 fps with the GPU idle half the time.

A governed rate never drops below 10 fps for a slow device, because below that the skeleton reads as broken; heat and idle may go lower. Low Power Mode on iOS and Battery Saver on Android cap it at 24.

An explicit targetFps is capped only by the camera and by what the device can finish at all, capacity(100%): feeding MediaPipe faster than it can finish only queues frames behind each other. Duty and fair heat leave it alone; serious and critical heat still apply unless thermalPolicy says otherwise.

Why the rate is what it is ​

Every reading of the rate comes with the constraint that set it, limitedBy:

limitedByMeaning
cameraThe camera's own rate: nothing faster exists to run on
deviceWhat this device finishes within its duty budget
targetYour targetFps
profileThe profile's ceiling, balanced or efficient
thermalHeat, including detection paused at critical
lowPowerLow Power Mode or Battery Saver
idleNobody has been in frame for a while
pausedDetection is off, or the camera is not running

It is on getState(), getProfile(), onReady and every onPerformanceChange.

Measuring the device ​

Before anything is measured, the rate is the camera's. An unknown device is not a slow one, and half a second at the camera's rate costs less than a start that looks slow.

The first estimate lands after 15 frames with a pose, about half a second. After that the median is refreshed every 15 frames over the last 60, moves of 2 fps or less are ignored, and a three-second cooldown after each change stops a device sitting between two answers from oscillating. A move that changes the rate fires onPerformanceChange({ reason: 'calibration' }), and so does a setProfile() call that changes it.

The measurement is cached, keyed by device model, model file, OS version and MediaPipe version, so the second launch starts from it. It is kept across camera restarts in the same session: a restart is not a new device. The GPU check is cached the same way, so it runs once per device and model rather than on every mount.

The tier (high, medium, low) is a label read off the same median, ≤ 22 ms high and ≤ 45 ms medium. It is reported so an app can reason about the device; it drives nothing. Before a measurement it comes from installed memory, and on Android from the number of cores too.

Inspecting it ​

ts
await cam.current?.getProfile();
// { profile: 'auto', phase: 'settled', source: 'measured', tier: 'high',
//   resolved: { delegate: 'GPU', targetFps: 30, preview: '1080p', analysis: '480p' },
//   p50InferenceMs: 16.2, measuredFps: 30, limitedBy: 'camera',
//   cameraFps: 30, thermalState: 'nominal', lowPower: false }

cam.current?.getState();
// { ..., fps: 30, limitedBy: 'camera' }

measuredFps is completed inferences over the last second, zero once results stop, so it is the number that exposes a device falling behind its target. getState().fps and limitedBy are read live from native on the JavaScript thread, with no hop to native's main thread.

Heat ​

Read from the OS once a second, and on iOS also whenever it notifies, never on the frame path:

StateResponse under auto
nominalthe rate above
fairduty drops to 70%
serioushalf the camera's rate at most
criticaldetection paused, preview continues, event emitted

Heat is adopted the moment it rises and cooling only after it has held for 30 seconds, at the warmest level seen meanwhile, so a device hovering on a boundary does not flap the rate.

Android names more states than iOS: NONE and LIGHT count as nominal, MODERATE as fair, SEVERE as serious, and CRITICAL and above as critical. On Android 11 and later the OS's forecast of where heat is heading counts too, 85% as fair and 95% as serious, so the rate backs off before the device starts throttling. Android 9 and older report no heat at all, so it never acts there.

thermalPolicy="critical-only" acts only at critical, and "off" never. Neither stops the reporting: onPerformanceChange still fires on every change of heat, with thermalState, so your app can decide for itself.

Camera geometry ​

Preview and analysis sizes are fixed for a session. Under the auto profile an auto preview is 1080p on a device with at least 5.5 GiB of memory and 720p otherwise, never 480p unless you ask. The other profiles fix it: 1080p for quality and unrestricted, 720p for balanced and efficient, and efficient analyzes at 360p rather than 480p. Nothing the governor learns changes geometry mid-session, so no measurement or heat reading ever restarts the camera. Only a resolution, analysisResolution or profile change does.

Precedence ​

text
1. profile        sets the ceiling, duty, idle and heat rows
2. targetFps      replaces the governed rate, capped by the camera and the device
3. heat           serious and critical apply to everything, unless thermalPolicy says otherwise
4. low power      caps the governed rate at 24

So this does exactly what it reads like:

tsx
<PoseCamera
  profile="quality"          // high ceiling, 95% duty
  targetFps={24}             // pinned: only what the device can finish caps it
  analysisResolution="auto"  // stays 480p
/>

Optimizations ​

What it does
Built while the camera opens (Android)The landmarker starts building on mount, beside the camera, not after it. iOS starts it once the camera is running
Pre-warmOne inference on a blank frame before the camera's first reaches the landmarker, so the first real frame is never the slow one
CPU first, then the GPU (Android)auto answers frames on the CPU landmarker, which builds in a fraction of the time, while the GPU one builds beside it; the GPU takes over mid-track once it has built. See budget Android phones
GPU check, onceOn Android the warm-up of the GPU landmarker is the check that the GPU works, on iOS a probe inference on a throwaway one; either runs once per device and model and is remembered. A GPU that fails at runtime is swapped for the CPU and the answer flips
Parked landmarkerTurning detection or the camera off, a trip to the background, or a screen pushed on top keeps the landmarker built for 30–60 s, so coming back is instant. On Android a camera screen closed and opened again within a minute takes back the landmarker it left
Visibility on a clockMediaPipe smooths each joint's visibility once per frame; it is re-timed to elapsed time, so a joint appears as fast at 10 fps as at 30
Idle searchNo person for 2 s drops to the profile's first idle rate, 20 s to its deep one; the frame that finds a pose ends it
Smoothing 'auto'Off for one pose, which MediaPipe already smooths; on for several, with MediaPipe's own constants
Lazy anglesComputes only the angles an angle condition, overlay.angles or data.angles asked for
Analysis ≠ previewModel sees a small frame; preview stays sharp
One overlay draw per resultShape layers the GPU composites on iOS; a hardware-accelerated view redrawn once per result on Android, which on a budget phone costs about 9 ms of each detection because the redraw shares the GPU with the model
Frames read on the JavaScript threadA drain reads the ring buffer directly, never queued behind native's main thread

Budget Android phones ​

What a low-end phone can do, measured on a Redmi Note 12 (Snapdragon 685, Adreno 610 GPU) with a person in view:

ModelDelegateBuildPer frameTop rate
fullGPU1.9 s89 msabout 11 fps
fullCPU0.7 s122 msabout 8 fps
liteGPU1.6 s63 msabout 15 fps
liteCPU0.7 s80 msabout 12 fps

The GPU is the faster and the cooler of the two once running, at half the CPU time, and the slower to build, so auto starts on the CPU and hands over. The first skeleton comes about 1.1 s after mount, the GPU takes over about 3 s in, and a screen opened again within a minute is back in 0.6 s. The governor then runs full at 10 fps on this phone: its 85% duty allows 9, and 10 is the floor a governed rate never goes below.

For more frames on budget phones, ship lite. It is one word in the plugin config and runs about half again as fast here, for the accuracy the model table describes. Keep maxPoses at 1 unless you need more: with more, MediaPipe runs its person detector on every frame in which it is tracking fewer people than maxPoses, and that detector costs about as much as the landmark model.

Tried on this phone, and not used:

  • OpenCL. Declaring libOpenCL.so lets MediaPipe load it, and its first attempt then took the GPU build from 1.9 s to 3.5 to 4 s and ran no faster. Without the declaration MediaPipe goes straight to OpenGL, which is what ships.
  • The GPU kernel cache. MediaPipe writes it only from its OpenCL path, so it saves nothing here.
  • More CPU threads. MediaPipe's Java API runs the CPU delegate on one thread and has no setting for more.
  • The NPU. MediaPipe's NPU delegate needs vendor dispatch libraries this chip does not have.
  • Gliding the skeleton between results. Drawing it moving toward each new result, instead of stepping to it, looked smoother and took detection from 9.8 fps to 7.2 at 30 redraws a second and to 4.2 at 60: every redraw is GPU work competing with the model's. The overlay redraws once per result.

Resource budgets ​

Targets, not enforced ceilings. The zero-allocation claims below are held by the code and its reviews; the memory and 10-minute sustained-run numbers are measured on one device so far and harden as the device matrix grows. They are the numbers a bug report should be filed against.

StateTarget above app baseline
Camera on, detection off< 40 MB
lite @ 480p analysis< 120 MB
full @ 480p analysis< 180 MB
Steady-state allocations per frame0, except the MPImage floor below
Return to idle after stopDetection()immediate: frames stop at once, the landmarker is freed after 60 s unused
10-minute sustained runthermal ≤ fair on mid-tier

The zero-allocation claim has one floor, and it is honest to name it. Handing a frame to MediaPipe requires an MPImage, and building one allocates about seven objects: the builder, the container, the image, its properties, and a small map inside the image. There is no API that takes a reusable one. Everything this package controls, the landmark buffers, the geometry, the filter, the evaluators and the ring buffer, allocates nothing per frame at the default configuration. The iOS overlay builds a few small paths per frame for its shape layers, a few hundred bytes, which is what replaced rasterizing a full-screen bitmap on the CPU.

Two configurations do allocate beyond that floor, both by choice: data.mode: 'live' allocates one direct buffer per drain, which is what carrying frames to JavaScript costs, and an angle overlay with decimals above zero formats a string per label per draw, as every angle label does on iOS.

App size ​

Exactly one model ships, whichever model your config selects.

Two things are measured, both read out of an assembled APK on the pinned 0.10.35. MediaPipe's native libraries come to 10.08 MB for arm64-v8a, 7.09 MB for armeabi-v7a, 14.31 MB for x86 and 12.48 MB for x86_64. The model files are 5.5 MB (lite), 8.96 MB (full) and 29.2 MB (heavy).

The native libraries are already compressed and do not shrink again inside the APK, so what is on disk is what is downloaded. The model compresses by about a tenth, 8.96 MB down to 8.03 MB for full, because float16 weights are close to incompressible.

The JavaScript is the part that rounds to nothing: about 70 KB of built output, and no runtime dependencies to pull in behind it.

Everything else is an estimate:

ModelAndroid install / Play downloadiOS
lite~19.7 MB / ~10.9 MB~26–41 MB
full~23.2 MB / ~14.2 MB~29–44 MB
heavy~43.4 MB / ~33.0 MB~49–64 MB

No release archive has been built and weighed yet, so these stay estimates until one has, per model and per platform.

Android requires an AAB. A universal APK carries all four ABI slices, 43.96 MB of native library where a phone loads 10.08 MB of it. Set abiFilters on your release build if you must ship an APK, and only there: dropping x86_64 from a debug build is what breaks the standard emulator on an Intel host.

Model files are ~93% incompressible (float16 weights), so they cost nearly full price on download.

MIT licensed. Built from the guides in the repository.