Annotating moving objects in video
Configure motion tracking, correct saved tracks, and try experimental hover discovery.
Core provides manual annotation on paused video. See can advance those boxes as the video plays when you supply compatible tracking models. Tracking is not enabled just by adding the default See extension.
Configure tracking
Host compatible NanoTrack ONNX weights and configure their URLs:
import { PopmeltProvider } from '@popmelt.com/core';
import { configureSee } from '@popmelt.com/see';
const vision = configureSee({
videoTracking: {
backboneUrl: '/models/nanotrack-backbone.onnx',
headUrl: '/models/nanotrack-head.onnx',
device: 'auto',
wasmPaths: '/onnx-runtime/',
},
});
<PopmeltProvider extensions={[vision]}>{children}</PopmeltProvider>;
The URLs are examples, not bundled files. Supply weights under terms appropriate for your application and serve ONNX Runtime's WASM assets when fallback is required. Cross-origin videos must permit CORS for frame crops and tracking.
Create and correct a track
- Pause on a clear frame.
- Enter Chat mode, drag around the object, and name it.
- Play the video to track the selected object.
- Double-click its label or box to rename it or correct its bounds at the current frame.
Tracking can drift. Correct the target when needed instead of assuming it remains accurate through every frame. Give the video a stable data-popmelt-media-id when its source may change.
Tracks and human edits are persisted through Core. Saved motion remains replayable after See is disabled; browser-local storage is the default unless you supply videoAnnotation.store.
Experimental hover discovery
See can also discover a video object from the point under your pointer using a SAM-compatible encoder and point-prompted mask decoder. This is experimental, separate from ordinary image recognition, and must be explicitly configured:
const vision = configureSee({
videoSegmentation: {
device: 'webgpu',
dtype: 'fp16',
},
videoTracking: {
backboneUrl: '/models/nanotrack-backbone.onnx',
headUrl: '/models/nanotrack-head.onnx',
},
});
The default segmentation model is Xenova/slimsam-77-uniform. Its FP16 encoder and decoder are about 21 MB together, fetched on the first eligible hover and cached by the browser—not bundled in See. WebGPU with FP16 is the tested path; WASM is a compatibility fallback, not a promise of instant hover response.
Manual box creation remains the fallback. Cross-origin video still needs CORS. The See README covers self-hosting and advanced segmentation settings.
Next: Creating annotations with the Chat tool or Annotating static objects in images.
Something you want to improve?
Leave a comment