Annotating static objects in images
Add browser-local object recognition to Core's image annotation workflow.
See identifies objects inside images using the same Chat interactions as the rest of the page. It extends Core; it does not replace the provider. Manual boxes on images and paused video only need Core—no recognition or segmentation model. See the media section of Creating annotations with the Chat tool.
Add See
In an existing Core project, ask your agent:
Add @popmelt.com/see to this Popmelt Core project
Or install it manually:
npm install @popmelt.com/see
See requires Core 0.17.0 or newer. Add it to the existing provider's extensions, preserving other extensions and props:
import { PopmeltProvider } from '@popmelt.com/core';
import { see } from '@popmelt.com/see';
<PopmeltProvider extensions={[see]}>{children}</PopmeltProvider>;
Work with recognized objects
In Chat mode, hover an object to reveal its label and bounds. Single-click to leave a comment; double-click to rename or resize it. Drag over an unrecognized object to add a box manually. Right-click an object's label to delete it.
All image elements are eligible by default. Use Core's imageAnnotation.selector to narrow discovery, and add a stable data-popmelt-media-id when URLs may change or an image appears more than once.
Recognition is model-dependent and can miss or mislabel objects. Manual annotation remains available while recognition loads or when See is disabled.
Choose recognition settings
Use configureSee instead of the default extension when you need different settings:
import { configureSee } from '@popmelt.com/see';
const vision = configureSee({
objectRecognition: {
minimumConfidence: 0.8,
runtime: { device: 'webgpu' },
},
});
<PopmeltProvider extensions={[vision]}>{children}</PopmeltProvider>;
Use a modern browser with WebAssembly support. WebGPU is preferred where available; compatible configurations can use WASM. Download size, latency, memory use, and available labels depend on the model and precision.
Local inference and model downloads
Inference runs in the browser. By default, model and runtime assets download from Hugging Face when first needed. Browser-local inference does not mean there are no external asset requests.
To self-host assets and disable remote model loading:
const vision = configureSee({
objectRecognition: {
model: 'detr-resnet-50',
runtime: {
allowRemoteModels: false,
localModelPath: '/models/',
wasmPaths: '/transformers-wasm/',
localFilesOnly: true,
},
},
});
Serve compatible model and runtime files at those paths. Remote images must permit CORS so See can read their pixels. Sending an annotation still passes the relevant evidence to your selected AI provider.
Keep objects without recognition
Remove See from extensions to turn recognition off. Saved objects and human edits remain available through Core. Results use browser-local storage by default; use Core's media stores for project-backed persistence.
Next: Annotating moving objects in video. Advanced model options are in the See README.
Something you want to improve?
Leave a comment