Promptable segmentation: turn an outline into a useful image mask
Evaluate image masks, fine edges, and editing effort.

A photograph contains objects, shadows, reflections, gaps, and surfaces that overlap. Asking an AI system to select one object requires more than recognizing its name. The system must decide which pixels belong to the requested region and which belong elsewhere. That decision becomes especially interesting when the request arrives as a click or a rough box instead of a carefully drawn outline.
Promptable segmentation makes this interaction useful: a person supplies a small amount of guidance, and a model proposes a mask. The proposal can save substantial manual work, but the quality of the result depends on what the mask will be used for. A selection that looks convincing in a colored overlay may still be unsuitable for compositing, measurement, or repeated editing.
Define the selected object before its boundary
Imagine a hypothetical designer preparing a photograph of a houseplant for a catalog. The desired selection could mean the ceramic pot, the leaves, the entire plant and pot, or everything in the foreground including its shadow. A point inside one leaf does not fully express these alternatives.
This ambiguity is part of the task, not necessarily a model failure. A useful interface lets the designer refine the instruction, inspect candidate masks, and undo an unsuccessful change. The design question is how quickly the person can reach the intended region with understandable controls.
Write down the selection policy for the workflow. Decide whether holes between leaves remain transparent, whether detached foliage counts, and whether shadows are selected separately. These choices also guide annotation and evaluation. Without them, reviewers may disagree while each follows a reasonable interpretation of the image.
What a prompt contributes
Alexander Kirillov and colleagues introduced Segment Anything around promptable segmentation, including spatial guidance such as points or boxes. The model produces masks from this guidance, and its design accounts for ambiguity in what a prompt might identify. The paper is a primary reference for the approach rather than a guarantee for every image domain.
Read the original research paper on arXiv
For the plant photograph, a foreground click says something about inclusion. An exclusion click can help distinguish the neighboring chair from a leaf. A box narrows the relevant area but does not necessarily settle every fine boundary inside it. Each interaction adds evidence about the intended selection.
Evaluate the interaction sequence, not just the first mask. If a model needs one extra click but produces a much more usable boundary, it may serve the designer better than one that looks impressive immediately and then becomes difficult to correct.
A mask is not an alpha matte
A binary mask classifies pixels as selected or unselected. An alpha matte represents degrees of opacity for compositing. Hair, fine fibers, motion blur, and transparent objects often require more than a clean division between inside and outside to look natural against a different background.
The distinction matters for the glass vase beside the plant. Selecting the vase's silhouette does not recover how the original background appears through the glass. Pasting the selected region onto a dark backdrop can therefore retain colors from the old scene even when the outline is correct.
Treat segmentation and compositing as connected but separate stages. The first can identify the region to work on; later processing may need edge refinement, color decontamination, or a dedicated matting method. A product should communicate these capabilities accurately rather than presenting every mask as a finished cutout.
Inspect the boundary at the intended scale
A small preview hides defects. Zoom in on thin stems, leaf tips, holes, and places where foreground and background have similar colors. Then inspect the output at its final display size. A boundary error that is obvious at extreme magnification may be irrelevant to a small thumbnail, while a halo may remain visible at any size.
Use several background colors during review. White can conceal pale fringes, and black can conceal dark ones. A checkerboard helps identify transparency but should not be the only inspection surface. Place the cutout into a realistic layout before deciding whether it is ready.
Preserve the original image and the editable mask separately. That allows a designer to correct an edge without repeatedly degrading the photograph through exports. Reversibility is a practical quality feature, especially when the model's first interpretation is close enough to be useful but still requires judgment.
Match metrics to the editing job
Intersection over union compares the overlap between a proposed region and a reference region. It is useful for evaluating masks, but a single area-based score can understate the importance of small structures. Losing a narrow stem may affect few pixels while visibly disconnecting a leaf from the plant.
For an editing workflow, also record boundary defects, missed holes, and how much human correction remains. Time to an acceptable result can reveal differences that a mask score misses. Define acceptable before timing the task so that reviewers are not quietly applying different standards.
Use a collection of challenging examples rather than one polished demonstration. Include clutter, soft edges, similar colors, partial occlusion, reflective surfaces, and small objects. Report results by these conditions so that a strong average does not conceal the cases your users encounter most often.
Avoid making the evaluation easier with hindsight
A prompt chosen after seeing a model's failure gives it information that an ordinary user might not have supplied initially. Interactive evaluation should distinguish a first attempt from later corrections. Both are useful measurements, but they answer different questions about the workflow.
For the plant example, record the initial click, every corrective click, any candidate-mask selection, and the stopping rule. Compare methods with the same interaction budget or report the effort needed to meet the same acceptance standard. This makes a favorable result easier to reproduce.
Also separate prompts supplied by a skilled evaluator from prompts produced by an automatic procedure. A professional retoucher may know exactly where to click to resolve an ambiguity. That success does not establish that a novice or a downstream automated pipeline will obtain the same outcome.
Video adds identity and time
A mask in one frame does not establish that the same object will remain selected throughout a clip. Objects move, change appearance, leave the frame, and pass behind other objects. Errors can spread if later predictions depend on earlier selections that were already wrong.
Nikhila Ravi and colleagues' SAM 2 extends promptable segmentation to images and videos, using a memory-based design to support video processing. Its paper offers a starting point for understanding the temporal problem. A deployment still needs evaluation on the kinds of motion, occlusion, and edits it will actually encounter.
Read the original research paper on arXiv
For a clip of the plant rotating on a stand, review more than the first and final frames. Check when a leaf crosses another leaf and when the pot passes behind a hand. Those transitions expose whether the selection tracks the intended object or merely follows a visually convenient region.
Separate perceived quality from measurement reliability
An attractive cutout and a reliable measurement mask have different acceptance criteria. A slightly expanded leaf outline may look fine in a catalog but distort an estimate of leaf area. A workflow intended for measurement must address scale, image geometry, annotation policy, and uncertainty in the boundary.
The model's confidence score does not independently validate the physical quantity being measured. Compare against references collected with a suitable procedure, and inspect systematic errors by object size and image condition. When the evidence is insufficient, preserve the uncertainty instead of turning a neat mask into an unjustified precise number.
This distinction also changes who should review the result. A designer can judge visual fit, while a measurement workflow needs someone who understands the quantity and collection method. The same model output can be helpful in both contexts without serving as the same kind of evidence.
Make corrections part of the saved work
Store the relationship between the source image, prompts, selected mask, and subsequent edits. If a source image is resized or rotated, coordinate handling must preserve the intended region. A forgotten transformation can make a saved click point at a different object when a project is reopened.
Keep review states explicit. A generated proposal, a manually corrected selection, and an approved export should not be indistinguishable files. This is particularly useful when several people collaborate or when a batch process sends uncertain cases to a human editor.
Promptable segmentation works best when it makes the person's intent easier to express and the result easier to inspect. Judge the full path from ambiguous image to accepted selection. The valuable outcome is an editable, appropriate mask that reduces real work while keeping the difficult boundaries visible.
Sources and rights
Segment Anything and SAM 2 are cited under their authors' copyrights and arXiv's non-exclusive distribution licence. Model weights, datasets, and software can have separate terms; the paper links do not grant rights to those artifacts. No source text, figures, or code are reproduced here.