PaperScope
LIVE · 2026-09-15 05:40 UTC

Adapting Open-Weight MLLMs to Generate Point Prompts for Electron Microscopy Segmentation

Samia Mohinta, Albert Cardona

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.14080 v1
Category
Submitted
2026-09-12

Abstract

Promptable models such as microSAM segment electron microscopy (EM) images from point prompts, but automation requires generating prompts without user input. We ask whether open-weight multimodal large language models (MLLMs) can generate them from natural-language requests by returning coordinates to a frozen segmenter. To that end, we convert masks from three mitochondria datasets into training examples, pairing images and instructions with centroid coordinates, then train LoRA adapters while freezing the MLLM backbone and microSAM. We find that Qwen3-VL reaches segmentation AP$_{50}$ $0.736$ after supervised fine-tuning and reward optimization, up from $0.247$ without adaptation, while automatic prompt generation (APG) achieves $0.773$. In addition, two other MLLMs improve, reaching or exceeding APG. When compared with a supervised centroid-heatmap detector that reaches AP$_{50}$ $0.904$ for this mitochondria task, Qwen3-VL more closely matches the annotated point set and instance counts. Moreover, training on two public datasets transfers to an unseen third, while training on all three transfers to an independent EM volume. Robustness tests show stable performance under unseen formulations of the natural-language request, while the coordinates can be reused by a second segmenter. To our knowledge, this is the first feasibility study of open-weight MLLMs as EM point generators, providing an inspectable, language-directed link between localization and mask decoding.

Comment: Accepted at the BioImage Computing (BIC) Workshop at ECCV 2026

arXiv abs page · PDF · same-day batch