PaperScope
LIVE · 2026-09-29 05:40 UTC

SCISSOR: Score-Conditioned Instrument Source Separation for Orchestral Recordings

Yiheng Lu, Hao-Wen Dong

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.33265 v1
Submitted
2026-09-27

Abstract

Orchestral separation recovers instrument sections from mixtures in which shared pitches, harmonics, and timbres obscure source identity. An aligned score provides instrument labels, note pitches, and activity times. A score-informed approach appends piano rolls to audio features before mask prediction. We introduce SCISSOR (Score-Conditioned Instrument Source Separation for Orchestral Recordings), which uses the score to form a frame-wise query for each source. Each query matches a shared audio representation, and a softmax over instrument and background slots jointly assigns overlapping time-frequency evidence. The queries retain instrument identity even when notes are missing from the score. After training on SynthSOD and a small set of URMP and PHENICX-Anechoic recordings, SCISSOR achieves the highest average SDR on held-out real recordings. With SynthSOD-only training, it leads on SynthSOD and zero-shot PHENICX-Anechoic, and improves on its audio-only control on zero-shot URMP. SCISSOR also degrades less under score corruption than the evaluated score-based baselines.

Comment: 5 pages, 1 figure, 3 tables. Submitted to ICASSP 2027

arXiv abs page · PDF · same-day batch