SAM Audio versus fixed stems: different inputs and goals
Distinguish Meta’s text, visual, and time-span prompts from a fixed six-stem music workflow.
Meta's SAM Audio uses text, visual, and time-span prompts to separate target sound from residual sound. It addresses general sound, speech, and music. See the official research page.
How this differs from fixed stems
A fixed set such as drums, bass, and vocals differs from requesting a particular sound in a time span. Choose according to how you can specify the source, rather than assuming either method always has better quality.
Design a useful test
Mix two recordings you control, keeping the originals as references. Change target level, overlap, or stereo placement one at a time. For a real recording without isolated references, listen for both unwanted leakage and missing target details.
Relationship to LA Studio
Do not describe LA Studio's Standard six-stem tool as SAM Audio. Its current Standard separation uses a Demucs-family music workflow. SAM Audio has its own distribution and runtime requirements; check Meta's official materials. This page does not announce an integration into LA Studio.
Judge both outputs
Check unwanted sound in the target and target leakage in the residual. Listen to quiet endings and decays as well as loud sections. For fixed instrument practice parts, see the six-stem guide.