Understanding Automatic Mixing

A Subtask-Oriented Analysis of Two-Stage Mixing Systems

Jinjie Shi1, Wei Hua2, Kunzhu Xie3, Make Li4, Yuchen Liu5, Joshua Reiss1 1 Queen Mary University of London, United Kingdom  ·  2 Guangxi Arts University, China  ·  3 Wuhan University of Communications, China  ·  4 Xinghai Conservatory of Music, China  ·  5 University College London, United Kingdom Correspondence: jinjie.shi@qmul.ac.uk

Abstract

Automatic mixing transforms multitrack recordings into perceptually coherent, balanced, and aesthetically consistent mixes. However, in real-world production, this task is challenging due to large track counts, diverse instrumentation, and strong inter-track dependencies. Two-stage systems address this complexity by separating intra-group processing from inter-group mixing, yet it remains unclear whether their gains arise from stronger component models or from explicit task decomposition. We present a subtask-oriented analysis of automatic mixing through three controlled listening experiments. We investigate whether full-mix models transfer to intra-group mixing, whether downstream models compensate for grouping and loudness errors, and whether two-stage decomposition improves full-mix quality. Across three dense pop, rock, and metal excerpts, transfer differs between the evaluated models; inappropriate grouping causes clear downstream degradation, while altered loudness relationships have weaker and model-dependent effects. Both two-stage variants significantly outperform their corresponding single-stage baselines. These findings support explicit separation of local balance and global mix coordination as a useful design principle for automatic mixing.

Three Research Questions

RQ1  ·  Transfer

Can models trained for full mixing transfer to intra-group mixing? See Experiment 1 demos.

RQ2  ·  Compensation

Can downstream models compensate for incorrect grouping and loudness relationships? See Experiment 2a and 2b.

RQ3  ·  Decomposition

Does explicit two-stage decomposition improve full-mix quality? See Experiment 3 demos.

Analysis framework

Grouping functions, intra-group processors, and inter-group models can be varied independently. See the grouping-rules documentation.

Two-Stage Framework

A grouping function partitions the input multitrack into functional groups. Intra-group processing produces group-level stems, which are then combined by an inter-group model into the final mix. This contrasts with a monolithic full-mix model that generates the final mix directly from the raw tracks. The framework is used purely as a controlled analysis scaffold — it is not itself proposed as a new mixing system.

Models Evaluated

ModelRoleDescription
ELLIntra-groupEqual Local Loudness — training-free, rule-based intra-group balancing.
Diff-MSTFull-mix / intra-group / inter-groupPredicts gain, EQ, compression, and panning via a differentiable mixing console.
MEGAMIFull-mix / intra-group / inter-groupConditional generative model producing coordinated track-level effect representations.
NoMixControlShared preprocessing only, no further processing — unprocessed control condition.

Citing

A camera-ready citation will be added once the paper is published. In the meantime, please reference the GitHub repository:

https://github.com/SparrowReivun/TwoStageMixingAnalysis