open-rvq-encoder-minimax-music3-41m โ pooled-corpus fine-tune (v3)
A fine-tune of SimpleTuner/open-rvq-encoder-minimax-music3-41m-v1
(checkpoint-17500) on a pooled reverse-distillation corpus 3.6x the size of the original:
the original 2,837 train tracks plus 8,122 new self-distillation tracks from
Mothersuperior/minimax-music3-rvq-distill-corpus-8k
(diversity-engineered captions: 355 genres with per-genre BPM/meter/key/ensemble priors, parity-with-floors sampling,
~31% instrumental, 30-150 s durations, original LLM lyrics). Architecture, loss (CE + 0.25 KL vs teacher top-50 @ T=1),
muP config, and training script are unchanged from the base โ only the data changed.
Benchmark (same harness, same 130 exact-alignment holdout records as the base model)
Running the base checkpoint through this harness reproduces its published numbers (0.762 / 0.411) exactly.
| metric | base ckpt-17500 | this model | delta |
|---|---|---|---|
| replay conditioning cosine (mean, ceiling 0.9999) | 0.7620 | 0.7843 | +0.022 |
| semantic top-1 (16,384-way) | 0.4109 | 0.4556 | +4.5 pts |
| semantic top-5 | 0.7844 | 0.8322 | +4.8 pts |
| acoustic top-1 (mean of 7 heads) | 0.0718 | 0.0853 | +19% rel |
| acoustic top-5 | 0.2097 | 0.2407 | +15% rel |
Raw eval outputs are in evaluation/.
Training
Warm-started from the base checkpoint-17500 weights (no optimizer state), then the base recipe
(--learning_rate 3e-4, cosine cycles, batch 64 windows of 128 frames) for 12 + 24 epochs on the pooled
10,959-track train split, validated every 500 steps on the base model's untouched 135-track holdout.
Best checkpoint selected by holdout semantic top-1 (step 54,500 of the continuation run). Single RTX PRO 6000.
Loading
Same format as the base model: rvq_encoder.safetensors + rvq_encoder_config.json + mup_base_shapes.bsh,
loaded with the RVQEncoderConfig / MiniMaxMusicRVQEncoder classes from SimpleTuner's
scripts/train_minimax_music_rvq_encoder.py (branch script/train-minimax-music-rvq-encoder).
Input: 128-ch DAV latents (~3.445 per 25 Hz semantic frame). Output: 8 codebook heads (c0 16,384; d1-d7 1,024).
Use is subject to the MiniMax Music 3 model terms and the reverse-distillation dataset terms.