audio
music
audio-tokenizer
rvq
minimax-music-3

RVQ encoder for MiniMax Music 3 (reverse distillation)

MiniMax Music 3 generates music as eight RVQ code streams at 25 Hz (one semantic stream with 16,384 codes, seven acoustic streams with 1,024) but ships no audio-to-code tokenizer. This encoder recovers it: it maps the frozen DAV autoencoder latents of any 44.1 kHz stereo audio to the eight codes per frame, so MiniMax Music 3 can be conditioned on a recording (reference-guided generation, continuation).

Folder Model
root the encoder: 169M parameters, 8-layer temporal transformer, depth decoder causal across codebooks
augmented/ the same model trained on degraded latent views; use it for lossy, band-limited or noisy audio

Code: rvq-reverse-distillation

from pathlib import Path

from rvq_ae.audio.io import load_audio
from rvq_ae.inference import CodeEncoder

codec = CodeEncoder.load(device="cuda", subfolder="augmented")   # or the root model
result = codec.encode(*load_audio(Path("song.flac")), topk=5)
result.codes                                               # [frames, 8]

Results

Condition replay cosine: the codes are teacher-forced through the official language model, depth decoder and condition encoder, and compared with the conditioning the generator used (1 = the generator's own codes; random codes 0.03, another track's codes 0.08).

Test set root augmented/
held-out generated tracks (130) 0.875 0.868
MM3-OOD, captions outside the training genres (240) 0.817 0.848
4 kHz low-pass input (32 held-out tracks) 0.493 0.831
MP3 at 32 kbit/s 0.530 0.816
pink noise at 10 dB SNR 0.586 0.767

Trained with jazz held out (three seeds), the same recipe encodes the unseen genre as well as the seen ones (0.859 against 0.865), with no prompt or lyric shared with training. On recorded music (Song Describer, MUSDB18-HQ), rendering MiniMax Music 3 from these codes preserves the recording's harmony (chroma 0.956), key (0.80) and tempo (0.85), and recovers 0.92 of the reachable conditioning against 0.91 on generated audio.

Which one to use: augmented/ was trained on degraded copies of every track (noise, lossy coding, filtering, reverberation, resampling) with the clean codes as targets. It costs 0.007 on clean generated audio and is far more robust to how real recordings arrive, so it is the better default for phone recordings, MP3s and anything that is not studio quality.

Training

Trained on the reverse distillation corpus (2,972 MiniMax Music 3 tracks with their sampled codes and the generator's top-50 distributions) with cross-entropy plus 0.25 KL to the generator's top-50, 20 epochs, AdamW 3e-4, effective batch 64, maximal update parametrisation, one RTX 3090, seed 1, published split; augmented/ also samples each window from the clean latents or one of two degraded copies.

Citation

@article{ashraf2026reverse,
  title  = {Recovering a Withheld Audio Tokenizer from Its Generator: Reverse Distillation of the
            {MiniMax Music 3} Encoder, In and Out of Distribution},
  author = {Munaza Ashraf and bghira},
  year   = {2026},
  url    = {https://github.com/MunazaAshraf10/rvq-reverse-distillation}
}

The weights follow the MiniMax Music 3 Community License, like anything generated with MiniMax Music 3.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Datasets used to train Munaza10/rvq-reverse-distillation