Add Ko-fi support blurb
Browse files
README.md
CHANGED
|
@@ -1,245 +1,249 @@
|
|
| 1 |
-
---
|
| 2 |
-
license: cc-by-nc-4.0
|
| 3 |
-
base_model:
|
| 4 |
-
- m-a-p/YuE2-3B
|
| 5 |
-
- Comfy-Org/YuE2
|
| 6 |
-
pipeline_tag: text-to-audio
|
| 7 |
-
language:
|
| 8 |
-
- en
|
| 9 |
-
- jam
|
| 10 |
-
tags:
|
| 11 |
-
- lora
|
| 12 |
-
- yue2
|
| 13 |
-
- music-generation
|
| 14 |
-
- text-to-music
|
| 15 |
-
- song-generation
|
| 16 |
-
- reggae
|
| 17 |
-
- roots-reggae
|
| 18 |
-
- dancehall
|
| 19 |
-
- steppers
|
| 20 |
-
- comfyui
|
| 21 |
-
- fs_audio
|
| 22 |
-
---
|
| 23 |
-
|
| 24 |
-
# MLTNT β Militant Roots Reggae LoRAs for YuE2
|
| 25 |
-
|
| 26 |
-
Four artist-style LoRAs that push [YuE2-3B](https://huggingface.co/m-a-p/YuE2-3B) into **modern militant roots reggae**: dark raspy male patois vocals, steppers and one-drop grooves, deep sub bass, bubbling Hammond, nyabinghi drums, horn stabs, dub sirens and spring reverb. Conscious, apocalyptic, anthemic.
|
| 27 |
-
|
| 28 |
-
Each file patches **both halves** of YuE2 in one go: the autoregressive planner (writes the score, decides the arrangement and the vocal lines) and the flow-matching decoder (the sound). Trigger word for all three: **`mltnt`**.
|
| 29 |
-
|
| 30 |
-
| | LoRA | File | Character | Start here |
|
| 31 |
-
|---|---|---|---|---|
|
| 32 |
-
| π₯ | **MLTNT Frontline** | `mltnt_frontline.safetensors` | The newest. Most distinctive voice tonality of the four, and the one trained with extra short excerpts of fast-delivery verses, so it is the pick for rapid-fire lyrics (see *Rapid-fire verses* below). Works on both recipes. | clip 1.0 / model 1.0, cfg 1.0 |
|
| 33 |
-
| π€ | **MLTNT Fusion** | `mltnt_fusion.safetensors` | Widest palette. Reggae hip-hop fusion: boom-bap over one-drop, deejay toasting, sampled roots hooks. Happily moves off the 76 BPM roots grid. | clip 1.0 / model 1.5, cfg 1.4 |
|
| 34 |
-
| π₯ | **MLTNT Steppers** | `mltnt_steppers.safetensors` | The flagship. Tight, hard-hitting steppers with a big anthemic chorus. Most reliable vocal density. | clip 1.0 / model 1.0, cfg 1.0 |
|
| 35 |
-
| πΏ | **MLTNT Roots** | `mltnt_roots.safetensors` | The purist. Straightest roots timbre, least steered toward any one arrangement. | clip 1.0 / model 1.0, cfg 1.0 |
|
| 36 |
-
|
| 37 |
-
## Listen
|
| 38 |
-
|
| 39 |
-
All demos use the same original lyric, seed 7, 32 steps `dpm_2` / `sgm_uniform`, no post-processing.
|
| 40 |
-
|
| 41 |
-
**MLTNT Frontline** β baseline recipe, prompt `prompts/steppers_baseline.txt`, **dense lyric** (verses written at ~17 words per line; planner chose a 160 BPM grid in F minor, 3:57, roughly double the words per minute of the standard lyric)
|
| 42 |
-
|
| 43 |
-
<audio controls src="https://huggingface.co/becausereasons/yue2-mltnt-militant-reggae/resolve/main/demos/mltnt_frontline__demo.mp3"></audio>
|
| 44 |
-
|
| 45 |
-
The same model, prompt, recipe and seed with the **standard 8-word-line lyric** β only the lyric differs, so this pair is the rapid-fire comparison (135 BPM F minor, 3:02):
|
| 46 |
-
|
| 47 |
-
<audio controls src="https://huggingface.co/becausereasons/yue2-mltnt-militant-reggae/resolve/main/demos/mltnt_frontline__standard_lyric.mp3"></audio>
|
| 48 |
-
|
| 49 |
-
Frontline on the Fusion recipe with the dense lyric, prompt `prompts/fusion_hiphop.txt` (160 BPM, 3:21):
|
| 50 |
-
|
| 51 |
-
<audio controls src="https://huggingface.co/becausereasons/yue2-mltnt-militant-reggae/resolve/main/demos/mltnt_frontline__fusion_dense.mp3"></audio>
|
| 52 |
-
|
| 53 |
-
Frontline on the Fusion recipe with the trap-dub prompt, `prompts/trap_dub.txt` with the phrase `rapid-fire double-time deejay flow` added to the vocal description, standard lyric (150 BPM F minor, 3:00). The cue phrase does not speed the delivery up (see *Rapid-fire verses*), but this is the prompt family that gives Frontline its wildest, most trap-leaning plans:
|
| 54 |
-
|
| 55 |
-
<audio controls src="https://huggingface.co/becausereasons/yue2-mltnt-militant-reggae/resolve/main/demos/mltnt_frontline__trap_dub.mp3"></audio>
|
| 56 |
-
|
| 57 |
-
Three renders kept for reference from **earlier Frontline builds** (not published as weights). First, the trap-dub prompt with the **dense lyric** on the build that preceded the current one, i.e. the same data without the fast-verse excerpts (160 BPM F minor, 3:13, about 110 words per minute):
|
| 58 |
-
|
| 59 |
-
<audio controls src="https://huggingface.co/becausereasons/yue2-mltnt-militant-reggae/resolve/main/demos/mltnt_frontline__trap_dub_dense_previous_build.mp3"></audio>
|
| 60 |
-
|
| 61 |
-
Second, the same earlier build on the Fusion recipe with a style prompt written in the strict native caption order (language β genre β vocal β instruments β mood β BPM β production, the shape of `prompts/steppers_native_format.txt`) and a different, longer original lyric sheet. The planner wrote a full 4:38 song at 145 BPM in E minor and sang it through:
|
| 62 |
-
|
| 63 |
-
<audio controls src="https://huggingface.co/becausereasons/yue2-mltnt-militant-reggae/resolve/main/demos/mltnt_frontline__native_caption_previous_build.mp3"></audio>
|
| 64 |
-
|
| 65 |
-
Third, the original Frontline demo: the plain trap-dub prompt on the first Frontline build (160 BPM F minor, 88 % of vocal bars sung):
|
| 66 |
-
|
| 67 |
-
<audio controls src="https://huggingface.co/becausereasons/yue2-mltnt-militant-reggae/resolve/main/demos/mltnt_frontline__trap_dub_previous_build.mp3"></audio>
|
| 68 |
-
|
| 69 |
-
**MLTNT Fusion** β Fusion recipe, prompt `prompts/fusion_hiphop.txt` (planner chose a 150 BPM grid in Gβ― minor, three verse-chorus rounds and an outro)
|
| 70 |
-
|
| 71 |
-
<audio controls src="https://huggingface.co/becausereasons/yue2-mltnt-militant-reggae/resolve/main/demos/mltnt_fusion__demo.mp3"></audio>
|
| 72 |
-
|
| 73 |
-
**MLTNT Steppers** β baseline recipe, prompt `prompts/steppers_baseline.txt`
|
| 74 |
-
|
| 75 |
-
<audio controls src="https://huggingface.co/becausereasons/yue2-mltnt-militant-reggae/resolve/main/demos/mltnt_steppers__demo.mp3"></audio>
|
| 76 |
-
|
| 77 |
-
**MLTNT Roots** β baseline recipe, prompt `prompts/steppers_baseline.txt` (planner stayed on the 76 BPM roots grid in A minor, 4:59 long, 149 of 190 vocal bars sung)
|
| 78 |
-
|
| 79 |
-
<audio controls src="https://huggingface.co/becausereasons/yue2-mltnt-militant-reggae/resolve/main/demos/mltnt_roots__demo.mp3"></audio>
|
| 80 |
-
|
| 81 |
-
## Quick start (ComfyUI)
|
| 82 |
-
|
| 83 |
-
These LoRAs are in the fused-key layout that Comfy's YuE2 implementation uses (`text_encoders.*` for the planner, `diffusion_model.*` for the decoder). They were trained with, and load through, the **[FS_Audio Suite](https://github.com/KytraScript/ComfyUI-FS_Audio_Suite)** node pack.
|
| 84 |
-
|
| 85 |
-
1. ComfyUI β₯ v0.36.0 (native YuE2 support) and the FS_Audio Suite custom node pack.
|
| 86 |
-
2. Base model: `yue2_3b_bf16.safetensors` from [Comfy-Org/YuE2](https://huggingface.co/Comfy-Org/YuE2) in `models/checkpoints/`.
|
| 87 |
-
3. Drop one `mltnt_*.safetensors` into `models/loras/`.
|
| 88 |
-
4. Chain the nodes:
|
| 89 |
-
|
| 90 |
-
```
|
| 91 |
-
π§© FS_Audio Lora Loader ββlorasβββΆ π€ FS_Audio Model Loader ββpipeβββΆ π΅ FS_Audio Sampler βββΆ πΏ FS_Audio Output
|
| 92 |
-
lora_name = mltnt_steppers.safetensors yue2_checkpoint = yue2_3b_bf16 style = <prompt, starts with "mltnt,">
|
| 93 |
-
strength_clip = 1.0 (planner) melody_transcriber = none lyrics = <tagged lyric blocks>
|
| 94 |
-
strength_model = 1.0 (decoder) score_mode = full
|
| 95 |
-
```
|
| 96 |
-
|
| 97 |
-
Sampler settings used for every demo:
|
| 98 |
-
|
| 99 |
-
| widget | value |
|
| 100 |
-
|---|---|
|
| 101 |
-
| steps / sampler / scheduler | 32 / `dpm_2` / `sgm_uniform` |
|
| 102 |
-
| score_mode | `full` |
|
| 103 |
-
| song_length_cap | 360 (a full song; 150 truncates mid-way) |
|
| 104 |
-
| repetition_penalty | 1.2 |
|
| 105 |
-
|
| 106 |
-
Two proven recipes:
|
| 107 |
-
|
| 108 |
-
| | Baseline (Steppers / Roots / Frontline) | Fusion |
|
| 109 |
-
|---|---|---|
|
| 110 |
-
| strength_clip (planner) | 1.0 | 1.0 |
|
| 111 |
-
| strength_model (decoder) | 1.0 | **1.5** |
|
| 112 |
-
| Weirdness (cfg) | 1.0 | **1.4** |
|
| 113 |
-
| score_temperature | 0.7 | **0.9** |
|
| 114 |
-
| music_temperature | 1.0 | **1.2** |
|
| 115 |
-
|
| 116 |
-
Headless, the same graph as an API prompt:
|
| 117 |
-
|
| 118 |
-
```python
|
| 119 |
-
prompt = {
|
| 120 |
-
"1": {"class_type": "FSAudioLoraLoader", "inputs": {"lora_name": "mltnt_steppers.safetensors", "strength_model": 1.0, "strength_clip": 1.0}},
|
| 121 |
-
"2": {"class_type": "FSAudioModelLoader", "inputs": {"yue2_checkpoint": "yue2_3b_bf16.safetensors", "melody_transcriber": "none", "loras": ["1", 0]}},
|
| 122 |
-
"3": {"class_type": "FSAudioSampler", "inputs": {"pipe": ["2", 0], "style": STYLE, "lyrics": LYRICS, "seed": 7,
|
| 123 |
-
"song_length_cap": 360, "score_mode": "full", "steps": 32, "sampler": "dpm_2", "scheduler": "sgm_uniform",
|
| 124 |
-
"Weirdness (cfg)": 1.0, "score_temperature": 0.7, "music_temperature": 1.0, "repetition_penalty": 1.2}},
|
| 125 |
-
"4": {"class_type": "FSAudioOutput", "inputs": {"song": ["3", 0], "score": ["3", 1], "info": ["3", 2], "filename_prefix": "mltnt"}},
|
| 126 |
-
}
|
| 127 |
-
```
|
| 128 |
-
|
| 129 |
-
## Prompting
|
| 130 |
-
|
| 131 |
-
### Style prompt
|
| 132 |
-
|
| 133 |
-
Start with the trigger, then write **one descriptive sentence** in this order: language β genre β vocal β instruments β mood β BPM β production. This is the format the planner responds to. A bare tag list or `[Intro]β¦[Outro]` scaffolding in the style field produces "odd, not reggae" output.
|
| 134 |
-
|
| 135 |
-
```
|
| 136 |
-
mltnt, Jamaican Patois English, modern militant roots reggae, dark raspy male vocal with heavy patois delivery, deep bassline with massive sub bass drops, skanking guitar, hammond organ bubble, nyabinghi percussion, horn stabs and dub sirens, apocalyptic conscious mood, 76 BPM, sparse hard-hitting arrangement, dub impact hits, harder drum impact, heavier low end, spring reverb and tape delay, space before an explosive anthemic chorus
|
| 137 |
-
```
|
| 138 |
-
|
| 139 |
-
Ready-made prompts in `prompts/`:
|
| 140 |
-
|
| 141 |
-
| file | what it gets you |
|
| 142 |
-
|---|---|
|
| 143 |
-
| `steppers_baseline.txt` | the approved baseline: sub bass drops, dub impact hits, steppers groove, explosive chorus |
|
| 144 |
-
| `steppers_native_format.txt` | the same brief rewritten in the strict caption order |
|
| 145 |
-
| `fusion_hiphop.txt` | reggae hip-hop fusion, deejay toasting, boom-bap over one-drop (best with **MLTNT Fusion**) |
|
| 146 |
-
| `trap_dub.txt` | trap and dubstep low end, gospel-stack chorus, tape-stop drops, 80 BPM written but the planner lands at 150β160. Behind three of the Frontline clips; results swing from great to chaotic between seeds, so treat it as the wild card |
|
| 147 |
-
| `trap_dubstep_fusion.txt` | militant roots fused with trap and dubstep low end, halftime breakdowns, gang-vocal chorus |
|
| 148 |
-
|
| 149 |
-
Production descriptors ("boom-bap", "deejay toasting", "steppers") move the planner's tempo grid and key far more than the numeric BPM in the same prompt.
|
| 150 |
-
|
| 151 |
-
### Lyrics
|
| 152 |
-
|
| 153 |
-
Tagged blocks, 4 lines per verse works best. Use `[Verse]`, `[Pre-Chorus]`, `[Chorus]`, `[Bridge]`, `[Outro]`. **Do not add an empty `[Intro]`** β the planner will happily write an 18-bar instrumental intro on its own. Shape:
|
| 154 |
-
|
| 155 |
-
```
|
| 156 |
-
[Verse]
|
| 157 |
-
four lines, roughly eight words each
|
| 158 |
-
|
| 159 |
-
[Verse]
|
| 160 |
-
four more lines
|
| 161 |
-
|
| 162 |
-
[Pre-Chorus]
|
| 163 |
-
two short lines that build tension
|
| 164 |
-
|
| 165 |
-
[Chorus]
|
| 166 |
-
four lines, the hook repeated at least twice across the song
|
| 167 |
-
|
| 168 |
-
[Bridge]
|
| 169 |
-
a chant or a two-word call, then four lines
|
| 170 |
-
|
| 171 |
-
[Outro]
|
| 172 |
-
two to four closing lines
|
| 173 |
-
```
|
| 174 |
-
|
| 175 |
-
**Rapid-fire verses.** The deejay's delivery speed comes from the *lyric*, not from the prompt. Words like "rapid-fire", "double-time" or "fast flow" in the style prompt do not speed the vocal up (tested: they mostly push the planner toward a shorter, trap-flavoured plan). What works is line density: write the verse at **15β17 words per line**, eight lines to the verse, and keep the chorus at the normal 7β8 words so the contrast lands. On the same prompt and seed that roughly doubles the words per minute and the LoRAs pack the lines intelligibly instead of overrunning them. Shape of a fast verse line:
|
| 176 |
-
|
| 177 |
-
```
|
| 178 |
-
Dem seh the future bright but mi seh where the people stand, where the promise and the plan
|
| 179 |
-
```
|
| 180 |
-
|
| 181 |
-
Keep the syllables simple (one- and two-syllable words); dense lines full of long words still garble. The two Frontline clips above are this exact comparison: same model, prompt, recipe and seed, standard lyric vs dense lyric.
|
| 182 |
-
|
| 183 |
-
A 2-word pre-chorus and a chant bridge give the planner clear pacing cues. Repeat the `[Chorus]` block verbatim wherever you want it sung again.
|
| 184 |
-
|
| 185 |
-
### Which knob does what
|
| 186 |
-
|
| 187 |
-
- **strength_model (decoder)** is the lever to push. 1.0 β 1.5 makes the sound darker, rougher and more present without touching the writing.
|
| 188 |
-
- **strength_clip (planner) above 1.0 collapses the vocal.** At 1.25 the planner wrote 1 sung bar against 424 rests. Keep it at 1.0.
|
| 189 |
-
- **Weirdness (cfg)** changes only the sound, not the arrangement. 1.0 is clean, 1.4 is the Fusion sweet spot, 1.7 gets noticeably wilder.
|
| 190 |
-
- **music_temperature / score_temperature** are what actually change the writing (tempo grid, key, section lengths). 1.2 / 0.9 gave the most adventurous still-coherent plans.
|
| 191 |
-
- **Seeds** can end a song early (seed 44 stopped at 4:30 with a 360 s cap). That is the seed, not the cap.
|
| 192 |
-
|
| 193 |
-
## Training
|
| 194 |
-
|
| 195 |
-
Trained on a single RTX 5090 with the FS_Audio Suite **Artist Trainer** on Comfy's own YuE2 weights. One run trains the planner LoRA and the decoder LoRA together; the file also carries full `vae2llm` / `llm2vae` projection diffs.
|
| 196 |
-
|
| 197 |
-
| | Steppers | Fusion | Roots | Frontline |
|
| 198 |
-
|---|---|---|---|---|
|
| 199 |
-
| checkpoint | step 300 (final) | **step 200** of 350 | step 300 (final) | step 350 (final) |
|
| 200 |
-
| planner / decoder steps | 300 / 800 | 350 / 1000 | 300 / 800 | 350 / 600 |
|
| 201 |
-
| final artist (planner) loss | 4.650 | 5.280 | 4.471 | 4.853 |
|
| 202 |
-
| decoder loss at checkpoint | 1.191 | 1.219 | 1.151 | 1.280 |
|
| 203 |
-
|
| 204 |
-
Shared hyper-parameters: rank 64 planner / rank 32 decoder, LR 3e-5 planner / 4e-5 decoder / 2e-5 I/O projections, artist fraction 0.5 vs regularizer pack, KL 0.1, batch 2 songs, 8192 max tokens, 750-frame windows, EMA 0.99, score-first 0, no transcription.
|
| 205 |
-
|
| 206 |
-
**Why step 200 for Fusion.** The trainer names its `_best` checkpoint by *planner* loss, which keeps falling. The decoder bottomed at steps 150β200 and then overfit hard (1.22 β 1.58 by step 350); the `_best` file sounds burnt. Pick checkpoints from the decoder curve. Frontline was trained with the decoder capped at 600 steps; its decoder loss bottomed at step 200 and drifted up by only 0.002 by step 350, so the final checkpoint is used. Frontline's set also adds eight short excerpts (20β45 s) of the fastest-delivery verses, each with only its own lyric lines β that did not make speed promptable from the caption, but it is where its voice comes from.
|
| 207 |
-
|
| 208 |
-
### Captioning your own dataset
|
| 209 |
-
|
| 210 |
-
The planner learns from the captions as much as from the audio, so the caption format decides whether your prompts work later. What worked here:
|
| 211 |
-
|
| 212 |
-
- **One descriptive sentence per song, trigger first**, in the same order you will prompt in: language β genre β vocal β instruments β mood β BPM β production. Example shape: `mltnt, Jamaican Patois English, modern militant roots reggae, dark raspy male vocal with heavy patois delivery, deep bassline, skanking guitar, bubbling hammond organ, nyabinghi percussion, horn stabs and dub sirens, urgent conscious mood, 76 BPM, sparse hard-hitting arrangement, spring reverb and dub delays`.
|
| 213 |
-
- **No tag lists, no section scaffolding** (`[Intro]β¦[Outro]`) in the caption. Tag-list captions produced a model that responds to tag-list prompts and writes odd, un-reggae plans.
|
| 214 |
-
- **Measure BPM, don't guess.** `librosa.beat.beat_track` on each file, rounded, written as `NN BPM` near the end of the sentence.
|
| 215 |
-
- A vision-language model can draft the captions from the audio, but **check the vocal gender by hand**: high male registers get labelled "female" often enough that every batch needs a grep-and-fix pass. A wrong gender in the caption shows up as a wrong voice at inference.
|
| 216 |
-
- **Lyrics as tagged blocks** in the same format you will prompt with: `[Verse]`, `[Pre-Chorus]`, `[Chorus]`, `[Bridge]`, `[Outro]`, patois spelling kept as sung, no empty intro tag. Transcribe at high confidence or leave the line out; a mis-heard lyric teaches the planner the wrong syllable count for the bar.
|
| 217 |
-
- **Audio prep**: FLAC, trimmed to β€ 320 s with a fade. MP3 rips with a damaged leading frame crash the dataset builder ("Header missing"), so transcode with `ffmpeg -err_detect ignore_err` first.
|
| 218 |
-
- **Trainer flags**: `score-first 0` and no automatic transcription. Letting the trainer transcribe melodies itself put the vocal melody in the instrument voice and rests in the vocal voice, and the decoder loss started climbing after step 200.
|
| 219 |
-
|
| 220 |
-
## Known limitations
|
| 221 |
-
|
| 222 |
-
- Vocals are patois-flavoured English only.
|
| 223 |
-
- Dense multi-syllable lines overrun their bars and come out garbled. Budget roughly words Γ· 2 seconds per line at ~120 wpm and shorten the line rather than fighting pronunciation.
|
| 224 |
-
- With the 150 s cap the planner still writes full-length intros and interludes, so the song truncates before the second chorus. Use 360.
|
| 225 |
-
- The planner occasionally writes a plan with almost no sung bars. Change the seed; do not raise planner strength.
|
| 226 |
-
- No instrumental-only mode is baked in; the LoRAs assume a lyric.
|
| 227 |
-
|
| 228 |
-
## Files
|
| 229 |
-
|
| 230 |
-
```
|
| 231 |
-
mltnt_steppers.safetensors 177 MB planner + decoder LoRA (bf16 weights, fp32 projection diffs)
|
| 232 |
-
mltnt_fusion.safetensors 177 MB
|
| 233 |
-
mltnt_roots.safetensors 177 MB
|
| 234 |
-
mltnt_frontline.safetensors 177 MB
|
| 235 |
-
demos/ mp3 renders (192 kbps from the FLAC masters), seed 7 throughout
|
| 236 |
-
prompts/ style prompts
|
| 237 |
-
```
|
| 238 |
-
|
| 239 |
-
##
|
| 240 |
-
|
| 241 |
-
|
| 242 |
-
|
| 243 |
-
|
| 244 |
-
|
| 245 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: cc-by-nc-4.0
|
| 3 |
+
base_model:
|
| 4 |
+
- m-a-p/YuE2-3B
|
| 5 |
+
- Comfy-Org/YuE2
|
| 6 |
+
pipeline_tag: text-to-audio
|
| 7 |
+
language:
|
| 8 |
+
- en
|
| 9 |
+
- jam
|
| 10 |
+
tags:
|
| 11 |
+
- lora
|
| 12 |
+
- yue2
|
| 13 |
+
- music-generation
|
| 14 |
+
- text-to-music
|
| 15 |
+
- song-generation
|
| 16 |
+
- reggae
|
| 17 |
+
- roots-reggae
|
| 18 |
+
- dancehall
|
| 19 |
+
- steppers
|
| 20 |
+
- comfyui
|
| 21 |
+
- fs_audio
|
| 22 |
+
---
|
| 23 |
+
|
| 24 |
+
# MLTNT β Militant Roots Reggae LoRAs for YuE2
|
| 25 |
+
|
| 26 |
+
Four artist-style LoRAs that push [YuE2-3B](https://huggingface.co/m-a-p/YuE2-3B) into **modern militant roots reggae**: dark raspy male patois vocals, steppers and one-drop grooves, deep sub bass, bubbling Hammond, nyabinghi drums, horn stabs, dub sirens and spring reverb. Conscious, apocalyptic, anthemic.
|
| 27 |
+
|
| 28 |
+
Each file patches **both halves** of YuE2 in one go: the autoregressive planner (writes the score, decides the arrangement and the vocal lines) and the flow-matching decoder (the sound). Trigger word for all three: **`mltnt`**.
|
| 29 |
+
|
| 30 |
+
| | LoRA | File | Character | Start here |
|
| 31 |
+
|---|---|---|---|---|
|
| 32 |
+
| π₯ | **MLTNT Frontline** | `mltnt_frontline.safetensors` | The newest. Most distinctive voice tonality of the four, and the one trained with extra short excerpts of fast-delivery verses, so it is the pick for rapid-fire lyrics (see *Rapid-fire verses* below). Works on both recipes. | clip 1.0 / model 1.0, cfg 1.0 |
|
| 33 |
+
| π€ | **MLTNT Fusion** | `mltnt_fusion.safetensors` | Widest palette. Reggae hip-hop fusion: boom-bap over one-drop, deejay toasting, sampled roots hooks. Happily moves off the 76 BPM roots grid. | clip 1.0 / model 1.5, cfg 1.4 |
|
| 34 |
+
| π₯ | **MLTNT Steppers** | `mltnt_steppers.safetensors` | The flagship. Tight, hard-hitting steppers with a big anthemic chorus. Most reliable vocal density. | clip 1.0 / model 1.0, cfg 1.0 |
|
| 35 |
+
| πΏ | **MLTNT Roots** | `mltnt_roots.safetensors` | The purist. Straightest roots timbre, least steered toward any one arrangement. | clip 1.0 / model 1.0, cfg 1.0 |
|
| 36 |
+
|
| 37 |
+
## Listen
|
| 38 |
+
|
| 39 |
+
All demos use the same original lyric, seed 7, 32 steps `dpm_2` / `sgm_uniform`, no post-processing.
|
| 40 |
+
|
| 41 |
+
**MLTNT Frontline** β baseline recipe, prompt `prompts/steppers_baseline.txt`, **dense lyric** (verses written at ~17 words per line; planner chose a 160 BPM grid in F minor, 3:57, roughly double the words per minute of the standard lyric)
|
| 42 |
+
|
| 43 |
+
<audio controls src="https://huggingface.co/becausereasons/yue2-mltnt-militant-reggae/resolve/main/demos/mltnt_frontline__demo.mp3"></audio>
|
| 44 |
+
|
| 45 |
+
The same model, prompt, recipe and seed with the **standard 8-word-line lyric** β only the lyric differs, so this pair is the rapid-fire comparison (135 BPM F minor, 3:02):
|
| 46 |
+
|
| 47 |
+
<audio controls src="https://huggingface.co/becausereasons/yue2-mltnt-militant-reggae/resolve/main/demos/mltnt_frontline__standard_lyric.mp3"></audio>
|
| 48 |
+
|
| 49 |
+
Frontline on the Fusion recipe with the dense lyric, prompt `prompts/fusion_hiphop.txt` (160 BPM, 3:21):
|
| 50 |
+
|
| 51 |
+
<audio controls src="https://huggingface.co/becausereasons/yue2-mltnt-militant-reggae/resolve/main/demos/mltnt_frontline__fusion_dense.mp3"></audio>
|
| 52 |
+
|
| 53 |
+
Frontline on the Fusion recipe with the trap-dub prompt, `prompts/trap_dub.txt` with the phrase `rapid-fire double-time deejay flow` added to the vocal description, standard lyric (150 BPM F minor, 3:00). The cue phrase does not speed the delivery up (see *Rapid-fire verses*), but this is the prompt family that gives Frontline its wildest, most trap-leaning plans:
|
| 54 |
+
|
| 55 |
+
<audio controls src="https://huggingface.co/becausereasons/yue2-mltnt-militant-reggae/resolve/main/demos/mltnt_frontline__trap_dub.mp3"></audio>
|
| 56 |
+
|
| 57 |
+
Three renders kept for reference from **earlier Frontline builds** (not published as weights). First, the trap-dub prompt with the **dense lyric** on the build that preceded the current one, i.e. the same data without the fast-verse excerpts (160 BPM F minor, 3:13, about 110 words per minute):
|
| 58 |
+
|
| 59 |
+
<audio controls src="https://huggingface.co/becausereasons/yue2-mltnt-militant-reggae/resolve/main/demos/mltnt_frontline__trap_dub_dense_previous_build.mp3"></audio>
|
| 60 |
+
|
| 61 |
+
Second, the same earlier build on the Fusion recipe with a style prompt written in the strict native caption order (language β genre β vocal β instruments β mood β BPM β production, the shape of `prompts/steppers_native_format.txt`) and a different, longer original lyric sheet. The planner wrote a full 4:38 song at 145 BPM in E minor and sang it through:
|
| 62 |
+
|
| 63 |
+
<audio controls src="https://huggingface.co/becausereasons/yue2-mltnt-militant-reggae/resolve/main/demos/mltnt_frontline__native_caption_previous_build.mp3"></audio>
|
| 64 |
+
|
| 65 |
+
Third, the original Frontline demo: the plain trap-dub prompt on the first Frontline build (160 BPM F minor, 88 % of vocal bars sung):
|
| 66 |
+
|
| 67 |
+
<audio controls src="https://huggingface.co/becausereasons/yue2-mltnt-militant-reggae/resolve/main/demos/mltnt_frontline__trap_dub_previous_build.mp3"></audio>
|
| 68 |
+
|
| 69 |
+
**MLTNT Fusion** β Fusion recipe, prompt `prompts/fusion_hiphop.txt` (planner chose a 150 BPM grid in Gβ― minor, three verse-chorus rounds and an outro)
|
| 70 |
+
|
| 71 |
+
<audio controls src="https://huggingface.co/becausereasons/yue2-mltnt-militant-reggae/resolve/main/demos/mltnt_fusion__demo.mp3"></audio>
|
| 72 |
+
|
| 73 |
+
**MLTNT Steppers** β baseline recipe, prompt `prompts/steppers_baseline.txt`
|
| 74 |
+
|
| 75 |
+
<audio controls src="https://huggingface.co/becausereasons/yue2-mltnt-militant-reggae/resolve/main/demos/mltnt_steppers__demo.mp3"></audio>
|
| 76 |
+
|
| 77 |
+
**MLTNT Roots** β baseline recipe, prompt `prompts/steppers_baseline.txt` (planner stayed on the 76 BPM roots grid in A minor, 4:59 long, 149 of 190 vocal bars sung)
|
| 78 |
+
|
| 79 |
+
<audio controls src="https://huggingface.co/becausereasons/yue2-mltnt-militant-reggae/resolve/main/demos/mltnt_roots__demo.mp3"></audio>
|
| 80 |
+
|
| 81 |
+
## Quick start (ComfyUI)
|
| 82 |
+
|
| 83 |
+
These LoRAs are in the fused-key layout that Comfy's YuE2 implementation uses (`text_encoders.*` for the planner, `diffusion_model.*` for the decoder). They were trained with, and load through, the **[FS_Audio Suite](https://github.com/KytraScript/ComfyUI-FS_Audio_Suite)** node pack.
|
| 84 |
+
|
| 85 |
+
1. ComfyUI β₯ v0.36.0 (native YuE2 support) and the FS_Audio Suite custom node pack.
|
| 86 |
+
2. Base model: `yue2_3b_bf16.safetensors` from [Comfy-Org/YuE2](https://huggingface.co/Comfy-Org/YuE2) in `models/checkpoints/`.
|
| 87 |
+
3. Drop one `mltnt_*.safetensors` into `models/loras/`.
|
| 88 |
+
4. Chain the nodes:
|
| 89 |
+
|
| 90 |
+
```
|
| 91 |
+
π§© FS_Audio Lora Loader ββlorasβββΆ π€ FS_Audio Model Loader ββpipeβββΆ π΅ FS_Audio Sampler βββΆ πΏ FS_Audio Output
|
| 92 |
+
lora_name = mltnt_steppers.safetensors yue2_checkpoint = yue2_3b_bf16 style = <prompt, starts with "mltnt,">
|
| 93 |
+
strength_clip = 1.0 (planner) melody_transcriber = none lyrics = <tagged lyric blocks>
|
| 94 |
+
strength_model = 1.0 (decoder) score_mode = full
|
| 95 |
+
```
|
| 96 |
+
|
| 97 |
+
Sampler settings used for every demo:
|
| 98 |
+
|
| 99 |
+
| widget | value |
|
| 100 |
+
|---|---|
|
| 101 |
+
| steps / sampler / scheduler | 32 / `dpm_2` / `sgm_uniform` |
|
| 102 |
+
| score_mode | `full` |
|
| 103 |
+
| song_length_cap | 360 (a full song; 150 truncates mid-way) |
|
| 104 |
+
| repetition_penalty | 1.2 |
|
| 105 |
+
|
| 106 |
+
Two proven recipes:
|
| 107 |
+
|
| 108 |
+
| | Baseline (Steppers / Roots / Frontline) | Fusion |
|
| 109 |
+
|---|---|---|
|
| 110 |
+
| strength_clip (planner) | 1.0 | 1.0 |
|
| 111 |
+
| strength_model (decoder) | 1.0 | **1.5** |
|
| 112 |
+
| Weirdness (cfg) | 1.0 | **1.4** |
|
| 113 |
+
| score_temperature | 0.7 | **0.9** |
|
| 114 |
+
| music_temperature | 1.0 | **1.2** |
|
| 115 |
+
|
| 116 |
+
Headless, the same graph as an API prompt:
|
| 117 |
+
|
| 118 |
+
```python
|
| 119 |
+
prompt = {
|
| 120 |
+
"1": {"class_type": "FSAudioLoraLoader", "inputs": {"lora_name": "mltnt_steppers.safetensors", "strength_model": 1.0, "strength_clip": 1.0}},
|
| 121 |
+
"2": {"class_type": "FSAudioModelLoader", "inputs": {"yue2_checkpoint": "yue2_3b_bf16.safetensors", "melody_transcriber": "none", "loras": ["1", 0]}},
|
| 122 |
+
"3": {"class_type": "FSAudioSampler", "inputs": {"pipe": ["2", 0], "style": STYLE, "lyrics": LYRICS, "seed": 7,
|
| 123 |
+
"song_length_cap": 360, "score_mode": "full", "steps": 32, "sampler": "dpm_2", "scheduler": "sgm_uniform",
|
| 124 |
+
"Weirdness (cfg)": 1.0, "score_temperature": 0.7, "music_temperature": 1.0, "repetition_penalty": 1.2}},
|
| 125 |
+
"4": {"class_type": "FSAudioOutput", "inputs": {"song": ["3", 0], "score": ["3", 1], "info": ["3", 2], "filename_prefix": "mltnt"}},
|
| 126 |
+
}
|
| 127 |
+
```
|
| 128 |
+
|
| 129 |
+
## Prompting
|
| 130 |
+
|
| 131 |
+
### Style prompt
|
| 132 |
+
|
| 133 |
+
Start with the trigger, then write **one descriptive sentence** in this order: language β genre β vocal β instruments β mood β BPM β production. This is the format the planner responds to. A bare tag list or `[Intro]β¦[Outro]` scaffolding in the style field produces "odd, not reggae" output.
|
| 134 |
+
|
| 135 |
+
```
|
| 136 |
+
mltnt, Jamaican Patois English, modern militant roots reggae, dark raspy male vocal with heavy patois delivery, deep bassline with massive sub bass drops, skanking guitar, hammond organ bubble, nyabinghi percussion, horn stabs and dub sirens, apocalyptic conscious mood, 76 BPM, sparse hard-hitting arrangement, dub impact hits, harder drum impact, heavier low end, spring reverb and tape delay, space before an explosive anthemic chorus
|
| 137 |
+
```
|
| 138 |
+
|
| 139 |
+
Ready-made prompts in `prompts/`:
|
| 140 |
+
|
| 141 |
+
| file | what it gets you |
|
| 142 |
+
|---|---|
|
| 143 |
+
| `steppers_baseline.txt` | the approved baseline: sub bass drops, dub impact hits, steppers groove, explosive chorus |
|
| 144 |
+
| `steppers_native_format.txt` | the same brief rewritten in the strict caption order |
|
| 145 |
+
| `fusion_hiphop.txt` | reggae hip-hop fusion, deejay toasting, boom-bap over one-drop (best with **MLTNT Fusion**) |
|
| 146 |
+
| `trap_dub.txt` | trap and dubstep low end, gospel-stack chorus, tape-stop drops, 80 BPM written but the planner lands at 150β160. Behind three of the Frontline clips; results swing from great to chaotic between seeds, so treat it as the wild card |
|
| 147 |
+
| `trap_dubstep_fusion.txt` | militant roots fused with trap and dubstep low end, halftime breakdowns, gang-vocal chorus |
|
| 148 |
+
|
| 149 |
+
Production descriptors ("boom-bap", "deejay toasting", "steppers") move the planner's tempo grid and key far more than the numeric BPM in the same prompt.
|
| 150 |
+
|
| 151 |
+
### Lyrics
|
| 152 |
+
|
| 153 |
+
Tagged blocks, 4 lines per verse works best. Use `[Verse]`, `[Pre-Chorus]`, `[Chorus]`, `[Bridge]`, `[Outro]`. **Do not add an empty `[Intro]`** β the planner will happily write an 18-bar instrumental intro on its own. Shape:
|
| 154 |
+
|
| 155 |
+
```
|
| 156 |
+
[Verse]
|
| 157 |
+
four lines, roughly eight words each
|
| 158 |
+
|
| 159 |
+
[Verse]
|
| 160 |
+
four more lines
|
| 161 |
+
|
| 162 |
+
[Pre-Chorus]
|
| 163 |
+
two short lines that build tension
|
| 164 |
+
|
| 165 |
+
[Chorus]
|
| 166 |
+
four lines, the hook repeated at least twice across the song
|
| 167 |
+
|
| 168 |
+
[Bridge]
|
| 169 |
+
a chant or a two-word call, then four lines
|
| 170 |
+
|
| 171 |
+
[Outro]
|
| 172 |
+
two to four closing lines
|
| 173 |
+
```
|
| 174 |
+
|
| 175 |
+
**Rapid-fire verses.** The deejay's delivery speed comes from the *lyric*, not from the prompt. Words like "rapid-fire", "double-time" or "fast flow" in the style prompt do not speed the vocal up (tested: they mostly push the planner toward a shorter, trap-flavoured plan). What works is line density: write the verse at **15β17 words per line**, eight lines to the verse, and keep the chorus at the normal 7β8 words so the contrast lands. On the same prompt and seed that roughly doubles the words per minute and the LoRAs pack the lines intelligibly instead of overrunning them. Shape of a fast verse line:
|
| 176 |
+
|
| 177 |
+
```
|
| 178 |
+
Dem seh the future bright but mi seh where the people stand, where the promise and the plan
|
| 179 |
+
```
|
| 180 |
+
|
| 181 |
+
Keep the syllables simple (one- and two-syllable words); dense lines full of long words still garble. The two Frontline clips above are this exact comparison: same model, prompt, recipe and seed, standard lyric vs dense lyric.
|
| 182 |
+
|
| 183 |
+
A 2-word pre-chorus and a chant bridge give the planner clear pacing cues. Repeat the `[Chorus]` block verbatim wherever you want it sung again.
|
| 184 |
+
|
| 185 |
+
### Which knob does what
|
| 186 |
+
|
| 187 |
+
- **strength_model (decoder)** is the lever to push. 1.0 β 1.5 makes the sound darker, rougher and more present without touching the writing.
|
| 188 |
+
- **strength_clip (planner) above 1.0 collapses the vocal.** At 1.25 the planner wrote 1 sung bar against 424 rests. Keep it at 1.0.
|
| 189 |
+
- **Weirdness (cfg)** changes only the sound, not the arrangement. 1.0 is clean, 1.4 is the Fusion sweet spot, 1.7 gets noticeably wilder.
|
| 190 |
+
- **music_temperature / score_temperature** are what actually change the writing (tempo grid, key, section lengths). 1.2 / 0.9 gave the most adventurous still-coherent plans.
|
| 191 |
+
- **Seeds** can end a song early (seed 44 stopped at 4:30 with a 360 s cap). That is the seed, not the cap.
|
| 192 |
+
|
| 193 |
+
## Training
|
| 194 |
+
|
| 195 |
+
Trained on a single RTX 5090 with the FS_Audio Suite **Artist Trainer** on Comfy's own YuE2 weights. One run trains the planner LoRA and the decoder LoRA together; the file also carries full `vae2llm` / `llm2vae` projection diffs.
|
| 196 |
+
|
| 197 |
+
| | Steppers | Fusion | Roots | Frontline |
|
| 198 |
+
|---|---|---|---|---|
|
| 199 |
+
| checkpoint | step 300 (final) | **step 200** of 350 | step 300 (final) | step 350 (final) |
|
| 200 |
+
| planner / decoder steps | 300 / 800 | 350 / 1000 | 300 / 800 | 350 / 600 |
|
| 201 |
+
| final artist (planner) loss | 4.650 | 5.280 | 4.471 | 4.853 |
|
| 202 |
+
| decoder loss at checkpoint | 1.191 | 1.219 | 1.151 | 1.280 |
|
| 203 |
+
|
| 204 |
+
Shared hyper-parameters: rank 64 planner / rank 32 decoder, LR 3e-5 planner / 4e-5 decoder / 2e-5 I/O projections, artist fraction 0.5 vs regularizer pack, KL 0.1, batch 2 songs, 8192 max tokens, 750-frame windows, EMA 0.99, score-first 0, no transcription.
|
| 205 |
+
|
| 206 |
+
**Why step 200 for Fusion.** The trainer names its `_best` checkpoint by *planner* loss, which keeps falling. The decoder bottomed at steps 150β200 and then overfit hard (1.22 β 1.58 by step 350); the `_best` file sounds burnt. Pick checkpoints from the decoder curve. Frontline was trained with the decoder capped at 600 steps; its decoder loss bottomed at step 200 and drifted up by only 0.002 by step 350, so the final checkpoint is used. Frontline's set also adds eight short excerpts (20β45 s) of the fastest-delivery verses, each with only its own lyric lines β that did not make speed promptable from the caption, but it is where its voice comes from.
|
| 207 |
+
|
| 208 |
+
### Captioning your own dataset
|
| 209 |
+
|
| 210 |
+
The planner learns from the captions as much as from the audio, so the caption format decides whether your prompts work later. What worked here:
|
| 211 |
+
|
| 212 |
+
- **One descriptive sentence per song, trigger first**, in the same order you will prompt in: language β genre β vocal β instruments β mood β BPM β production. Example shape: `mltnt, Jamaican Patois English, modern militant roots reggae, dark raspy male vocal with heavy patois delivery, deep bassline, skanking guitar, bubbling hammond organ, nyabinghi percussion, horn stabs and dub sirens, urgent conscious mood, 76 BPM, sparse hard-hitting arrangement, spring reverb and dub delays`.
|
| 213 |
+
- **No tag lists, no section scaffolding** (`[Intro]β¦[Outro]`) in the caption. Tag-list captions produced a model that responds to tag-list prompts and writes odd, un-reggae plans.
|
| 214 |
+
- **Measure BPM, don't guess.** `librosa.beat.beat_track` on each file, rounded, written as `NN BPM` near the end of the sentence.
|
| 215 |
+
- A vision-language model can draft the captions from the audio, but **check the vocal gender by hand**: high male registers get labelled "female" often enough that every batch needs a grep-and-fix pass. A wrong gender in the caption shows up as a wrong voice at inference.
|
| 216 |
+
- **Lyrics as tagged blocks** in the same format you will prompt with: `[Verse]`, `[Pre-Chorus]`, `[Chorus]`, `[Bridge]`, `[Outro]`, patois spelling kept as sung, no empty intro tag. Transcribe at high confidence or leave the line out; a mis-heard lyric teaches the planner the wrong syllable count for the bar.
|
| 217 |
+
- **Audio prep**: FLAC, trimmed to β€ 320 s with a fade. MP3 rips with a damaged leading frame crash the dataset builder ("Header missing"), so transcode with `ffmpeg -err_detect ignore_err` first.
|
| 218 |
+
- **Trainer flags**: `score-first 0` and no automatic transcription. Letting the trainer transcribe melodies itself put the vocal melody in the instrument voice and rests in the vocal voice, and the decoder loss started climbing after step 200.
|
| 219 |
+
|
| 220 |
+
## Known limitations
|
| 221 |
+
|
| 222 |
+
- Vocals are patois-flavoured English only.
|
| 223 |
+
- Dense multi-syllable lines overrun their bars and come out garbled. Budget roughly words Γ· 2 seconds per line at ~120 wpm and shorten the line rather than fighting pronunciation.
|
| 224 |
+
- With the 150 s cap the planner still writes full-length intros and interludes, so the song truncates before the second chorus. Use 360.
|
| 225 |
+
- The planner occasionally writes a plan with almost no sung bars. Change the seed; do not raise planner strength.
|
| 226 |
+
- No instrumental-only mode is baked in; the LoRAs assume a lyric.
|
| 227 |
+
|
| 228 |
+
## Files
|
| 229 |
+
|
| 230 |
+
```
|
| 231 |
+
mltnt_steppers.safetensors 177 MB planner + decoder LoRA (bf16 weights, fp32 projection diffs)
|
| 232 |
+
mltnt_fusion.safetensors 177 MB
|
| 233 |
+
mltnt_roots.safetensors 177 MB
|
| 234 |
+
mltnt_frontline.safetensors 177 MB
|
| 235 |
+
demos/ mp3 renders (192 kbps from the FLAC masters), seed 7 throughout
|
| 236 |
+
prompts/ style prompts
|
| 237 |
+
```
|
| 238 |
+
|
| 239 |
+
## Support
|
| 240 |
+
|
| 241 |
+
These LoRAs are trained on my own GPU and released free. If they're useful to you and you'd like to chip in for compute, there's a Ko-fi: **[ko-fi.com/becausereasons](https://ko-fi.com/becausereasons)** <3
|
| 242 |
+
|
| 243 |
+
## License and credits
|
| 244 |
+
|
| 245 |
+
Weights are released under **CC BY-NC 4.0**, inherited from the YuE2-3B base model. Non-commercial use only; attribute "MLTNT LoRAs by becausereasons".
|
| 246 |
+
|
| 247 |
+
- [YuE2](https://huggingface.co/m-a-p/YuE2-3B) by the Multimodal Art Projection (m-a-p) team; ComfyUI repack by [Comfy-Org](https://huggingface.co/Comfy-Org/YuE2).
|
| 248 |
+
- [ComfyUI-FS_Audio_Suite](https://github.com/KytraScript/ComfyUI-FS_Audio_Suite) by KytraScript / The Fixed Seed Company: inference nodes and the artist trainer.
|
| 249 |
+
- Trained and documented by becausereasons, September 2026.
|