CNZN — Canzone Italiana LoRAs for YuE2

Artist-style LoRAs that push YuE2-3B into classic Italian pop: the Sanremo-era ballad with a warm, raspy male baritone, piano, strings and a sweeping singalong chorus, and on request Italo disco, new wave and funk-pop. Vocals are sung in Italian. Trigger word: cnzn.

September 2026 update: second generation. Four new checkpoints trained with a different trainer (the experimental YuE2 support in Ostris AI Toolkit). They capture the voices and the vocal tonality of the idiom far more closely than the first release: a darker, raspier baritone, choir backings, whispered and spoken passages, and songs that end on their own. The first release, cnzn_sanremo, stays available below; it is looser in voice but very forgiving.

Last updated: 20 September 2026 (second-generation files, demos and the trainer comparison below).

Every file patches both halves of YuE2: the autoregressive planner (writes the score, the arrangement and the vocal lines, and carries most of the singer identity) and the flow-matching decoder (the sound).

LoRA File Character Start here
🌙 CNZN Notte cnzn_notte.safetensors The identity checkpoint (step 600). Dark, raspy, close-miked baritone; opens and closes songs with whispered lines when the prompt asks for spoken delivery. clip 0.5 / model 1.0
🎭 CNZN Teatro cnzn_teatro.safetensors The ballad checkpoint (step 650, final). Most vocal drama and dynamics; holds a full ballad at full strength and lands the ending. clip 1.0 / model 1.0 for ballads, clip 0.5 for up-tempo
🤫 CNZN Sussurro cnzn_sussurro.safetensors The whisper checkpoint (step 350). Early and light: at full planner strength it follows a whispered / spoken-word prompt most faithfully, close-miked low baritone whispers on the verses. clip 1.0 / model 1.0 with prompts/whisper_tagged.txt
🎶 CNZN Coro cnzn_coro.safetensors Step 550. Sweeter, with the backing choir up front; the best of the three for Italo disco. clip 1.0 for ballads, clip 0.5 for disco
🇮🇹 CNZN Sanremo cnzn_sanremo.safetensors First release (FS_Audio trainer). Warm, more generic baritone, very stable, writes long full-structure songs. clip 1.0 / model 1.0

clip is strength_clip (the planner), model is strength_model (the decoder).

Listen: second generation

Same original Italian lyric, seed 7, 32 steps dpm_2 / sgm_uniform, 360 s cap, no post-processing. Every song below ended on its own.

Notte, modern baritone — prompts/modern_baritone_spoken.txt, clip 0.5 / model 1.0. Whispered opening and close, dark raspy verses, sung chorus (90 BPM, D minor, 3:44):

Teatro, classic ballad — prompts/classic_ballad.txt, full strength (96 BPM, D major, 4:19):

Coro, classic ballad — the same prompt and seed one checkpoint earlier: the choir backings move forward (96 BPM, D major, 4:51):

Coro, Italo disco — prompts/italo_disco_falsetto.txt, clip 0.5 / model 1.0. Short spoken-sung verses, falsetto chorus (115 BPM, D major, 4:52):

Sussurro, whisper-tagged prompt — prompts/whisper_tagged.txt with the verses tagged [Spoken], full strength. Whispered verses throughout, sung chorus (93 BPM, D minor, 3:34):

Teatro, whisper-tagged prompt — prompts/whisper_tagged.txt, clip 0.5 / model 1.0 (93 BPM, D minor, 3:28):

Range: the same prompts as the first release

The second-generation files are not ballad-only. These use the first release's own range prompts, same lyric and seed, so they compare one to one with the range demos further down. Up-tempo styles run at clip 0.5 / model 1.0.

Notte, electro-funk — prompts/electro_funk_registers.txt, clip 0.5 (125 BPM, D minor, 4:10):

Coro, new wave — prompts/new_wave_baritone.txt, clip 0.5. Low baritone verses, tenor chorus, and unlike the first release it ends on its own (115 BPM, D minor, 5:26):

Teatro, Italo pop — prompts/italo_pop.txt, clip 0.5 (115 BPM, D major, 4:22):

Notte, Sanremo ballad — prompts/sanremo_ballad.txt, the first release's headline prompt, full strength (73 BPM, D major, 5:10):

Listen: first release, the range demos

All demos use the same original Italian lyric (ten tagged sections), seed 7, baseline recipe, 32 steps dpm_2 / sgm_uniform, no post-processing.

The demo — prompt prompts/sanremo_ballad.txt, 360 s cap. The planner chose 72 BPM in D major and wrote the whole lyric through: intro, two verse / pre-chorus / chorus rounds, bridge, final chorus, outro (5:56, 84 of 108 vocal bars sung):

Classic ballad, prompts/classic_ballad.txt (96 BPM, D major, 360 s cap, 101 of 155 vocal bars sung):

The same prompt and seed with a 240 s cap. The planner wrote the identical score, but the decoder renders the whole song as one piece, so a different cap is a different performance of the same plan (cut in the second chorus):

Italo disco, prompts/italo_disco_falsetto.txt: falsetto chorus over baritone spoken-sung verses, arpeggiated synth bass, vocoder harmonies (115 BPM, D major, 4:57):

New wave, prompts/new_wave_baritone.txt: low baritone verses rising to a strained tenor chorus, cold pads, sequenced bass, LinnDrum (115 BPM, D minor, hits the 6:00 cap, 153 of 196 vocal bars sung):

Using the second-generation files

They load exactly like the first release (same fused-key layout, text_encoders.* planner and diffusion_model.* decoder, all 224 modules; the LoRA matrices are stored as lora_A / lora_B). Tested through the FS_Audio Lora Loader with the graph in Quick start below.

  • Ballads and slow songs: clip 1.0 / model 1.0.
  • Italo disco, new wave, anything up-tempo, or whenever a render runs away: clip 0.5 / model 1.0. The half-strength planner keeps the voice and the tonality and hands song structure back to the base model.
  • Whispers on demand: cnzn_sussurro at clip 1.0 with the whisper phrase in the style sentence and [Spoken] on the sections you want whispered. The whisper lives in the planner at full strength; clip 0.5 dilutes it.
  • Spoken or whispered delivery on the other files: describe it in the style sentence ("speaks the verses in a low intimate close-miked spoken word, then sings the chorus"). That is what triggers it; section tags in the lyric matter less.
  • Do not stack this LoRA's planner with another planner LoRA trained on similar material. On in-distribution prompts the two pushes add up and the render burns; clip 0.5 + 0.5 only behaved on prompts far from the training data.

First release: CNZN Sanremo (FS_Audio trainer)

An artist-style LoRA that pushes YuE2-3B into classic Italian pop: the Sanremo-era ballad with a warm, raspy male lead, lush analog synth pads, piano, fretless bass, gated-reverb drums and a sweeping singalong chorus. The same weights happily go the other way, into Italo disco, new wave and electro-funk, when the prompt asks for it. Vocals are sung in Italian.

The file patches both halves of YuE2 in one go: the autoregressive planner (writes the score, decides the arrangement and the vocal lines) and the flow-matching decoder (the sound). Trigger word: cnzn.

LoRA File Character Start here
🇮🇹 CNZN Sanremo cnzn_sanremo.safetensors Warm raspy male baritone, romantic and nostalgic, big melodic chorus. Writes full songs with intro, verses, pre-chorus, chorus, bridge and outro when the lyric has them. clip 1.0 / model 1.0, cfg 1.0

Quick start (ComfyUI)

The LoRA is in the fused-key layout that Comfy's YuE2 implementation uses (text_encoders.* for the planner, diffusion_model.* for the decoder). It was trained with, and loads through, the FS_Audio Suite node pack.

  1. ComfyUI ≥ v0.36.0 (native YuE2 support) and the FS_Audio Suite custom node pack.
  2. Base model: yue2_3b_bf16.safetensors from Comfy-Org/YuE2 in models/checkpoints/.
  3. Drop cnzn_sanremo.safetensors into models/loras/.
  4. Chain the nodes:
🧩 FS_Audio Lora Loader  ──loras──▶  🎤 FS_Audio Model Loader  ──pipe──▶  🎵 FS_Audio Sampler  ──▶  💿 FS_Audio Output
   lora_name      = cnzn_sanremo.safetensors        yue2_checkpoint = yue2_3b_bf16       style   = <prompt, starts with "cnzn,">
   strength_clip  = 1.0   (planner)                 melody_transcriber = none            lyrics  = <tagged Italian lyric blocks>
   strength_model = 1.0   (decoder)                                                      score_mode = full

Sampler settings used for every demo:

widget value
steps / sampler / scheduler 32 / dpm_2 / sgm_uniform
score_mode full
song_length_cap 360 (the planner writes 5 to 7 minute songs for a ten-section lyric; 240 cuts the second chorus)
repetition_penalty 1.2

Two recipes:

Baseline Wild
strength_clip (planner) 1.0 1.0
strength_model (decoder) 1.0 (1.5 for a rougher, more present voice) 1.0
Weirdness (cfg) 1.0 1.4
score_temperature 0.7 0.9
music_temperature 1.0 1.2

Headless, the same graph as an API prompt:

prompt = {
  "1": {"class_type": "FSAudioLoraLoader", "inputs": {"lora_name": "cnzn_sanremo.safetensors", "strength_model": 1.0, "strength_clip": 1.0}},
  "2": {"class_type": "FSAudioModelLoader", "inputs": {"yue2_checkpoint": "yue2_3b_bf16.safetensors", "melody_transcriber": "none", "loras": ["1", 0]}},
  "3": {"class_type": "FSAudioSampler", "inputs": {"pipe": ["2", 0], "style": STYLE, "lyrics": LYRICS, "seed": 7,
        "song_length_cap": 360, "score_mode": "full", "steps": 32, "sampler": "dpm_2", "scheduler": "sgm_uniform",
        "Weirdness (cfg)": 1.0, "score_temperature": 0.7, "music_temperature": 1.0, "repetition_penalty": 1.2}},
  "4": {"class_type": "FSAudioOutput", "inputs": {"song": ["3", 0], "score": ["3", 1], "info": ["3", 2], "filename_prefix": "cnzn"}},
}

Prompting

Style prompt

Start with the trigger, then write one descriptive sentence in this order: language → genre → vocal → instruments → mood → production → BPM. This is the format the planner was trained on and responds to. Bare tag lists produce odd plans.

cnzn, Italian, classic 80s Italian pop ballad in the Sanremo style, warm expressive male baritone lead vocal, lush analog synth pads, acoustic piano, fretless bass, clean electric guitar, gated reverb drums, sweeping melodic singalong chorus, romantic, nostalgic, resilient and cinematic, elegant open-vowel phrasing with strong repetition, emotionally direct, dry intimate vocal mix with minimal echo and reverb, 72 BPM

Ready-made prompts in prompts/:

file what it gets you
sanremo_ballad.txt the demo prompt: the Sanremo-era ballad, 72 BPM, dry intimate vocal
classic_ballad.txt a plainer 96 BPM ballad with choir harmonies and synth strings
italo_disco_falsetto.txt Italo disco at 118 BPM written (planner lands at 115), falsetto chorus, vocoder, handclaps
new_wave_baritone.txt 80s new wave, baritone verses to tenor chorus, cold pads, LinnDrum, D minor
electro_funk_registers.txt electro-funk Italo pop at 124 BPM, register slides, slap synth bass, FM piano (rendered, not in the demos)
italo_pop.txt upbeat Italo pop with gated snare and chant backing vocals (written in the same format, not rendered at release)

Register words work: "falsetto", "baritone", "head voice", "spoken-sung verses" all show up in the render. Synth vocabulary ("arpeggiated synth bass", "vocoder harmonies", "LinnDrum", "gated snare") moves the planner off the ballad grid and up to 115–125 BPM together with the numeric BPM.

Lyrics

Tagged Italian blocks. Use [Intro], [Verse 1], [Verse 2], [Pre-Chorus], [Chorus], [Bridge], [Outro]; number the verses and write a repeated chorus out again where you want it sung. Standard Italian orthography with accents. Shape that made the demo:

[Intro]
two short lines and the hook phrase

[Verse 1]
two four-line stanzas, roughly eight to ten syllables per line

[Pre-Chorus]
four lines that build

[Chorus]
three four-line stanzas, the hook line repeated at the top of each

[Verse 2]
two more four-line stanzas

[Pre-Chorus]
(repeat)

[Chorus]
(repeat, written out)

[Bridge]
two four-line stanzas

[Chorus]
(repeat, written out)

[Outro]
the hook phrase, twice, trailing off

The planner honours that structure section for section. Open vowels and short words sing cleanest; long dense lines overrun their bars.

Which knob does what

  • strength_model (decoder) is the lever to push. 1.0 → 1.5 makes the voice rougher and more present without touching the writing.
  • strength_clip (planner) above 1.0 collapses the vocal on YuE2 artist LoRAs in general. Keep it at 1.0.
  • Weirdness (cfg) changes only the sound. music_temperature / score_temperature change the writing (tempo grid, key, section lengths, vocal density); 1.2 / 0.9 on the Italo disco prompt added a pre-chorus and raised sung bars from 62 % to 79 % of the plan.
  • Song length is not promptable. A ten-section lyric plans to 5–7 minutes regardless of "3 minutes" in the prompt. Shorten the lyric or lower the cap.

Two trainers: what we learned

Both generations were trained on the same core set of songs, on one RTX 5090, and rendered with the same prompts, lyric and seed, so the differences below come from the trainer.

Old trainer versus new trainer. The first release was trained with the FS_Audio Suite Artist Trainer, the second generation with AI Toolkit. Comparing them on the same songs explained something we had noticed across all our earlier LoRAs: they learned a genre convincingly but never a voice.

  • What we found in the old pipeline. In the trainer version we used (September 2026), the decoder-training path conditions the decoder on raw codec indices, while at inference the decoder is conditioned on vocabulary ids (index plus an offset of 151853). The planner path adds the offset correctly. We checked it three ways: teacher-forcing real recordings' tokens through the base decoder tracks the original closely only with the offset (onset correlation 0.83 to 0.94, against about 0.05 without it, which is the shuffled baseline); the trainer's step-0 decoder loss matches the no-offset figure on three datasets; and a decoder LoRA trained that way did not improve teacher-forced reconstruction.
  • What that means in practice. The decoder half of those LoRAs still learns something real, a whole-dataset shift in sound (we measured 2 to 8 dB of spectral change, clearly audible), but not the token-to-sound detail that carries a particular singer. That matches what our ears had been telling us.
  • What we did not solve. Adding the offset in the trainer lowered validation loss but made held-out reconstruction steadily worse, so the training forward differs from the inference path in some further way we have not identified. We reverted to the stock trainer, and we are describing an observation, not shipping a fix. Check the current FS_Audio Suite before assuming this still applies.
  • What the old trainer does better. Its planner trains on whole songs (8192 tokens) with a regulariser pack and a held-out song, so it sees every ending. That is why cnzn_sanremo writes long, complete, well-structured songs at full strength with no tuning, and why it remains the forgiving choice.
  • What the new trainer changes. It applies the offset on both paths and trains planner and decoder together in one network, which is where the voices come from. Its weakness is the mirror image: memory limits the planner to the first 160 s of each song, so structure and endings have to be taught with short whole items and protected with planner strength 0.5.

What the AI Toolkit trainer does differently. One LoRA network spans both experts and is trained jointly on the Comfy-Org checkpoint (the int8 ConvRot repack here): next-token loss on the planner over each song from its start, flow-matching loss on the decoder over a random 60 s window. Audio is tokenised with the community real-audio tokenizer (Mothersuperior/yue2-mothersuperior-realaudio-tokenizer-v4, a MERT-based head), and every song gets a SheetSage2 lead sheet so the planner also learns the score-first path. A KL term (ar_kl_weight) anchors the planner to the base model.

What it buys. Singer identity. The first-generation file learned the genre; these learned the voices: register, rasp, phrasing, choir textures, spoken and whispered asides. In same-seed A/B listening the difference is not subtle.

What it costs. The planner over-commits easily. Observed over two runs and roughly 130 test renders:

  • Full planner strength is only safe up to a point. In the first toolkit run (20 songs, 1000 steps) full strength was good at steps 400 to 600, stopped ending songs from step 700, and was a mess by 1000. Half planner strength (clip 0.5) was good at every checkpoint.
  • The failure is visible in the score before you render. In score mode the planner writes the whole ABC score first, and song length is simply bars times tempo. An over-committed planner writes a runaway score: hundreds of bars stuck inside "intro > verse", which renders as minutes of groove with no singing. A healthy score walks through its sections. Parsing the .abc sidecar (bar count, tempo, section list) catches these without listening.
  • Hitting the length cap is two different things. Either a healthy score that is simply longer than the cap (a ten-section lyric at 96 BPM wants about six and a half minutes: raise the cap or drop a chorus), or a runaway. Only the second is a fault.
  • More data and whole short items fixed most of it. The second run added three songs and five short excerpts (15 to 40 s: spoken, whispered and low-register passages cut from songs already in the set, captioned as such, repeated twice). Because those excerpts fit in memory whole, the planner finally trained on how a piece ends; full songs are truncated to their first 160 s by the memory budget and never show it an ending. Result: 45 of 48 evaluation renders planned cleanly within the cap, full strength holds a ballad at the final checkpoint, and spoken passages stopped garbling.
  • Thin material stays thin. About 75 seconds of whispered and spoken source was enough to make spoken delivery stable and to get whispered openings, not enough to make it dependable on every prompt. Distinct passages beat repeats; the planner memorises exact sentences quickly.
  • Signal meters did not detect the failures. Spectral flatness, high-frequency share, clipping and loudness all tracked the arrangement, not the garbling: approved renders and broken ones scored alike. Compare renders that differ in one setting, by ear, and use the score check for structure.

Practical notes for anyone training YuE2 with AI Toolkit on Windows with 32 GB

  • Steps start at 3 to 5 s and degrade to minutes with VRAM "full": the caching allocator fragments on variable song lengths, and because the Windows driver spills to system RAM instead of raising out-of-memory, the cache never flushes. One torch.cuda.empty_cache() per training step fixed it outright (32 GB used down to about 18 GB, a steady 4 to 5 s per step). Allocator garbage-collection settings did not help.
  • In-training samples use CUDA graphs, and the next training step then fails with "Offset increment outside graph capture". Set ar_cuda_graphs: false in the model kwargs or disable sampling, and evaluate in ComfyUI instead.
  • Whole-song planner loss does not fit in 32 GB with lead sheets (songs here are 5 to 8 k codec tokens plus a 2 to 3 k token sheet). ar_max_tokens: 4000 does.
  • Save often in the 300 to 700 step range and audition a ladder of checkpoints: the keepers of both runs sat between steps 400 and 650, and checkpoints differ in flavour as much as in quality.
  • Stop and resume works from the last saved checkpoint; an on-demand save before stopping loses nothing.
second generation (Notte / Teatro / Coro)
trainer Ostris AI Toolkit, YuE2 audio extension (experimental), joint planner + decoder LoRA
base weights Comfy-Org yue2_3b_int8_convrot for training, yue2_3b_bf16 for the demos
dataset 23 songs (Italian pop from the late 1970s to the 2000s, including two AI-generated anchor songs) plus five short spoken / whispered / low-register excerpts, repeated twice
steps / checkpoints 650, saved every 50; published 350, 550, 600, 650
network rank 32 / alpha 32 on both experts
learning rate 5e-5 decoder, x0.6 on the planner, AdamW 8-bit
planner settings ar_kl_weight 0.2, cot: full, abc_dropout 0.5, ar_max_tokens 4000, 1500-frame decoder window
wall time 46 minutes for 650 steps

Training

Trained on a single RTX 5090 with the FS_Audio Suite Artist Trainer on Comfy's own YuE2 weights. One run trains the planner LoRA and the decoder LoRA together; the file also carries full vae2llm / llm2vae projection diffs.

CNZN Sanremo
dataset 20 songs, 87 minutes: Italian pop from the late 1970s to the 2000s, mostly Sanremo-style ballads and Italo pop with a few synth-pop and funk-pop cuts, plus two AI-generated anchor songs in the target style
checkpoint step 350 (final)
planner / decoder steps 350 / 600
artist (planner) loss 5.154 → 4.423
decoder loss 1.193 → 1.176, monotone, no upturn
wall time 29 minutes including the dataset build

Hyper-parameters: rank 64 planner / rank 32 decoder, LR 3e-5 planner / 4e-5 decoder / 2e-5 I/O projections, artist fraction 0.5 vs regularizer pack, KL 0.1, batch 2 songs, 8192 max tokens, 750-frame windows, EMA 0.99, one song held out, score-first 0, no transcription.

The anchor trick. The two AI-generated songs in the set were captioned with the exact prompt published above as sanremo_ballad.txt, not with an automatic description. Doing the same on an earlier reggae LoRA was the single largest quality jump in that series, so this run was restarted to apply it. The planner learns the prompt-to-plan mapping from those captions; the rest of the set teaches the voice and the idiom.

Captioning your own dataset

The planner learns from the captions as much as from the audio, so the caption format decides whether your prompts work later. What worked here:

  • One descriptive sentence per song, trigger first, in the order you will prompt in: language → genre → vocal → instruments → mood → production → BPM. Decade words ("80s Italo disco") are useful descriptors for this material and are worth keeping.
  • No tag lists, no section scaffolding in the caption.
  • Measure BPM, don't guess, and sanity-check it: beat trackers double slow ballads (one 66 BPM ballad came out as 136). If the caption says "slow" and the number says 130+, halve it.
  • A vision-language model can draft the captions from the audio, but check the vocal gender by hand: a high male tenor in a duet got labelled "female" here. A wrong gender in the caption shows up as a wrong voice at inference.
  • Check the language by ear or with a speech model, not from the file name. Two files in this set carried Spanish titles and turned out to be the Italian versions.
  • Lyrics as tagged blocks with standard orthography and accents. One song has a multilingual chorus; foreign passages were written in their own language (Russian in Latin transliteration). Transcribe at high confidence; a mis-heard lyric teaches the planner the wrong syllable count for the bar.
  • Audio prep: FLAC 44.1 kHz / 16-bit, trimmed to ≤ 320 s with a fade. MP3 rips with a damaged leading frame crash the dataset builder ("Header missing"); transcode with ffmpeg -err_detect ignore_err. One file was an MP3 wrapped in a RIFF header and needed -f mp3 forced.
  • Trainer flags: score-first 0 and no automatic transcription.

Known limitations

  • Italian vocals only. The training data has one multilingual chorus; do not expect other languages to hold.
  • Songs run long: 5 to 7 minutes for a full ten-section lyric. Use the 360 s cap and a shorter lyric for radio length.
  • A few prompts leave the bridge or outro instrumental. Change the seed or use the wild recipe, which raised vocal density on the same prompt and seed.
  • Female voices are in the training data only as duet partners and backing harmonies; a female lead is not a tested direction.
  • No instrumental-only mode is baked in; the LoRA assumes a lyric.

Files

cnzn_notte.safetensors          118 MB   second generation, step 600 (planner + decoder LoRA, bf16)
cnzn_teatro.safetensors         118 MB   second generation, step 650
cnzn_coro.safetensors           118 MB   second generation, step 550
cnzn_sanremo.safetensors        177 MB   first release: planner + decoder LoRA (bf16 weights, fp32 projection diffs)
demos/                          mp3 renders (192 kbps from the FLAC masters), seed 7 throughout
prompts/                        style prompts

Support

These LoRAs are trained on my own GPU and released free. If they're useful to you and you'd like to chip in for compute, there's a Ko-fi: ko-fi.com/becausereasons <3

License and credits

Weights are released under CC BY-NC 4.0, inherited from the YuE2-3B base model. Non-commercial use only; attribute "CNZN LoRA by becausereasons".

  • YuE2 by the Multimodal Art Projection (m-a-p) team; ComfyUI repack by Comfy-Org.
  • AI Toolkit by Ostris: the trainer behind the second-generation files; real-audio tokenizer by Kytra.
  • ComfyUI-FS_Audio_Suite by KytraScript / The Fixed Seed Company: inference nodes and the artist trainer.
  • Sister release: MLTNT militant roots reggae LoRAs.
  • Trained and documented by becausereasons, September 2026.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for becausereasons/yue2-cnzn-canzone-italiana

Finetuned
Comfy-Org/YuE2
Adapter
(9)
this model