--- license: other license_name: circlestone-labs-non-commercial-license license_link: LICENSE.md base_model: circlestone-labs/Anima pipeline_tag: text-to-image tags: - text-to-image - diffusion - anime - anima - comfyui - qwen - t5-free - safetensors - experimental --- # Anima T5-Free Base **Anima T5-Free Base** is an experimental derivative of [`circlestone-labs/Anima`](https://huggingface.co/circlestone-labs/Anima) that replaces Anima's original text-conditioning path with a Qwen-based conditioning system. The project has gone through several conditioning designs. The current **v0.3** branch removes the remaining T5 row-mapping bottleneck and exposes the full native Qwen token sequence directly to F128. > **Important:** the current **v0.3 native-rows path uses neither the T5 model > nor the T5 tokenizer**. Qwen3.5-2B-Base is the semantic and tokenization > source end-to-end. > > Earlier v0.2 F128 checkpoints still used a T5 tokenizer only as an > offset/row-map. Those files are kept as legacy research checkpoints and are > now marked `-old`. **Base model:** [CircleStone Labs / Anima](https://huggingface.co/circlestone-labs/Anima) **ComfyUI node:** [ComfyUI-AnimaT5Free](https://github.com/Disya123/ComfyUI-AnimaT5Free) **Support:** [Boosty](https://boosty.to/neotavern/donate) --- # Current checkpoints ## `anima-t5free-fused-v0.3-native-rows.safetensors` Current **working v0.3 native-rows checkpoint**. This version removes the T5-derived row map: ```text Qwen3.5-2B-Base ↓ all 25 hidden-state levels ↓ one carrier row per native Qwen token ↓ X [Lq, 25 × 2048] ↓ X [Lq, 51200] ↓ F128 receiver ↓ 28 block-specific K/V sets ``` The current `native-rows` checkpoint keeps the legacy fixed 512-slot cross-attention grid for compatibility with the already-trained receiver/DiT weights: ```text live K/V: [16, Lq, 128] ↓ zero-pad to 512 rows ↓ cross-attention over 512 rows ``` This is important: the zero rows are not a neutral implementation detail for the old weights. Because the historical attention path had no padding mask, those zero-key rows contributed to the softmax denominator. Removing them changes the attention function substantially. This checkpoint is therefore the **working transition baseline** for native Qwen rows. The same trained weights can also be used with the **VSINK compatibility runtime**. VSINK removes the physical zero-padded K/V rows and keeps only the `Lq` live native-Qwen rows, but analytically restores the denominator mass that the historical `(512-Lq)` zero-key rows contributed to attention. In other words, VSINK is **no physical grid**, but it intentionally preserves the legacy 512-slot softmax normalization. It is an inference/runtime change; the model weights themselves do not need to be retrained for this compatibility mode. --- ## `anima-t5free-fused-v0.3-native-rows-no-t5grid.safetensors` Experimental **no-physical-grid** checkpoint / training initialization. It uses the same native Qwen carrier: ```text H [25, Lq, 2048] ↓ X [Lq, 51200] ↓ F128 receiver ↓ 28 × K/V [16, Lq, 128] ``` No physical zero-padding to 512 rows is required. However, there are now **two different attention-normalization modes** that must not be conflated. ### VSINK compatibility mode ```text live K/V only: [16, Lq, 128] ↓ attention over live rows ↓ analytic legacy null-mass for (512-Lq) missing zero rows ``` This mode is inference-compatible with the existing trained receiver/DiT weights. It reproduces the useful effect of the old 512-slot grid without materializing those zero K/V rows. ### True no-grid mode ```text live K/V only: [16, Lq, 128] ↓ softmax only over the live rows ↓ no legacy null-mass ``` This changes the attention function itself. The existing receiver/DiT weights were trained in the legacy 512-slot normalization regime, so **true no-grid is not a drop-in inference conversion**. Direct true-no-grid inference with the current weights produces severe static/noise-like failures. Therefore, if true no-grid normalization is the goal, **continuation retraining is required**. Current experiments also suggest that a tiny CA-side adapter is not enough for a reliable recovery; the practical path is a substantial continuation/full fine-tune of the affected conditioning path and DiT. The same `.safetensors` weight package can therefore be used in two very different runtime regimes: - **VSINK:** no physical grid, legacy normalization preserved, no retraining required for compatibility; - **true no-grid:** no physical grid and no legacy null-mass, retraining required. # Legacy checkpoints ## `anima-t5free-fused-v0.2-f128-lite-ft-old.safetensors` Later v0.2 **F128 full-finetune** checkpoint. This is the checkpoint on which the main F128 DiT/receiver fine-tuning was performed. It uses the full 25-level Qwen trajectory, but still maps Qwen tokens into T5-derived rows before F128. Conceptually: ```text Qwen hidden trajectory ↓ T5-tokenizer-derived row map ↓ [n, 51200] ↓ F128 receiver ↓ 28 block-specific K/V ↓ fixed 512-slot attention grid ``` This checkpoint is kept for comparison and as the trained weight source for the v0.3 migration. ## `anima-t5free-fused-v0.2-experimental-old.safetensors` Original **F128** checkpoint before the later full-finetuning stage. It already uses the F128 design and the full 25-level Qwen trajectory, but uses the original FP32 receiver and predates the later full DiT fine-tune. ## `anima-t5free-fused-v0.1-old.safetensors` Legacy pre-F128 compatibility architecture. v0.1 used selected Qwen layers, a learned row planner and a fixed `512 × 1024` compatibility carrier. It is architecturally different from v0.2/v0.3 F128. --- # F128 v0.3 architecture The current native-rows carrier is deliberately simple: ```text prompt ↓ Qwen3.5-2B-Base ↓ H [25, Lq, 2048] ↓ per-layer RMS normalization ↓ transpose / flatten depth ↓ X [Lq, 51200] ↓ F128 receiver ↓ 28 × (K, V) ``` For each Qwen token, all 25 hidden-state levels are preserved: ```text 25 × 2048 = 51200 features per native Qwen token ``` There is no T5 tokenizer in this path and no sequence-axis pooling before the receiver. This means v0.3 preserves both: - **depth:** all 25 Qwen hidden-state levels; - **sequence:** all native Qwen token rows up to the current runtime limit. The receiver remains row-wise. Its learned projections operate on each `51200`-wide row independently and do not require a fixed sequence length. --- ## Why native rows were introduced The v0.2 T5-derived row map could collapse many Qwen tokens into only a few conditioning rows. A concrete diagnostic example: ```text Japanese prompt: 14 Qwen tokens → 3 T5-derived rows ``` The same semantic source therefore lost substantial sequence resolution before the trainable F128 receiver saw it. v0.3 removes that bottleneck: ```text 14 Qwen tokens → 14 F128 carrier rows ``` Testing confirmed that native Qwen rows remain semantically active with the existing weights: controlled prompt changes such as eye color still change the generated image while preserving the rest of an img2img source. --- # Physical grid, VSINK, and true no-grid These are three separate runtime behaviors. ### 1. Physical 512-slot grid (legacy reference) ```text Lq live K/V rows + (512-Lq) zero K/V rows → attention over 512 rows ``` Historically, the zero rows were included without a padding mask. For a zero key, the attention logit is exactly zero, so every unused row contributes ```text exp(0) = 1 ``` to the softmax denominator. ### 2. VSINK: no physical grid, legacy normalization preserved VSINK keeps only the live rows: ```text K/V shape: [16, Lq, 128] ``` but analytically reproduces the denominator contribution of the missing zero rows. For live attention logits `s_i`, let: ```text Z = sum_i exp(s_i) M = 512 - Lq ``` Ordinary live-only attention produces: ```text O_live = sum_i exp(s_i) v_i / Z ``` VSINK returns: ```text O_vsink = O_live * Z / (Z + M) ``` which is algebraically the same function as the physical 512-slot zero-padded grid under the current Anima attention contract. This equivalence has been verified numerically: - pre-`W_O` relative error: `6.87e-07`; - post-`W_O` FP32 relative error: `1.82e-06`; - the remaining BF16 difference is ordinary quantization/dithering noise. Therefore: ```text VSINK = no physical grid + virtual legacy null-mass ``` The number `512` is now a **compatibility constant in the attention normalization**, not the number of physically materialized K/V rows. ### 3. True no-grid ```text Lq live K/V rows → softmax only over those Lq rows → no legacy null-mass ``` This is the mathematically clean live-only attention path, but it is a different function from the one the current receiver/DiT weights were trained on. The trained model does **not** currently tolerate this conversion as a drop-in runtime change. True no-grid therefore requires continuation retraining. ## Why VSINK exists VSINK is a compatibility mechanism, not a claim that the historical normalization disappeared conceptually. It removes the physical cost of the 512-row zero-padded K/V grid while keeping the old attention behavior that the trained model expects. This distinction is important: ```text physical grid removed: yes legacy 512 null-mass removed: no ``` If the project later wants to eliminate the legacy null-mass itself, that is a separate training objective and should be treated as a genuine model migration, not as an inference-only conversion. ## Resolution behavior observed so far Same-seed PHYS vs VSINK tests indicate that high-resolution behavior is approximately preserved through at least `1536×1536` in the current test set. Measured end-to-end timings: | resolution | PHYS | VSINK | observed speedup | | --- | ---: | ---: | ---: | | 512² | 0.32-0.33 s | 0.31-0.32 s | 2-6% | | 1024² | 1.42 s | 1.35-1.40 s | 1.5-5% | | 1536² | 3.91 s | 3.75-3.86 s | 1.2-4% | The wall-clock gain is modest because text cross-attention is only a small part of total DiT cost, especially at high image resolutions. VSINK still removes the physical `(512-Lq)` K/V padding work from that text-attention path. # Current sequence-length limit The F128 receiver itself no longer requires a 512-row resampler or a fixed sequence-axis grid. However, the current runtime still has a loud **1..512 native Qwen row guard**. In `native-rows-v2`, one carrier row corresponds to one Qwen token position, so the current limit is effectively: ```text Lq <= 512 Qwen token positions ``` For VSINK specifically, the compatibility mass is defined as: ```text M = 512 - Lq ``` Therefore: - `Lq < 512`: VSINK adds the analytic legacy null-mass; - `Lq = 512`: `M = 0`, so VSINK becomes ordinary live-only attention; - `Lq > 512`: the current VSINK compatibility definition is not valid and must not be silently extrapolated. This `512` is **not the image resolution**. A `1024×1024` or `1536×1536` image still uses the same text-side compatibility constant; only the number of image queries changes. F128 can in principle support longer text sequences, but raising the runtime limit above 512 now requires an explicit design decision: either define a new compatibility normalization or retrain toward a true no-grid regime. It should not be implemented by silently changing the guard. # Latent coordinates The v0.3 native-rows branch uses the **stock Anima / Wan21 latent coordinate convention**. Do not apply the earlier raw-latent compatibility patch to v0.3 unless a checkpoint explicitly declares a different latent contract. --- # What v0.3 does not use The current native-rows conditioning path does **not** use: - the original T5 text encoder/model weights; - the T5 tokenizer; - T5 hidden states; - T5-derived row geometry; - Anima's original `llm_adapter` semantic path; - the old v0.1 learned row planner / OOV segmenter; - the old fixed `512 × 1024` compatibility carrier. The working `v0.3-native-rows` checkpoint still retains the historical **512 K/V attention grid** only as an attention-compatibility mechanism for existing trained weights. --- # Training status The strongest trained F128 weights currently come from the earlier v0.2 full-parameter DiT fine-tune. The trainable system consisted of: - the F128 receiver; - the Anima DiT. Qwen3.5-2B-Base remained frozen. The v0.3 files are an architectural migration of those learned weights: ```text v0.2 trained F128 ↓ native Qwen rows ↓ v0.3-native-rows ↓ remove physical 512-row K/V padding ├─ VSINK: keep legacy null-mass analytically → works with existing weights └─ true no-grid: remove null-mass entirely → requires retraining ``` The important result is that **removing the physical grid does not itself require retraining** if VSINK preserves the legacy attention normalization. By contrast, **true no-grid normalization does require continuation retraining**. Directly removing the legacy null-mass from the already-trained receiver/DiT produces severe static/noise-like failures. Current CA-only recovery experiments did not restore a reliable useful model, so true no-grid should be treated as a substantial continuation/full-finetune project rather than as a cheap inference patch. No retraining is required merely to use the VSINK compatibility runtime. # Prompt behavior observed so far The trained F128 path can respond to concrete prompt attributes including: - eye color; - hair color; - hair length; - character appearance; - broad scene/environment cues; - some natural-language descriptions. With v0.3 native rows, controlled img2img tests still show direct semantic control. For example, changing only `red eyes` to `green eyes` changes the eye color while preserving the same source composition. This is evidence that removing the T5 row bottleneck did not destroy the Qwen semantic signal. Larger global changes, binding and long multi-requirement prompts remain less reliable with the current trained weights. --- # Multilingual notes Qwen3.5 itself is multilingual. The old v0.2 row mapper introduced a particularly severe failure mode for CJK text because T5-derived row geometry could collapse many Qwen tokens into very few conditioning rows. v0.3 removes this specific architectural bottleneck by using native Qwen rows. However, the receiver/DiT training data was not balanced for multilingual instruction following, so **robust multilingual behavior is still not claimed yet**. In other words: ```text CJK row-collapse bug: removed in v0.3 multilingual training coverage: still limited ``` --- # Image editing The current checkpoints can be used in ordinary img2img workflows, but they are **not dedicated instruction-edit models**. A source image supplied through the diffusion latent path can preserve structure while text conditioning changes attributes. This is not yet multimodal vision-language editing: Qwen currently receives text, not the source image as visual context. Dedicated source-image + instruction training remains future work. --- # ComfyUI This repository requires: **[Disya123/ComfyUI-AnimaT5Free](https://github.com/Disya123/ComfyUI-AnimaT5Free)** Stock Anima loaders do not understand the F128 conditioning path. Typical current layout: ```text ComfyUI/ └── models/ ├── diffusion_models/ │ ├── anima-t5free-fused-v0.3-native-rows.safetensors │ └── anima-t5free-fused-v0.3-native-rows-no-t5grid.safetensors ├── text_encoders/ │ └── qwen_35_2b_base.safetensors └── vae/ └── qwen_image_vae.safetensors ``` Do **not** additionally load the original Anima `llm_adapter`. For normal inference today, use the trained native-row weights: ```text anima-t5free-fused-v0.3-native-rows.safetensors ``` Two compatibility runtimes are valid for those trained weights: - the legacy **physical 512-slot grid** reference path; - **VSINK**, if supported by the installed custom-node revision, which keeps only live `Lq` K/V rows while preserving the same legacy null-mass analytically. The `no-t5grid` weight package should not be interpreted as proof that true live-only no-grid normalization is inference-ready. **True no-grid still requires retraining.** With VSINK, however, the same no-physical-grid K/V shape can be used without retraining because the legacy normalization is retained. --- # Suggested starting settings For the currently working native-rows checkpoint: ```text sampler: Euler scheduler: simple CFG: ~3 steps: 20-32 ``` These are practical starting points, not universal optimal settings. For img2img, lower `denoise` preserves more source structure while values closer to `1.0` allow stronger rewriting. --- # Reproducibility For meaningful comparisons, keep fixed: ```text checkpoint revision Qwen text encoder revision custom-node revision carrier mode attention mode: physical-grid / VSINK / true-no-grid VSINK compatibility size (`N=512` for the current model) VAE latent coordinate convention sampler scheduler step count CFG seed / initial noise resolution / aspect ratio precision img2img denoise (if used) ``` Changing any of these can materially change the result. --- # Compatibility Compatibility with upstream Anima LoRAs, ControlNets, merges, training scripts and other extensions is **not guaranteed**. Anything that assumes the original Anima text adapter, original conditioning tensor format, original text-encoder path, or fixed text-conditioning behavior requires explicit adaptation. F128 checkpoints use custom state-dict keys and runtime hooks, so generic upstream Anima tooling should not be assumed to work unchanged. --- # Research status F128 is an ongoing research path. The current v0.3 work establishes that: - frozen Qwen3.5-2B-Base can provide the text representation source; - all 25 Qwen hidden-state levels can be retained; - native Qwen token rows can be passed directly to F128 without T5 row conversion; - the receiver itself does not structurally require a fixed sequence length; - block-specific K/V can drive all 28 native Anima cross-attention sites; - native-row conditioning remains semantically active with the existing trained weights; - the historical fixed 512 **physical K/V grid** is not required for native Qwen rows; - the historical 512-slot **softmax null-mass**, however, is part of the function learned by the current receiver/DiT; - VSINK can reproduce that legacy null-mass analytically while keeping only the live `Lq` K/V rows; - PHYS and VSINK match to FP32-rounding accuracy in direct attention tests; - high-resolution behavior is approximately preserved through at least `1536×1536` in the current same-seed tests; - true no-grid normalization remains a separate model-migration target and requires retraining. The current practical path is therefore: ```text native Qwen rows + live-only physical K/V + VSINK legacy normalization ``` A major continuation/full-finetune step is required only if the project decides to remove the legacy null-mass itself and move to **true no-grid** attention. # Known limitations The current trained model may still fail on: - exact object counts; - negation; - spatial relations; - ownership / attribute binding; - multi-character binding; - small accessories; - exact clothing details; - complex multi-object composition; - long chains of simultaneous requirements; - text rendering. The project remains experimental. --- # Support If you find this project useful and want to support further training and experiments: **[Support me on Boosty](https://boosty.to/neotavern/donate)** Support is optional. There are no exclusive model files, early-access checkpoints, or gated model content attached to the subscription. --- # Attribution This model is derived from: **CircleStone Labs - Anima** https://huggingface.co/circlestone-labs/Anima This repository contains a modified text-conditioning architecture and fine-tuned diffusion checkpoints. It does not claim authorship of the original Anima model. --- # License The CircleStone model components are licensed under the **CircleStone Non-Commercial License**. See [`LICENSE.md`](./LICENSE.md) for the license text included with this repository. Use of this derivative model remains subject to the applicable upstream license terms. Third-party components, including Qwen and runtime dependencies, remain subject to their own applicable licenses. This model card is descriptive and is not legal advice.