BulmaX Gen2 - Weights Only (No Source)
Public weights-only repository. No architecture or training source is stored in this repository.
BulmaX Gen2 is a custom generative multimodal research model with a language backbone, audio and vision pathways, a persistent learned coherence graph, and adaptive model-growth systems. This repository is a public weights and continuation-state mirror. It does not contain the private architecture or training source and is not a drop-in Transformers model.
Graph-backed context: 5,242,880 tokens
BulmaX's graph-context evaluation reached 5,242,880 tokens (the 5M stage), with all registered stages passing across three seeds. The design combines short-range neural attention with long-range coherence memory: the local attention window stays at 1,024 tokens, while context is streamed through 5,120 chunks with isolated per-document graph state.
This is graph-backed context, not a 5M-token dense-attention or KV-cache window. The graph carries information across chunks; document boundaries reset transient state so separate documents do not share their context. Persistent learned graph weights and transient document state are distinct.
Verified context results
| Measurement | Result |
|---|---|
| Largest tested stream | 5,242,880 tokens |
| Local attention / chunk size | 1,024 tokens |
| Chunks at the largest stage | 5,120 |
| Stages | 8K, 16K, 32K, 64K, 128K, 256K, 512K, 1M, 2M, 5M |
| Seeds per stage | 3: seeds 0, 1, and 2 |
| Completed stage-seed measurements | 30 / 30 |
| 5M-stage exact recall | 1.0 for each seed |
| 5M-stage retrieval precision / source-span recall | 1.0 / 1.0 for each seed |
| 5M-stage temporal-order accuracy / interference error | 1.0 / 0.0 for each seed |
| Document reset and carry | Verified for each 5M-stage seed |
| Registered campaign verdict | PROMOTED through 5M |
These are targeted graph-memory recall measurements, not a general-purpose 5M-token reasoning benchmark. Each 5M-stage seed tests four queries against eight source facts, with retrieval top-k 8. Model-weight hashes were unchanged during the evaluation, and each stage restarted from the frozen base. The separate production-recall evaluation also passed its graph-versus-baseline and graph-versus-shuffled controls. Broader long-document generation, reasoning, and adversarial recall quality require additional evaluation. The recorded memory telemetry does not establish end-to-end GPU memory use for 5M-token generation.
Native and Rust performance
The context path uses the native coherence engine with Rust host-side work and state-only reads that avoid copying the entire persistent graph for each read. In a paired, warm eight-token response measurement on the same checkpoint, response time fell from 48.27 seconds to 8.22 seconds, and first-token time fell from 6.49 seconds to 1.63 seconds, with identical generated token IDs. This measured a short 18-token prompt, not a complete 5M-token request; cold loading and full-context ingestion are excluded. The newly initialized graph injection gates were zero in that comparison, so it establishes output parity and latency for that configuration, not the quality of learned graph injection.
Evidence identity: context campaign fb38694d6220c44eda144728e411952288dfa3750ce46346a87045885d2505a4;
measurement index d56b8ec7f1e8f15556410bdcbf026b010463667739529458556bdcd388e73268.
The tested model hash is 4364e1f5db4630d37af91d740b5cf86d057d4b48767a54577102925ce540f7fe
and graph hash is be0f05bbd24b9098a25e46742bde82011f83c8b73758ecaad4f186a162963b45.
These identify the evaluated pair. The step-155505 inventory below is a
different artifact and is not a claim that the evaluated context capability
carries over to it.
Step-155505 lineage and shape
The recovery lineage preserves every tensor learned through training step 149,200 and reconstructs the model's intended runtime shape around that checkpoint with audited zero-contribution grafts. Those grafts are present, wired, and trainable, and began as exact identities so they could not overwrite the checkpoint's learned behavior at initialization. Step 155,505 is a later checkpoint in the same lineage: the grafted systems have since received training, and the values below describe its fully expanded stored shape.
| Property | Verified value |
|---|---|
| Production class | BulmaXGenerativeMultimodal |
| Resume checkpoint | step 155,505 |
| Training-source revision | ee62e0813dd2be4e19d243fa80f07d227da28c0f |
| Trainable parameter elements | 9,565,932,416 |
| Parameter tensors | 16,795 |
| Buffer elements / tensors | 2,638,192,257 / 3,358 |
| Total reconstructed state elements / entries | 12,204,124,673 / 20,153 |
| Stored step-155505 state elements / tensors | 12,204,124,673 / 20,153 |
| Stored checkpoint dtypes | 19,457 BF16 tensors, 412 I64 tensors, 280 BOOL tensors and 4 U8 tensors |
The 9.566B figure is the exact trainable-parameter count of the reconstructed grafted runtime. Total state also includes persistent buffers; it is not a second parameter-count estimate.
The published step-155505 bundle is the first file that physically stores the fully expanded grafted shape after resumed training.
Backbone
| Component | Shape |
|---|---|
| Vocabulary | 32,004 tokens: 32,000 SentencePiece tokens plus four modality delimiters at IDs 32000-32003 |
| Hidden width | 4,096 |
| Transformer depth | 32 layers |
| Local attention window | 1,024 tokens |
| Evaluated graph-backed context | 5,242,880 tokens; scope described above |
| Full-preset nominal context | 2,048 tokens |
| Attention head dimension | 128 |
| FFN width | 11,008 |
| Base low-rank width | 320 |
| Static rank-growth capacity | 384 |
| Viral MoE | 16 of 32 layers, eight experts, top-2 routing |
| Low-rank SwiGLU | Remaining 16 layers |
| Swarm stages | After layers 8, 16, 24, and 32 |
| Swarm shape | 16 cells, top-4 routing, cell hidden width 32, clone interval 500 |
The four swarm stages, embedding, output, modality paths, recursive model, and other top-level organs are all owned by the same trainable runtime. Every named parameter in the reconstructed production model requires gradients.
The four modality tokens are <audio> (32000), </audio> (32001), <image>
(32002), and </image> (32003).
Attention topology
Each transformer layer reserves eight heterogeneous attention slots, for 256 slots across the model. The exact checkpoint has 216 active heads:
- Layers 0-23 have eight active heads each, with ranks
[320, 320, 128, 128, 128, 128, 128, 128]. - Layers 24-27 have four active heads each, with ranks
[320, 320, 128, 128]. - Layers 28-31 have two active heads each, with ranks
[320, 320]. - Every layer's reserved slot ranks are
[320, 320, 128, 128, 128, 128, 128, 128].
The attention path combines a complex quantum anchor (stored as separate real and imaginary tensors) with a synchronized real neurogenic companion. The companion rank is capped at 64, its contribution starts at 0.02 and is capped at 0.20, and the synchronization auxiliary weight is 0.10.
Every block also contains live short-term plasticity, a Hadamard spectral path with phase rotation, quaternion interference, hyperbolic geometry, and a neural-cellular-automata complement. The NCA complement uses rank 128, four steps, a four-step ceiling, and a neighborhood size of five.
Multimodal pathways
Audio
Native BulmaX audio path (always in the checkpoint):
- 16 kHz, 80-bin mel front end.
- FFT size 400 and hop length 160.
- Convolution widths 256 and 512.
- Four eight-head Conformer blocks.
- Up to 128 audio tokens projected into the shared 4,096-wide space.
- Continuous waveform decoder plus a 256-class mu-law codec head.
YuE2 residual audio graft (optional, flag-gated — not “native YuE2”):
From step 156750 onward the live IBM resume line trains a small residual bridge into YuE2-sized latents. Call it a graft, not native YuE2 audio.
| Piece | Detail |
|---|---|
| Geometry | BulmaX hidden 4096 → 64 → 4096 (audio_graft.yue_encoder_proj / yue_decoder_bridge) plus per-layer audio_graft.alpha (32), zero-init so alpha=0 is identity |
| Teacher | Frozen YuE2-Vae latents aligned to BulmaX audio tokens (BULMAX_AUDIO_GRAFT_TEACHER=yue2) |
| Train loss | Dedicated audio_graft_loss (logged in trainer/vitals JSON and TensorBoard as train/audio_graft_loss) plus ` |
| A/B listen | Rank-0 writes four FLACs under /workspace/artifacts/audio/listen/<run>/ (outside the run dir): input, vae_roundtrip, teacher, graft — reconstruction A/B, not text→song generation yet |
| Flags | BULMAX_AUDIO_GRAFT=0|1, BULMAX_AUDIO_GRAFT_TRAINABLE=0|1, BULMAX_AUDIO_GRAFT_TEACHER=yue2, BULMAX_AUDIO_GRAFT_LOSS_WEIGHT (default 0.1) |
| License | YuE2-3B / YuE2-Vae weights are CC BY-NC 4.0 (attribution required; commercial use needs m-a-p permission). Repo licenses/ MIT files cover third-party code only, not the weights. Sources: m-a-p/YuE2-3B, m-a-p/YuE2-Vae |
Older Hub pins (e.g. 156001 and below) do not contain audio_graft.* tensors.
Enabling the graft on a pin that lacks those keys cold-inits them; training then
fills them. Full text-prompt → YuE2 song generation through the graft is not
wired yet — only the reconstruction / latent-alignment path is.
Vision
- 224 x 224 input resolution with 16 x 16 patches (196 patches).
- Patch width 1,024.
- Six 16-head transformer blocks projected into the shared 4,096-wide space.
- Continuous 256 x 256 RGB decoder plus a frozen FSQ vision-codec head.
The language, audio, and vision pathways train through the shared multimodal objective rather than as unrelated side models.
Persistent coherence and long-term context
BulmaX's long-term context is a learned graph and is part of the checkpoint, not a disposable cache. A resumable checkpoint requires both coherence files:
.coherence.pt: 59 tensors, 4,327,310 elements..coherence_graph.bin:COHRformat v4, 131,840 nodes and 600,078,097 edges.
The graph uses an architecture-matched node capacity of
768 + (4096 x 32) = 131,840. The production runtime restores it through the
custom native coherence engine and verifies its structure before creating the
optimizer.
Required runtime systems
The production registry requires all 27 source modules below. The launch contract rejects nonempty disable lists and requires every subsystem switch to remain enabled.
Core and modality: bulmax_model, bulmax_generative,
bulmax_multimodal, bulmax_audio_encoder, bulmax_codec_head,
preprocess_media.
Attention and geometry: neural_cellular_automata,
multidim_interference, hadamard_attention, hyperbolic,
short_term_plasticity, synchronized_attention.
Adaptive and cognitive: recursive_self_model, liquid_network,
valence, autopoietic, criticality, noprop, viral_expert,
swarm_immune, stigmergy, evolution, ewc, meta_optimizer, r_zero.
Coherence and data lineage: coherence_multimodal_loss, resumable_data.
bulmax_vision_encoder.py is also a direct dependency of the multimodal model,
although it is not a separate row in the 27-entry production registry.
Together these provide the recursive self-model, emotional valence/urgency/attachment state, context hypernetwork, liquid cross-layer connectivity, autotelic curiosity, four autopoietic layers, NoProp paths in every block, criticality control, differentiable CUDA coherence, elastic weight consolidation, meta learning-rate control, R-Zero curriculum, stigmergic routing with FlowMutator/Red Queen dynamics, structural evolution, swarm adaptation, and stochastic-rounding optimization.
SBP data augmentation and exact resumable data cursors are part of the same training lifecycle. Cadenced systems remain armed between their scheduled updates; cadence does not mean disabled.
Exact step-155505 bundle
A checkpoint is valid only when all four same-step members are present.
| File | Bytes | SHA-256 |
|---|---|---|
bulmax_sft_step_155505.safetensors |
24,410,636,342 | d221921192d34420d7c5b1852150ccaea191ca2315db6dd71a027a9e5a53dcf9 |
bulmax_sft_step_155505.training_state.pt |
62,924,899,535 | 1d7dcfbc673a2f3c7ef77a640192b2d838bf09a8146890780e5b967fd564b5f8 |
bulmax_sft_step_155505.coherence.pt |
15,228,402 | d0d8978bba3c404981e2afe906f872eaa026150d17daebeb4950291e5647cba0 |
bulmax_sft_step_155505.coherence_graph.bin |
10,205,019,245 | d6bfcf8b26a2a948e909f2053f62fef26787952002acf9c5eaf824bbfd60fdc4 |
| Complete bundle | 97,555,783,524 | Four-file verification required |
The training-state sidecar contains 16,793 optimizer-state entries for 16,795
saved parameter IDs, five parameter groups of sizes [7428, 32, 32, 9296, 7], their
learning rates, scheduler and data-cursor state, and per-rank Python, NumPy,
CPU Torch, and CUDA RNG continuation state. The two intentionally cold saved
entries belong to the coherence sensory projection. This file is part of exact
resumption, not an optional training log.
Loading and continuation
This repository cannot be loaded with AutoModel.from_pretrained or any stock
Transformers architecture. Exact continuation requires:
- The matching custom BulmaX source and production registry.
- The 32k SentencePiece assets and FSQ tokenizer.
- All four same-step checkpoint files listed above.
- A compatible PyTorch/CUDA build and the custom native coherence runtime.
- Receipt-bound validation of source, checkpoint hashes, data lineage, optimizer ownership, graph structure, and subsystem configuration.
The historical single-B200 resume profile uses BF16 FSDP, activation checkpointing, an EMA teacher, and the CUDA/Triton fused AdamW optimizer with stochastic rounding. The optimizer keeps BF16 parameter and moment updates unbiased without allocating a separate FP32 master-weight copy.
Status and limitations
- The step-155505 checkpoint and its sidecars are exact training artifacts, and the grafted systems they contain have been trained rather than left at zero contribution.
- The YuE2 audio graft is a separate, flag-gated residual path. It does not replace the native mel/Conformer/mu-law stack, and it is not a claim of built-in YuE2 song generation. Commercial use of YuE2 weights requires a license beyond CC BY-NC 4.0.
- This model uses custom research code and checkpoint formats. Standard model loaders and hosted inference APIs are not compatible.
- No standardized benchmark, safety, or state-of-the-art claim is made by this card. Those require separate evaluation receipts.
- Possession of the public weights does not provide the private source or a complete runnable environment.