Instructions to use AlexWortega/tinyvla with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- LeRobot
How to use AlexWortega/tinyvla with LeRobot:
- Notebooks
- Google Colab
- Kaggle
TinyVLA: a dual-rate VLA that beats OpenVLA-7B on LIBERO-Spatial at 0.64B params
Headline result: 86% success on LIBERO-Spatial (n=100) in dual-rate mode (semantic latent refreshed at 1 Hz), 79% full-rate β vs OpenVLA-7B's 84.7% β with 140M parameters on the online control path. Built and trained end-to-end in ~5 days on a single RTX A6000.
1. Architecture
SLOW / SMART path (0.5β2 Hz, cached)
instruction ββββββββββββββββββββββ
camera β Qwen3.5-0.8B vision tower β Qwen3.5 language stack (first 12 of 24 layers)
β β
β 8 readout tokens β linear β semantic latent (8Γ512)
β β cached between refreshes
βββββββββββββββββββΌββββββββββββββββββββββββΌββββββββββββββββββββββββββ
β FAST / MOTOR path (10β20 Hz, ~140M online)
βββ 64 spatial tokens per camera ββ
proprio (padded 16d) ββββββββββββββββββββββββββββββββ€
embodiment-id embedding βββββββββββββββββββββββββββββ€
βΌ
flow-matching action expert (38M, 12 blocks,
interleaved cross-/self-attention, AdaLN-free)
βΌ
50-step action chunk (padded 8d + validity mask)
- Single shared encoder: the Qwen3.5-0.8B native ViT (100.6M) feeds both paths.
- Slow path: image + instruction + 8 trailing pad-token "readout" slots through the truncated language stack (L/2 layer-skip a la SmolVLA); their hidden states β latent 8Γ512.
- Fast path: raw vision-tower spatial tokens (bypassing the LM) + proprio + embodiment token, cross-attended by the flow expert. 10 Euler steps at inference.
- Flow matching: linear-OT path, t ~ Beta(1.5, 1), velocity target, masked MSE.
- Totals: 643M stored / 39.5Mβ575M trainable depending on stage / ~140M online per tick.
- Implemented as an out-of-tree LeRobot policy plugin
(
--policy.type=tinyvla --policy.discover_packages_path=tinyvla, lerobot 0.6.1).
2. Training recipe
| Stage | Data | Steps | Trainable | Result |
|---|---|---|---|---|
| 1. Overfit smoke | svla_so100_stacking (56 eps) | 5k | expert+proj (40M) | loss 1.885β0.099 |
| 2. Mixture pretrain | 40 SO100/101 community datasets (~10.7K eps, weight 0.65, per-dataset proportional) + Bridge V2 10K eps (0.35) | 40k, bs 64Γ2, lr 1e-4 | 39.5M (backbone frozen) | held-out RMSE@t0 0.30Ο; stale-latent 2s == fresh |
| 2b. Staleness aug | same | from step 20k | β | latent from frame tβk, k~U(0,2s), p=0.5 |
| 3. Spatial distillation | + 132K teacher labels (Qwen3.5-4B, center-point β 32Γ32 grid), CE Γ0.15 on cam0 tokens | 20k, lr 5e-5 | same + head | head top-1 20.8%, top-5 44.8% (random 0.1%) |
| 4. LIBERO FT (ft4) | HuggingFaceVLA/libero (1693 eps, all suites) | 40k, bs 48, lr 1e-4 / backbone 0.1Γ | 575M (LM layers + vision unfrozen; vocab embeddings frozen) | loss 0.066 |
Ablation FTs: ft1 (all frozen) β 0% closed-loop; ft2 (vision unfrozen) β 60%; ft3 (= ft2 from distilled ckpt) β 61%; ft4 (LM+vision unfrozen) β 79/86%. Unfreezing the LM was the single largest gain; distillation was neutral on LIBERO (likely overwritten by vision unfreezing during FT β bake it into pretraining next time).
Cross-embodiment handling: canonical schema (cam0=scene/cam1=wrist via regex mapping, per-dataset mean/std normalization, action pad-to-8 + validity mask, state pad-to-16, learned embodiment-id token). Weighted mixture sampler with per-source retry on corrupt frames (IndexError/KeyError/decode RuntimeError).
3. Results
LIBERO β all four suites (dual-rate, latent @ 1 Hz)
| Suite | TinyVLA | OpenVLA-7B (published) |
|---|---|---|
| Spatial | 86% (n=100; full-rate 79%) | 84.7% |
| Object | 94% (n=50) | 88.4% |
| Goal | 86% (n=50) | 79.2% |
| Long (libero_10) | 70% (n=50) | 53.7% |
| Average | 84.0% | 76.5% |
Beats OpenVLA-7B on every suite at 1/11 the parameters and ~1/50 the online compute. XS-VLA (90.3 avg) remains ahead β their latent flow matching is future work for us.
LIBERO-Spatial (100 episodes, 10 per task, init states 0-9)
| Mode | Success |
|---|---|
| Full-rate (latent every chunk) | 79/100 = 79% |
| Dual-rate (latent @ 1 Hz) | 86/100 = 86% |
Dual-rate β₯ full-rate consistently across all evals (also 65% vs 58% at ft3, n=50) β a stable cached latent appears to act as temporal regularization.
Context (published numbers, their own eval protocols)
| Model | Params | LIBERO-Spatial | 4-suite avg |
|---|---|---|---|
| MolmoAct2-LIBERO | 5B | 98.4% (97.2% re-measured on our rig, n=500) | 98.25% |
| Ο0 | 3.3B | 98% | β |
| SmolVLA-2.25B | 2.25B | 93% | β |
| XS-VLA | 0.26B | ~93% | 90.3% |
| MolmoAct-7B (v1) | 7B | 87.0% | 86.6% |
| TinyVLA (this) | 0.64B (140M online) | 86% dual-rate | 84.0% |
| OpenVLA | 7B | 84.7% | 76.5% |
MolmoAct2 context (from their cards/repo, no numeric tables published outside the LIBERO-LeRobot card): 5B params on Molmo2-ER backbone + flow expert; claims to beat Ο0.5; Molmo2-ER reportedly surpasses GPT-5 / Gemini-Robotics-ER-1.5 on 13 embodied-reasoning benchmarks; #1 VLA on the MolmoSpace leaderboard. Inference needs ~26GB fp32 (16GB bf16), cuda-graphs for ~2Γ speedup β H100-class serving vs our Raspberry-Pi-class 140M fast path. Notably they ship MolmoAct2-SO100_101 (absolute joint control) β a 5B competitor on our target embodiment; a head-to-head offline eval on SO101 held-out data is an obvious next test.
Honesty notes: our n=100 (Β±8pp) vs their 500; single suite; eval covers init states 0β9 per task (demos of same tasks/init states are in training data β this is the standard LIBERO protocol used by all listed models). Multi-suite evals (object/goal/long) and an identically-protocoled SmolVLA baseline run are in progress.
Dual-rate validation (the architectural claim)
- Offline: held-out chunk MSE with latent from a frame 2s earlier == fresh (0.772 vs 0.772).
- Closed-loop: 86% @ 1 Hz refresh vs 79% full-rate.
- Deployment implication (Raspberry Pi 5): the 0.5B slow path can run at 0.5β1 Hz (llama.cpp GGUF Q4 territory), only the 140M fast path needs real-time rates.
4. Pitfalls we hit (read before reproducing)
- robosuite/LIBERO renders are 180Β°-rotated vs recorded datasets. A vertical flip
alone leaves a horizontal MIRROR β the policy reaches to the wrong side and scores 0%.
Use
img[::-1, ::-1]. (lerobot'sLiberoEnv.render()flips for visualization but_format_raw_obsfeeds policies raw rotated frames β check your eval path.) - eef orientation antipodal ambiguity: env quat β rotvec gives βΟ-ish x-component, datasets store +Ο-ish. Canonicalize (flip antipodal representation when rotvec_x < 0) or your normalized state input is off by ~17Ο.
- HuggingFaceVLA/libero meta says fps=10 but frames are native 20 Hz (~161 frames/episode β 8s demos). Do NOT action-repeat at eval.
- GT-replay harness is the decisive convention test: replay dataset actions in the env (first episode per task β init_state 0). 4/5 success = conventions right; policy-independent.
- HuggingFaceVLA community_dataset v1/v3 are aggregates of v2.1 sub-datasets β convert with
lerobot's
convert_dataset_v21_to_v30.py; 4/44 sources have metaβ parquet row counts (broken at source) β validatelen(reader.hf_dataset) == meta.total_frames. - Qwen3.5 custom forwards need
mm_token_type_idsfor M-RoPE; patchify must match the official processor bit-exact (verify!). Linear-attention needs fla-core + causal-conv1d or you get a 5Γ slower torch fallback. - Flow-matching loss can drop while conditioning is ignored β A/B with shuffled observations to verify the model actually uses them.
5. Repository layout
src/tinyvla/ # LeRobot plugin: configuration/modeling/processor + modules
scripts/train.py # accelerate loop: weighted mixture, staleness aug, distill loss, resume
scripts/eval_libero.py # closed-loop LIBERO eval (all convention fixes, dual-rate, video)
scripts/replay_libero.py, diag_obs_gap.py # convention diagnostics
scripts/prepare_community_v3.py, label_spatial.py, eval_offline.py
configs/ # stage2/stage3/libero_ft* yaml
checkpoints/ # stage2_final, stage3_final, libero_ft4_final (main)
videos/ # sample eval episodes
6. Quick start
from tinyvla.modeling_tinyvla import TinyVLAPolicy
policy = TinyVLAPolicy.from_pretrained("AlexWortega/tinyvla") # ft4 checkpoint at root
# closed-loop LIBERO: python scripts/eval_libero.py --checkpoint <ckpt> \
# --suite libero_spatial --episodes 100 --n-action-steps 10 --refresh-s 1.0
Trained 2026-08-14 β 2026-08-19 on 1Γ RTX A6000 (48GB), lerobot 0.6.1, torch 2.10, transformers 5.5.4.
- Downloads last month
- 291