LucaCerovaz commited on
Commit
f5687aa
·
verified ·
1 Parent(s): e078b53

Add Violetto model card and evaluation table

Browse files

Document the model, published evaluations, Apache-2.0 license, and vLLM plugin usage with the current GitHub installation commands.

Files changed (3) hide show
  1. .gitattributes +1 -0
  2. README.md +92 -0
  3. assets/evaluations.png +3 -0
.gitattributes CHANGED
@@ -34,3 +34,4 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
  tokenizer.json filter=lfs diff=lfs merge=lfs -text
 
 
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
  tokenizer.json filter=lfs diff=lfs merge=lfs -text
37
+ assets/evaluations.png filter=lfs diff=lfs merge=lfs -text
README.md ADDED
@@ -0,0 +1,92 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ pipeline_tag: text-generation
4
+ tags:
5
+ - mathematics
6
+ - reasoning
7
+ - vllm
8
+ ---
9
+
10
+ # Limite 1B - Violetto
11
+
12
+ Paradigma's first model, built for high-throughput mathematical reasoning.
13
+
14
+ [GitHub & vLLM plugin](https://github.com/paradigma-inc/limite-violetto) · [Release blog](https://paradigma.inc/blog/limite-1b-violetto/)
15
+
16
+ ## Overview
17
+
18
+ Limite 1B - Violetto is Paradigma’s first model, designed for high-throughput solutions of difficult mathematical problems.
19
+
20
+ Limite is a 1-billion parameter dense autoregressive transformer, trained from scratch on a mixture of less than 300B highly curated tokens, capable of handling sequences up to 131k tokens long.
21
+
22
+ Apart from data curation, we achieve such high sample efficiency by equipping Limite with an architecture strongly inspired by recent advancements made by the community on pre-training speedrun competitions. The base model itself scores well on math benchmarks with few-shot prompting.
23
+
24
+ We leverage a mix of synthetic data generation, curated SFT and RL post-training to achieve results on competition-level math that rivals recent models tens of times larger, trained with orders of magnitudes more FLOPs. As an example, Limite achieves an average of 74.25% on BeyondAIME, with MUSE-Glimmer-30B scoring 70%. The full table with evaluations is available below.
25
+
26
+ ## Evaluation
27
+
28
+ Results across seven mathematical benchmarks, including AIME 2026, HMMT February 2026, APEX Shortlist, and BeyondAIME. Scores are percentages; higher is better.
29
+
30
+ ![Full evaluation table comparing Limite 1B - Violetto with other models on seven mathematical benchmarks.](assets/evaluations.png)
31
+
32
+ † Marked results are taken from model cards or MathArena and were not reevaluated by our team. A dash indicates an unreported result; “n.d.” indicates an undisclosed parameter count. Evaluation configurations may differ across models and sources.
33
+
34
+ See the [release blog](https://paradigma.inc/blog/limite-1b-violetto/) for the model overview and examples, and the [GitHub repository](https://github.com/paradigma-inc/limite-violetto) for serving code and usage details.
35
+
36
+ ## Run with vLLM
37
+
38
+ This repository contains the Violetto checkpoint and tokenizer. To run the model, use the [Limite vLLM plugin on GitHub](https://github.com/paradigma-inc/limite-violetto).
39
+
40
+ **Python 3.12 · vLLM 0.26.0 · PyTorch 2.11.0 (CUDA 13.0) · Transformers 5.6.2**
41
+
42
+ On Linux x86-64 with an NVIDIA GPU and a CUDA 13.0-compatible NVIDIA driver, install [uv](https://docs.astral.sh/uv/getting-started/installation/), then follow the GitHub quickstart:
43
+
44
+ ```bash
45
+ git clone https://github.com/paradigma-inc/limite-violetto.git
46
+ cd limite-violetto
47
+ uv sync --locked
48
+ ```
49
+
50
+ This installs the plugin and its locked serving dependencies, including vLLM, CUDA-enabled PyTorch, and Transformers. The NVIDIA driver must already be installed on the host; a separate CUDA toolkit installation is not required.
51
+
52
+ From the same directory, start Violetto:
53
+
54
+ ```bash
55
+ VLLM_PLUGINS=limite uv run --locked vllm serve paradigma-inc/limite-1b-violetto
56
+ ```
57
+
58
+ The checkpoint is downloaded from Hugging Face on first use. Keep tensor and pipeline parallel sizes at 1. For full setup instructions and the option to install only the plugin into an existing compatible environment, see the [GitHub README](https://github.com/paradigma-inc/limite-violetto#quickstart).
59
+
60
+ **Hugging Face Transformers compatibility is coming soon.** We will release support for loading and running Violetto directly with Transformers.
61
+
62
+ ## Prompting and intended use
63
+
64
+ Send one mathematical problem as a text-only user message through the chat API. The checkpoint's bundled chat template inserts the canonical mathematical system prompt, which asks for step-by-step reasoning and a final answer inside `\boxed{}`. Use this template when formatting prompts locally as well.
65
+
66
+ The bundled template enforces the fixed system prompt and does not accept custom system prompts or tool calls.
67
+
68
+ Limite is designed to be as lightly instruction-tuned as possible, to challenge the assumption that models need to be embedded in an assistant persona to function well. As a result, Limite is designed to be used to respond in single turns, with an extremely high mathematical capability per parameter count.
69
+
70
+ Limite is designed to be a strong reasoner of mathematical problems. As a result, its response patterns are radically different than that of regular assistants, and shouldn’t be used expecting instruction following in the same form as other, more general-purpose language models.
71
+
72
+ Limite can lose the scope of a prompt and reinterpret it as a different—often mathematical—task.
73
+
74
+ See the [release blog](https://paradigma.inc/blog/limite-1b-violetto/) for examples.
75
+
76
+ ## License
77
+
78
+ The model weights are released under **Apache-2.0**. The [vLLM serving code](https://github.com/paradigma-inc/limite-violetto/blob/main/LICENSE) is also licensed under Apache-2.0.
79
+
80
+ ## Citation
81
+
82
+ ```bibtex
83
+ @misc{paradigma2026limite,
84
+ title = {{Limite 1B - Violetto}},
85
+ author = {Prignano, Mario and Cirillo, Gabriele and
86
+ Morosini, Alessio and Cerovaz, Luca and
87
+ Bartolocci, Alessandro and Rodolà, Emanuele and
88
+ Starace, Giulio and Pappone, Francesco},
89
+ year = {2026},
90
+ howpublished = {\url{https://paradigma.inc/blog/limite-1b-violetto/}}
91
+ }
92
+ ```
assets/evaluations.png ADDED

Git LFS Details

  • SHA256: a11ad8cdea8ea471f82aad7061405582bede0c8a8ee7b13e5273be6dd04ad802
  • Pointer size: 131 Bytes
  • Size of remote file: 514 kB