Text Generation
Transformers
Safetensors
GGUF
English
granite
granite-4.2
formal-logic
reasoning
lora
model-merging
wise-ft
reinforcement-learning
grpo
conversational
Instructions to use webAI-Official/TwIL-LM3-Pro with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use webAI-Official/TwIL-LM3-Pro with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="webAI-Official/TwIL-LM3-Pro") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("webAI-Official/TwIL-LM3-Pro") model = AutoModelForCausalLM.from_pretrained("webAI-Official/TwIL-LM3-Pro", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use webAI-Official/TwIL-LM3-Pro with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf webAI-Official/TwIL-LM3-Pro:Q4_K_M # Run inference directly in the terminal: llama cli -hf webAI-Official/TwIL-LM3-Pro:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf webAI-Official/TwIL-LM3-Pro:Q4_K_M # Run inference directly in the terminal: llama cli -hf webAI-Official/TwIL-LM3-Pro:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf webAI-Official/TwIL-LM3-Pro:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf webAI-Official/TwIL-LM3-Pro:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf webAI-Official/TwIL-LM3-Pro:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf webAI-Official/TwIL-LM3-Pro:Q4_K_M
Use Docker
docker model run hf.co/webAI-Official/TwIL-LM3-Pro:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use webAI-Official/TwIL-LM3-Pro with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "webAI-Official/TwIL-LM3-Pro" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "webAI-Official/TwIL-LM3-Pro", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/webAI-Official/TwIL-LM3-Pro:Q4_K_M
- SGLang
How to use webAI-Official/TwIL-LM3-Pro with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "webAI-Official/TwIL-LM3-Pro" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "webAI-Official/TwIL-LM3-Pro", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "webAI-Official/TwIL-LM3-Pro" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "webAI-Official/TwIL-LM3-Pro", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Ollama
How to use webAI-Official/TwIL-LM3-Pro with Ollama:
ollama run hf.co/webAI-Official/TwIL-LM3-Pro:Q4_K_M
- Unsloth Desktop
- Pi
How to use webAI-Official/TwIL-LM3-Pro with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf webAI-Official/TwIL-LM3-Pro:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "webAI-Official/TwIL-LM3-Pro:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use webAI-Official/TwIL-LM3-Pro with Docker Model Runner:
docker model run hf.co/webAI-Official/TwIL-LM3-Pro:Q4_K_M
- Lemonade
How to use webAI-Official/TwIL-LM3-Pro with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull webAI-Official/TwIL-LM3-Pro:Q4_K_M
Run and chat with the model
lemonade run user.TwIL-LM3-Pro-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use webAI-Official/TwIL-LM3-Pro with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf webAI-Official/TwIL-LM3-Pro:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default webAI-Official/TwIL-LM3-Pro:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use webAI-Official/TwIL-LM3-Pro with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf webAI-Official/TwIL-LM3-Pro:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "webAI-Official/TwIL-LM3-Pro:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Rename VibeThinker-trained to VibeThinker-webAI-trained
Browse files- README.md +14 -14
- benchmarks.png +2 -2
README.md
CHANGED
|
@@ -35,9 +35,9 @@ gate, strict-7 and six-lane average of any arm in the tables below for which eac
|
|
| 35 |
computed, including Qwen3-8B, Qwen3.5-4B, VibeThinker-3B and gpt-oss-120b (the 120B has no gate or strict-7
|
| 36 |
value).
|
| 37 |
|
| 38 |
-

|
| 39 |
|
| 40 |
-
*VibeThinker-trained is the public VibeThinker-3B after the same post-training pipeline (SLERP, d = 0.5, t = 0.5); see [Against the tuned VibeThinker-3B](#against-the-tuned-vibethinker-3b). It has no strict-7 or six-lane-average value, so those bars show n/a.*
|
| 41 |
|
| 42 |
## Highlights
|
| 43 |
|
|
@@ -66,7 +66,7 @@ value).
|
|
| 66 |
comes entirely from BBH-logic (0.9540 against 0.6107); without that row TwIL-LM3-Pro trails.
|
| 67 |
* **The pipeline is not tied to one model.** The same recipe was run on five base models. On
|
| 68 |
VibeThinker-3B it lifts the Track A macro gate from 0.374 to 0.508 (SLERP, the
|
| 69 |
-
*VibeThinker-trained* column) while the Track B 10-dataset macro moves from 0.815 to 0.802 — see
|
| 70 |
[One pipeline, several models](#one-pipeline-several-models).
|
| 71 |
* **Structured formal output.** Tuned for the objects rather than the prose: FOL translation,
|
| 72 |
entailment labels, semantic parses, Lean statements and Lean proof critique.
|
|
@@ -115,7 +115,7 @@ dedicated decode-throughput protocol: `ans/s` is defined throughout as
|
|
| 115 |
`tok/s ÷ mean generation length`, so it measures completed answers rather than raw decode rate.
|
| 116 |
Cells marked † need the engine note below.
|
| 117 |
|
| 118 |
-
| lane / metric | TwIL-LM3-Pro | Granite-4.2-3B base | VibeThinker-3B | VibeThinker-trained ★ | Qwen3.5-4B | TwIL-LM3 | Llama-3.2-3B | LFM2-2.6B | LFM2.5-8B-A1B | Qwen3-8B | gpt-oss-120b ‡ |
|
| 119 |
|---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
|
| 120 |
| lean_formalize token_f1 | 0.5092 | 0.2943 | 0.2087 | 0.527 | 0.4996 | 0.5869 | 0.3690 | 0.1321 | 0.4655 | 0.4022 | **0.6306** |
|
| 121 |
| rule_induction derivation | 0.4195 | 0.2267 | 0.2038 | 0.227 | 0.5078 | 0.3192 | 0.0825 | 0.0615 | 0.1936 | 0.3680 | **0.6518** |
|
|
@@ -142,7 +142,7 @@ average cannot be computed for it; that is what the — cells mean, not a zero.
|
|
| 142 |
format and tokenizer make the corpus lanes score a different quantity. The number is reported
|
| 143 |
for completeness but is not a comparable measurement, and is excluded from the bolding.
|
| 144 |
|
| 145 |
-
★ **VibeThinker-trained** is the public VibeThinker-3B after the same post-training pipeline as
|
| 146 |
TwIL-LM3-Pro, in its SLERP (d = 0.5, t = 0.5) configuration; it is not a checkpoint in this
|
| 147 |
repository, and [the section below](#against-the-tuned-vibethinker-3b) explains how it differs
|
| 148 |
from the base model. Its figures are taken from the internal family comparison tables, to three
|
|
@@ -223,7 +223,7 @@ spots in absolute terms are `procedural` (strict 0.1200, loose 0.2350) and FOL t
|
|
| 223 |
|
| 224 |
### Track B — held-out benchmarks
|
| 225 |
|
| 226 |
-
| dataset | TwIL-LM3-Pro | Granite-4.2-3B base | VibeThinker-3B | VibeThinker-trained ★ | Qwen3.5-4B | TwIL-LM3 | Llama-3.2-3B | LFM2-2.6B | LFM2.5-8B-A1B | Qwen3-8B | gpt-oss-120b ‡ |
|
| 227 |
|---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
|
| 228 |
| gsm8k | 0.9433 | 0.9533 | 0.9600 | 0.930 | 0.8633 | 0.8733 | 0.8300 | 0.8767 | 0.9133 | 0.9567 | **0.9767** |
|
| 229 |
| svamp | 0.9500 | 0.9200 | 0.9367 | **0.953** | 0.8867 | 0.8500 | 0.8200 | 0.9000 | 0.9133 | 0.9400 | 0.9400 |
|
|
@@ -292,19 +292,19 @@ MuSR-team sit slightly above the 2% cap-hit threshold (3.0%, 3.0% and 2.8%).
|
|
| 292 |
The tables above use the public VibeThinker-3B checkpoint. The same post-training pipeline was also
|
| 293 |
applied to it, and two of its tuned configurations are the closest same-scale comparisons to
|
| 294 |
TwIL-LM3-Pro: WiSE-FT (λ = 0.50), and the SLERP merge (d = 0.5, t = 0.5) that appears in the
|
| 295 |
-
tables and plot as **VibeThinker-trained**. These values come from the family comparison tables
|
| 296 |
rather than from a per-lane raw report, so they are shown as a summary only:
|
| 297 |
|
| 298 |
| model | macro gate | macro_primary | B10 | B14 | Track A truncation |
|
| 299 |
|---|---:|---:|---:|---:|---:|
|
| 300 |
| TwIL-LM3-Pro | **0.554** | **0.588** | 0.790 | **0.743** | 24.2% |
|
| 301 |
-
| VibeThinker-trained (SLERP, d = 0.5, t = 0.5) | 0.508 | 0.579 | **0.802** | 0.728 | — |
|
| 302 |
| VibeThinker-3B, WiSE-FT λ = 0.50 | 0.541 | **0.588** | **0.802** | 0.728 ◊ | 14.3% |
|
| 303 |
|
| 304 |
◊ There is no B14 row for the λ = 0.50 configuration; the figure is the one recorded for the SLERP
|
| 305 |
configuration in the row above. Truncation was not recorded for the SLERP configuration.
|
| 306 |
|
| 307 |
-
Against VibeThinker-trained, TwIL-LM3-Pro is ahead on Track A (macro gate 0.554 against 0.508,
|
| 308 |
`macro_primary` 0.588 against 0.579) and on the 14-dataset macro (0.743 against 0.728), and behind
|
| 309 |
on the 10-dataset macro (0.790 against 0.802). Against the WiSE-FT λ = 0.50 configuration the two
|
| 310 |
are effectively tied on Track A (gate 0.554 against 0.541, `macro_primary` equal at 0.588, both
|
|
@@ -312,10 +312,10 @@ within sampling noise at n = 200), and the tuned VibeThinker-3B is ahead on the
|
|
| 312 |
with a lower truncation rate. TwIL-LM3-Pro's edge is the 14-dataset macro, a gap that cannot be
|
| 313 |
broken down per dataset from the summary values.
|
| 314 |
|
| 315 |
-
#### How VibeThinker-trained differs from the base VibeThinker-3B
|
| 316 |
|
| 317 |
**Base VibeThinker-3B** is WeiboAI's public checkpoint, unmodified, and is what the untuned columns
|
| 318 |
-
in the tables above measure. **VibeThinker-trained** starts from those same weights and changes
|
| 319 |
them in two ways:
|
| 320 |
|
| 321 |
* **Formal-logic post-training.** A rank-64 LoRA is trained on the same synthetic formal-logic
|
|
@@ -328,13 +328,13 @@ them in two ways:
|
|
| 328 |
held-out capability in TwIL-LM3-Pro.
|
| 329 |
|
| 330 |
The family tables record no reinforcement-learning (MGPO) run for VibeThinker-3B, so
|
| 331 |
-
VibeThinker-trained reflects the supervised and merging stages only, whereas TwIL-LM3-Pro also has
|
| 332 |
the MGPO stage. It is a reference point for the pipeline, not a checkpoint shipped in this
|
| 333 |
repository.
|
| 334 |
|
| 335 |
What that changes, on the family tables (one source, so the comparison is like for like):
|
| 336 |
|
| 337 |
-
| metric | VibeThinker-3B (base) | VibeThinker-trained | change |
|
| 338 |
|---|---:|---:|---:|
|
| 339 |
| macro gate | 0.374 | 0.508 | +0.134 |
|
| 340 |
| macro_primary | 0.444 | 0.579 | +0.135 |
|
|
@@ -509,7 +509,7 @@ card, which describes the same harness. For Track B, the arms checked (including
|
|
| 509 |
and Qwen3.5-4B, on all 18 tasks) share the same sampled rows and decoding, but the serving engine differs between
|
| 510 |
arms (vLLM 0.19.1 for TwIL-LM3-Pro, its base, Qwen3.5-4B, Qwen3-8B and LFM2.5-8B-A1B; vLLM 0.11.2 for
|
| 511 |
TwIL-LM3, Llama-3.2-3B and VibeThinker-3B), and the engine version is part of the protocol hash.
|
| 512 |
-
The VibeThinker-trained column (★) comes from the internal family comparison tables and is not
|
| 513 |
covered by the manifest checks described here. Throughput has its own, separate engine caveat (see the † note under the Track A table). With
|
| 514 |
n = 200 per lane on Track A and n = 300 per dataset on Track B, differences of two to three points
|
| 515 |
are within sampling noise.
|
|
|
|
| 35 |
computed, including Qwen3-8B, Qwen3.5-4B, VibeThinker-3B and gpt-oss-120b (the 120B has no gate or strict-7
|
| 36 |
value).
|
| 37 |
|
| 38 |
+

|
| 39 |
|
| 40 |
+
*VibeThinker-webAI-trained is the public VibeThinker-3B after the same post-training pipeline (SLERP, d = 0.5, t = 0.5); see [Against the tuned VibeThinker-3B](#against-the-tuned-vibethinker-3b). It has no strict-7 or six-lane-average value, so those bars show n/a.*
|
| 41 |
|
| 42 |
## Highlights
|
| 43 |
|
|
|
|
| 66 |
comes entirely from BBH-logic (0.9540 against 0.6107); without that row TwIL-LM3-Pro trails.
|
| 67 |
* **The pipeline is not tied to one model.** The same recipe was run on five base models. On
|
| 68 |
VibeThinker-3B it lifts the Track A macro gate from 0.374 to 0.508 (SLERP, the
|
| 69 |
+
*VibeThinker-webAI-trained* column) while the Track B 10-dataset macro moves from 0.815 to 0.802 — see
|
| 70 |
[One pipeline, several models](#one-pipeline-several-models).
|
| 71 |
* **Structured formal output.** Tuned for the objects rather than the prose: FOL translation,
|
| 72 |
entailment labels, semantic parses, Lean statements and Lean proof critique.
|
|
|
|
| 115 |
`tok/s ÷ mean generation length`, so it measures completed answers rather than raw decode rate.
|
| 116 |
Cells marked † need the engine note below.
|
| 117 |
|
| 118 |
+
| lane / metric | TwIL-LM3-Pro | Granite-4.2-3B base | VibeThinker-3B | VibeThinker-webAI-trained ★ | Qwen3.5-4B | TwIL-LM3 | Llama-3.2-3B | LFM2-2.6B | LFM2.5-8B-A1B | Qwen3-8B | gpt-oss-120b ‡ |
|
| 119 |
|---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
|
| 120 |
| lean_formalize token_f1 | 0.5092 | 0.2943 | 0.2087 | 0.527 | 0.4996 | 0.5869 | 0.3690 | 0.1321 | 0.4655 | 0.4022 | **0.6306** |
|
| 121 |
| rule_induction derivation | 0.4195 | 0.2267 | 0.2038 | 0.227 | 0.5078 | 0.3192 | 0.0825 | 0.0615 | 0.1936 | 0.3680 | **0.6518** |
|
|
|
|
| 142 |
format and tokenizer make the corpus lanes score a different quantity. The number is reported
|
| 143 |
for completeness but is not a comparable measurement, and is excluded from the bolding.
|
| 144 |
|
| 145 |
+
★ **VibeThinker-webAI-trained** is the public VibeThinker-3B after the same post-training pipeline as
|
| 146 |
TwIL-LM3-Pro, in its SLERP (d = 0.5, t = 0.5) configuration; it is not a checkpoint in this
|
| 147 |
repository, and [the section below](#against-the-tuned-vibethinker-3b) explains how it differs
|
| 148 |
from the base model. Its figures are taken from the internal family comparison tables, to three
|
|
|
|
| 223 |
|
| 224 |
### Track B — held-out benchmarks
|
| 225 |
|
| 226 |
+
| dataset | TwIL-LM3-Pro | Granite-4.2-3B base | VibeThinker-3B | VibeThinker-webAI-trained ★ | Qwen3.5-4B | TwIL-LM3 | Llama-3.2-3B | LFM2-2.6B | LFM2.5-8B-A1B | Qwen3-8B | gpt-oss-120b ‡ |
|
| 227 |
|---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
|
| 228 |
| gsm8k | 0.9433 | 0.9533 | 0.9600 | 0.930 | 0.8633 | 0.8733 | 0.8300 | 0.8767 | 0.9133 | 0.9567 | **0.9767** |
|
| 229 |
| svamp | 0.9500 | 0.9200 | 0.9367 | **0.953** | 0.8867 | 0.8500 | 0.8200 | 0.9000 | 0.9133 | 0.9400 | 0.9400 |
|
|
|
|
| 292 |
The tables above use the public VibeThinker-3B checkpoint. The same post-training pipeline was also
|
| 293 |
applied to it, and two of its tuned configurations are the closest same-scale comparisons to
|
| 294 |
TwIL-LM3-Pro: WiSE-FT (λ = 0.50), and the SLERP merge (d = 0.5, t = 0.5) that appears in the
|
| 295 |
+
tables and plot as **VibeThinker-webAI-trained**. These values come from the family comparison tables
|
| 296 |
rather than from a per-lane raw report, so they are shown as a summary only:
|
| 297 |
|
| 298 |
| model | macro gate | macro_primary | B10 | B14 | Track A truncation |
|
| 299 |
|---|---:|---:|---:|---:|---:|
|
| 300 |
| TwIL-LM3-Pro | **0.554** | **0.588** | 0.790 | **0.743** | 24.2% |
|
| 301 |
+
| VibeThinker-webAI-trained (SLERP, d = 0.5, t = 0.5) | 0.508 | 0.579 | **0.802** | 0.728 | — |
|
| 302 |
| VibeThinker-3B, WiSE-FT λ = 0.50 | 0.541 | **0.588** | **0.802** | 0.728 ◊ | 14.3% |
|
| 303 |
|
| 304 |
◊ There is no B14 row for the λ = 0.50 configuration; the figure is the one recorded for the SLERP
|
| 305 |
configuration in the row above. Truncation was not recorded for the SLERP configuration.
|
| 306 |
|
| 307 |
+
Against VibeThinker-webAI-trained, TwIL-LM3-Pro is ahead on Track A (macro gate 0.554 against 0.508,
|
| 308 |
`macro_primary` 0.588 against 0.579) and on the 14-dataset macro (0.743 against 0.728), and behind
|
| 309 |
on the 10-dataset macro (0.790 against 0.802). Against the WiSE-FT λ = 0.50 configuration the two
|
| 310 |
are effectively tied on Track A (gate 0.554 against 0.541, `macro_primary` equal at 0.588, both
|
|
|
|
| 312 |
with a lower truncation rate. TwIL-LM3-Pro's edge is the 14-dataset macro, a gap that cannot be
|
| 313 |
broken down per dataset from the summary values.
|
| 314 |
|
| 315 |
+
#### How VibeThinker-webAI-trained differs from the base VibeThinker-3B
|
| 316 |
|
| 317 |
**Base VibeThinker-3B** is WeiboAI's public checkpoint, unmodified, and is what the untuned columns
|
| 318 |
+
in the tables above measure. **VibeThinker-webAI-trained** starts from those same weights and changes
|
| 319 |
them in two ways:
|
| 320 |
|
| 321 |
* **Formal-logic post-training.** A rank-64 LoRA is trained on the same synthetic formal-logic
|
|
|
|
| 328 |
held-out capability in TwIL-LM3-Pro.
|
| 329 |
|
| 330 |
The family tables record no reinforcement-learning (MGPO) run for VibeThinker-3B, so
|
| 331 |
+
VibeThinker-webAI-trained reflects the supervised and merging stages only, whereas TwIL-LM3-Pro also has
|
| 332 |
the MGPO stage. It is a reference point for the pipeline, not a checkpoint shipped in this
|
| 333 |
repository.
|
| 334 |
|
| 335 |
What that changes, on the family tables (one source, so the comparison is like for like):
|
| 336 |
|
| 337 |
+
| metric | VibeThinker-3B (base) | VibeThinker-webAI-trained | change |
|
| 338 |
|---|---:|---:|---:|
|
| 339 |
| macro gate | 0.374 | 0.508 | +0.134 |
|
| 340 |
| macro_primary | 0.444 | 0.579 | +0.135 |
|
|
|
|
| 509 |
and Qwen3.5-4B, on all 18 tasks) share the same sampled rows and decoding, but the serving engine differs between
|
| 510 |
arms (vLLM 0.19.1 for TwIL-LM3-Pro, its base, Qwen3.5-4B, Qwen3-8B and LFM2.5-8B-A1B; vLLM 0.11.2 for
|
| 511 |
TwIL-LM3, Llama-3.2-3B and VibeThinker-3B), and the engine version is part of the protocol hash.
|
| 512 |
+
The VibeThinker-webAI-trained column (★) comes from the internal family comparison tables and is not
|
| 513 |
covered by the manifest checks described here. Throughput has its own, separate engine caveat (see the † note under the Track A table). With
|
| 514 |
n = 200 per lane on Track A and n = 300 per dataset on Track B, differences of two to three points
|
| 515 |
are within sampling noise.
|
benchmarks.png
CHANGED
|
Git LFS Details
|
|
Git LFS Details
|