anurag051194 commited on
Commit
a8247e4
·
verified ·
1 Parent(s): 69d1b3f

Rename VibeThinker-trained to VibeThinker-webAI-trained

Browse files
Files changed (2) hide show
  1. README.md +14 -14
  2. benchmarks.png +2 -2
README.md CHANGED
@@ -35,9 +35,9 @@ gate, strict-7 and six-lane average of any arm in the tables below for which eac
35
  computed, including Qwen3-8B, Qwen3.5-4B, VibeThinker-3B and gpt-oss-120b (the 120B has no gate or strict-7
36
  value).
37
 
38
- ![TwIL-LM3-Pro formal and general reasoning benchmarks against VibeThinker-3B, VibeThinker-trained, Qwen3.5-4B, Qwen3-8B, gpt-oss-120b, LFM2.5-8B-A1B and TwIL-LM3](benchmarks.png)
39
 
40
- *VibeThinker-trained is the public VibeThinker-3B after the same post-training pipeline (SLERP, d = 0.5, t = 0.5); see [Against the tuned VibeThinker-3B](#against-the-tuned-vibethinker-3b). It has no strict-7 or six-lane-average value, so those bars show n/a.*
41
 
42
  ## Highlights
43
 
@@ -66,7 +66,7 @@ value).
66
  comes entirely from BBH-logic (0.9540 against 0.6107); without that row TwIL-LM3-Pro trails.
67
  * **The pipeline is not tied to one model.** The same recipe was run on five base models. On
68
  VibeThinker-3B it lifts the Track A macro gate from 0.374 to 0.508 (SLERP, the
69
- *VibeThinker-trained* column) while the Track B 10-dataset macro moves from 0.815 to 0.802 — see
70
  [One pipeline, several models](#one-pipeline-several-models).
71
  * **Structured formal output.** Tuned for the objects rather than the prose: FOL translation,
72
  entailment labels, semantic parses, Lean statements and Lean proof critique.
@@ -115,7 +115,7 @@ dedicated decode-throughput protocol: `ans/s` is defined throughout as
115
  `tok/s ÷ mean generation length`, so it measures completed answers rather than raw decode rate.
116
  Cells marked † need the engine note below.
117
 
118
- | lane / metric | TwIL-LM3-Pro | Granite-4.2-3B base | VibeThinker-3B | VibeThinker-trained ★ | Qwen3.5-4B | TwIL-LM3 | Llama-3.2-3B | LFM2-2.6B | LFM2.5-8B-A1B | Qwen3-8B | gpt-oss-120b ‡ |
119
  |---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
120
  | lean_formalize token_f1 | 0.5092 | 0.2943 | 0.2087 | 0.527 | 0.4996 | 0.5869 | 0.3690 | 0.1321 | 0.4655 | 0.4022 | **0.6306** |
121
  | rule_induction derivation | 0.4195 | 0.2267 | 0.2038 | 0.227 | 0.5078 | 0.3192 | 0.0825 | 0.0615 | 0.1936 | 0.3680 | **0.6518** |
@@ -142,7 +142,7 @@ average cannot be computed for it; that is what the — cells mean, not a zero.
142
  format and tokenizer make the corpus lanes score a different quantity. The number is reported
143
  for completeness but is not a comparable measurement, and is excluded from the bolding.
144
 
145
- ★ **VibeThinker-trained** is the public VibeThinker-3B after the same post-training pipeline as
146
  TwIL-LM3-Pro, in its SLERP (d = 0.5, t = 0.5) configuration; it is not a checkpoint in this
147
  repository, and [the section below](#against-the-tuned-vibethinker-3b) explains how it differs
148
  from the base model. Its figures are taken from the internal family comparison tables, to three
@@ -223,7 +223,7 @@ spots in absolute terms are `procedural` (strict 0.1200, loose 0.2350) and FOL t
223
 
224
  ### Track B — held-out benchmarks
225
 
226
- | dataset | TwIL-LM3-Pro | Granite-4.2-3B base | VibeThinker-3B | VibeThinker-trained ★ | Qwen3.5-4B | TwIL-LM3 | Llama-3.2-3B | LFM2-2.6B | LFM2.5-8B-A1B | Qwen3-8B | gpt-oss-120b ‡ |
227
  |---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
228
  | gsm8k | 0.9433 | 0.9533 | 0.9600 | 0.930 | 0.8633 | 0.8733 | 0.8300 | 0.8767 | 0.9133 | 0.9567 | **0.9767** |
229
  | svamp | 0.9500 | 0.9200 | 0.9367 | **0.953** | 0.8867 | 0.8500 | 0.8200 | 0.9000 | 0.9133 | 0.9400 | 0.9400 |
@@ -292,19 +292,19 @@ MuSR-team sit slightly above the 2% cap-hit threshold (3.0%, 3.0% and 2.8%).
292
  The tables above use the public VibeThinker-3B checkpoint. The same post-training pipeline was also
293
  applied to it, and two of its tuned configurations are the closest same-scale comparisons to
294
  TwIL-LM3-Pro: WiSE-FT (λ = 0.50), and the SLERP merge (d = 0.5, t = 0.5) that appears in the
295
- tables and plot as **VibeThinker-trained**. These values come from the family comparison tables
296
  rather than from a per-lane raw report, so they are shown as a summary only:
297
 
298
  | model | macro gate | macro_primary | B10 | B14 | Track A truncation |
299
  |---|---:|---:|---:|---:|---:|
300
  | TwIL-LM3-Pro | **0.554** | **0.588** | 0.790 | **0.743** | 24.2% |
301
- | VibeThinker-trained (SLERP, d = 0.5, t = 0.5) | 0.508 | 0.579 | **0.802** | 0.728 | — |
302
  | VibeThinker-3B, WiSE-FT λ = 0.50 | 0.541 | **0.588** | **0.802** | 0.728 ◊ | 14.3% |
303
 
304
  ◊ There is no B14 row for the λ = 0.50 configuration; the figure is the one recorded for the SLERP
305
  configuration in the row above. Truncation was not recorded for the SLERP configuration.
306
 
307
- Against VibeThinker-trained, TwIL-LM3-Pro is ahead on Track A (macro gate 0.554 against 0.508,
308
  `macro_primary` 0.588 against 0.579) and on the 14-dataset macro (0.743 against 0.728), and behind
309
  on the 10-dataset macro (0.790 against 0.802). Against the WiSE-FT λ = 0.50 configuration the two
310
  are effectively tied on Track A (gate 0.554 against 0.541, `macro_primary` equal at 0.588, both
@@ -312,10 +312,10 @@ within sampling noise at n = 200), and the tuned VibeThinker-3B is ahead on the
312
  with a lower truncation rate. TwIL-LM3-Pro's edge is the 14-dataset macro, a gap that cannot be
313
  broken down per dataset from the summary values.
314
 
315
- #### How VibeThinker-trained differs from the base VibeThinker-3B
316
 
317
  **Base VibeThinker-3B** is WeiboAI's public checkpoint, unmodified, and is what the untuned columns
318
- in the tables above measure. **VibeThinker-trained** starts from those same weights and changes
319
  them in two ways:
320
 
321
  * **Formal-logic post-training.** A rank-64 LoRA is trained on the same synthetic formal-logic
@@ -328,13 +328,13 @@ them in two ways:
328
  held-out capability in TwIL-LM3-Pro.
329
 
330
  The family tables record no reinforcement-learning (MGPO) run for VibeThinker-3B, so
331
- VibeThinker-trained reflects the supervised and merging stages only, whereas TwIL-LM3-Pro also has
332
  the MGPO stage. It is a reference point for the pipeline, not a checkpoint shipped in this
333
  repository.
334
 
335
  What that changes, on the family tables (one source, so the comparison is like for like):
336
 
337
- | metric | VibeThinker-3B (base) | VibeThinker-trained | change |
338
  |---|---:|---:|---:|
339
  | macro gate | 0.374 | 0.508 | +0.134 |
340
  | macro_primary | 0.444 | 0.579 | +0.135 |
@@ -509,7 +509,7 @@ card, which describes the same harness. For Track B, the arms checked (including
509
  and Qwen3.5-4B, on all 18 tasks) share the same sampled rows and decoding, but the serving engine differs between
510
  arms (vLLM 0.19.1 for TwIL-LM3-Pro, its base, Qwen3.5-4B, Qwen3-8B and LFM2.5-8B-A1B; vLLM 0.11.2 for
511
  TwIL-LM3, Llama-3.2-3B and VibeThinker-3B), and the engine version is part of the protocol hash.
512
- The VibeThinker-trained column (★) comes from the internal family comparison tables and is not
513
  covered by the manifest checks described here. Throughput has its own, separate engine caveat (see the † note under the Track A table). With
514
  n = 200 per lane on Track A and n = 300 per dataset on Track B, differences of two to three points
515
  are within sampling noise.
 
35
  computed, including Qwen3-8B, Qwen3.5-4B, VibeThinker-3B and gpt-oss-120b (the 120B has no gate or strict-7
36
  value).
37
 
38
+ ![TwIL-LM3-Pro formal and general reasoning benchmarks against VibeThinker-3B, VibeThinker-webAI-trained, Qwen3.5-4B, Qwen3-8B, gpt-oss-120b, LFM2.5-8B-A1B and TwIL-LM3](benchmarks.png)
39
 
40
+ *VibeThinker-webAI-trained is the public VibeThinker-3B after the same post-training pipeline (SLERP, d = 0.5, t = 0.5); see [Against the tuned VibeThinker-3B](#against-the-tuned-vibethinker-3b). It has no strict-7 or six-lane-average value, so those bars show n/a.*
41
 
42
  ## Highlights
43
 
 
66
  comes entirely from BBH-logic (0.9540 against 0.6107); without that row TwIL-LM3-Pro trails.
67
  * **The pipeline is not tied to one model.** The same recipe was run on five base models. On
68
  VibeThinker-3B it lifts the Track A macro gate from 0.374 to 0.508 (SLERP, the
69
+ *VibeThinker-webAI-trained* column) while the Track B 10-dataset macro moves from 0.815 to 0.802 — see
70
  [One pipeline, several models](#one-pipeline-several-models).
71
  * **Structured formal output.** Tuned for the objects rather than the prose: FOL translation,
72
  entailment labels, semantic parses, Lean statements and Lean proof critique.
 
115
  `tok/s ÷ mean generation length`, so it measures completed answers rather than raw decode rate.
116
  Cells marked † need the engine note below.
117
 
118
+ | lane / metric | TwIL-LM3-Pro | Granite-4.2-3B base | VibeThinker-3B | VibeThinker-webAI-trained ★ | Qwen3.5-4B | TwIL-LM3 | Llama-3.2-3B | LFM2-2.6B | LFM2.5-8B-A1B | Qwen3-8B | gpt-oss-120b ‡ |
119
  |---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
120
  | lean_formalize token_f1 | 0.5092 | 0.2943 | 0.2087 | 0.527 | 0.4996 | 0.5869 | 0.3690 | 0.1321 | 0.4655 | 0.4022 | **0.6306** |
121
  | rule_induction derivation | 0.4195 | 0.2267 | 0.2038 | 0.227 | 0.5078 | 0.3192 | 0.0825 | 0.0615 | 0.1936 | 0.3680 | **0.6518** |
 
142
  format and tokenizer make the corpus lanes score a different quantity. The number is reported
143
  for completeness but is not a comparable measurement, and is excluded from the bolding.
144
 
145
+ ★ **VibeThinker-webAI-trained** is the public VibeThinker-3B after the same post-training pipeline as
146
  TwIL-LM3-Pro, in its SLERP (d = 0.5, t = 0.5) configuration; it is not a checkpoint in this
147
  repository, and [the section below](#against-the-tuned-vibethinker-3b) explains how it differs
148
  from the base model. Its figures are taken from the internal family comparison tables, to three
 
223
 
224
  ### Track B — held-out benchmarks
225
 
226
+ | dataset | TwIL-LM3-Pro | Granite-4.2-3B base | VibeThinker-3B | VibeThinker-webAI-trained ★ | Qwen3.5-4B | TwIL-LM3 | Llama-3.2-3B | LFM2-2.6B | LFM2.5-8B-A1B | Qwen3-8B | gpt-oss-120b ‡ |
227
  |---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
228
  | gsm8k | 0.9433 | 0.9533 | 0.9600 | 0.930 | 0.8633 | 0.8733 | 0.8300 | 0.8767 | 0.9133 | 0.9567 | **0.9767** |
229
  | svamp | 0.9500 | 0.9200 | 0.9367 | **0.953** | 0.8867 | 0.8500 | 0.8200 | 0.9000 | 0.9133 | 0.9400 | 0.9400 |
 
292
  The tables above use the public VibeThinker-3B checkpoint. The same post-training pipeline was also
293
  applied to it, and two of its tuned configurations are the closest same-scale comparisons to
294
  TwIL-LM3-Pro: WiSE-FT (λ = 0.50), and the SLERP merge (d = 0.5, t = 0.5) that appears in the
295
+ tables and plot as **VibeThinker-webAI-trained**. These values come from the family comparison tables
296
  rather than from a per-lane raw report, so they are shown as a summary only:
297
 
298
  | model | macro gate | macro_primary | B10 | B14 | Track A truncation |
299
  |---|---:|---:|---:|---:|---:|
300
  | TwIL-LM3-Pro | **0.554** | **0.588** | 0.790 | **0.743** | 24.2% |
301
+ | VibeThinker-webAI-trained (SLERP, d = 0.5, t = 0.5) | 0.508 | 0.579 | **0.802** | 0.728 | — |
302
  | VibeThinker-3B, WiSE-FT λ = 0.50 | 0.541 | **0.588** | **0.802** | 0.728 ◊ | 14.3% |
303
 
304
  ◊ There is no B14 row for the λ = 0.50 configuration; the figure is the one recorded for the SLERP
305
  configuration in the row above. Truncation was not recorded for the SLERP configuration.
306
 
307
+ Against VibeThinker-webAI-trained, TwIL-LM3-Pro is ahead on Track A (macro gate 0.554 against 0.508,
308
  `macro_primary` 0.588 against 0.579) and on the 14-dataset macro (0.743 against 0.728), and behind
309
  on the 10-dataset macro (0.790 against 0.802). Against the WiSE-FT λ = 0.50 configuration the two
310
  are effectively tied on Track A (gate 0.554 against 0.541, `macro_primary` equal at 0.588, both
 
312
  with a lower truncation rate. TwIL-LM3-Pro's edge is the 14-dataset macro, a gap that cannot be
313
  broken down per dataset from the summary values.
314
 
315
+ #### How VibeThinker-webAI-trained differs from the base VibeThinker-3B
316
 
317
  **Base VibeThinker-3B** is WeiboAI's public checkpoint, unmodified, and is what the untuned columns
318
+ in the tables above measure. **VibeThinker-webAI-trained** starts from those same weights and changes
319
  them in two ways:
320
 
321
  * **Formal-logic post-training.** A rank-64 LoRA is trained on the same synthetic formal-logic
 
328
  held-out capability in TwIL-LM3-Pro.
329
 
330
  The family tables record no reinforcement-learning (MGPO) run for VibeThinker-3B, so
331
+ VibeThinker-webAI-trained reflects the supervised and merging stages only, whereas TwIL-LM3-Pro also has
332
  the MGPO stage. It is a reference point for the pipeline, not a checkpoint shipped in this
333
  repository.
334
 
335
  What that changes, on the family tables (one source, so the comparison is like for like):
336
 
337
+ | metric | VibeThinker-3B (base) | VibeThinker-webAI-trained | change |
338
  |---|---:|---:|---:|
339
  | macro gate | 0.374 | 0.508 | +0.134 |
340
  | macro_primary | 0.444 | 0.579 | +0.135 |
 
509
  and Qwen3.5-4B, on all 18 tasks) share the same sampled rows and decoding, but the serving engine differs between
510
  arms (vLLM 0.19.1 for TwIL-LM3-Pro, its base, Qwen3.5-4B, Qwen3-8B and LFM2.5-8B-A1B; vLLM 0.11.2 for
511
  TwIL-LM3, Llama-3.2-3B and VibeThinker-3B), and the engine version is part of the protocol hash.
512
+ The VibeThinker-webAI-trained column (★) comes from the internal family comparison tables and is not
513
  covered by the manifest checks described here. Throughput has its own, separate engine caveat (see the † note under the Track A table). With
514
  n = 200 per lane on Track A and n = 300 per dataset on Track B, differences of two to three points
515
  are within sampling noise.
benchmarks.png CHANGED

Git LFS Details

  • SHA256: 4770042ab5b5b49ac64f1a1229c85eb85331f01c618655802375a15f64968375
  • Pointer size: 131 Bytes
  • Size of remote file: 135 kB

Git LFS Details

  • SHA256: 97640dc089cdcf2894f1f5c0f9277a40773d842d777e6f14dd7a59f76fb66780
  • Pointer size: 131 Bytes
  • Size of remote file: 136 kB