EldanRing commited on
Commit
580a75a
·
verified ·
1 Parent(s): f0f9931

Publish approved Winnow card, usage and evidence

Browse files
.gitattributes CHANGED
@@ -37,3 +37,4 @@ gguf/Winnow-E2B-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
37
  gguf/mmproj-Winnow-E2B.gguf filter=lfs diff=lfs merge=lfs -text
38
  gguf/Winnow-E2B-BF16.gguf filter=lfs diff=lfs merge=lfs -text
39
  gguf/Gemma-4-E2B-IT-Assistant-BF16.gguf filter=lfs diff=lfs merge=lfs -text
 
 
37
  gguf/mmproj-Winnow-E2B.gguf filter=lfs diff=lfs merge=lfs -text
38
  gguf/Winnow-E2B-BF16.gguf filter=lfs diff=lfs merge=lfs -text
39
  gguf/Gemma-4-E2B-IT-Assistant-BF16.gguf filter=lfs diff=lfs merge=lfs -text
40
+ assets/e2b-e4b-adaptive-portrait.png filter=lfs diff=lfs merge=lfs -text
README.md CHANGED
@@ -13,47 +13,51 @@ tags:
13
 
14
  # Winnow-E2B
15
 
16
- **Download the Q8 target for the reported results; a matching BF16 target is also available for higher-memory setups.** Add the projector for vision or the E2B assistant for optional MTP. Verify each downloaded file against [SHA256SUMS](SHA256SUMS). Generic Hub model snippets do not provide Winnow's `/v1/systemone` decision API.
17
 
18
- Winnow-E2B is an EldanRing decision fine-tune of [Google's Gemma 4 E2B IT](https://huggingface.co/google/gemma-4-E2B-it). It scores the candidate answers supplied for each question through Winnow's native `/v1/systemone` API. It also supports ordinary text chat; image input requires the matching vision projector. Candidate probabilities are conditional on the supplied options, not a general guarantee of correctness.
19
 
20
- The Q8 and BF16 GGUFs contain the same fine-tune merged into the language-model weights; no separate adapter or base-model download is needed. The source base revision is [`3e22461f65e89153144f8adb70e3b8c2cc9845a7`](https://huggingface.co/google/gemma-4-E2B-it/tree/3e22461f65e89153144f8adb70e3b8c2cc9845a7). The projector is the corresponding F16 vision export. Audio support has not been established for this package.
21
 
22
- ## Exact model files
23
 
24
- | Role | File | Bytes | SHA-256 |
25
- |---|---|---:|---|
26
- | Q8 target | [gguf/Winnow-E2B-Q8_0.gguf](https://huggingface.co/EldanRing/Winnow-E2B/resolve/main/gguf/Winnow-E2B-Q8_0.gguf?download=true) | 4,954,594,336 | `cb37414b46b5c4ac54ea9a44f71bb6147991c34e6be414a8cdd359cc10f0dfb1` |
27
- | BF16 target | [gguf/Winnow-E2B-BF16.gguf](https://huggingface.co/EldanRing/Winnow-E2B/resolve/main/gguf/Winnow-E2B-BF16.gguf?download=true) | 9,311,304,736 | `bcebcf44720a5409bed1960fc7caf511b9506b33b232ef46072f3f1713f61f7f` |
28
- | F16 vision projector | [gguf/mmproj-Winnow-E2B.gguf](https://huggingface.co/EldanRing/Winnow-E2B/resolve/main/gguf/mmproj-Winnow-E2B.gguf?download=true) | 985,653,600 | `7faa5282cff7250380c191bdea4e13a7b397ad851e2d883873cd9c323a5e192f` |
29
- | Optional BF16 MTP assistant | [gguf/Gemma-4-E2B-IT-Assistant-BF16.gguf](https://huggingface.co/EldanRing/Winnow-E2B/resolve/main/gguf/Gemma-4-E2B-IT-Assistant-BF16.gguf?download=true) | 170,194,016 | `0772dcc50761a47a2da87151d3105ad2b82994d9dd1dae3ad3c5c0cd26086b16` |
30
 
31
- The Q8 target, projector, and assistant hashes identify the **original files used in the E2B tests**. The BF16 target is a separate export of the same merged checkpoint; the results below do **not** evaluate or calibrate it. The shorter target and projector filenames name byte-identical copies. See the small [checksum file](SHA256SUMS) and [release manifest](release-manifest.json). Check each downloaded file; do not substitute an E4B or 12B model, projector, or assistant.
 
 
 
 
32
 
33
- The optional MTP draft is a CPU BF16 GGUF conversion of Google's [official E2B IT assistant](https://huggingface.co/google/gemma-4-E2B-it-assistant/tree/2d874ef7d29f9a30599a1e4b3c1cbc9595f005df), revision `2d874ef7d29f9a30599a1e4b3c1cbc9595f005df`. Its pinned source `model.safetensors` has SHA-256 `93682eb1c97639d18f007704dc880bd74cbe530adaf7b1bb561213863fdad2a6`. For MTP, download the bundled assistant alongside the Q8 target:
34
 
35
- ```sh
36
- huggingface-cli download EldanRing/Winnow-E2B gguf/Winnow-E2B-Q8_0.gguf gguf/Gemma-4-E2B-IT-Assistant-BF16.gguf --local-dir Winnow-E2B
37
- ```
38
 
39
- The evaluated GGUF was produced with the pinned [`llama.cpp` converter revision `911f6cdc8ab8a530b2bee09ee61471a6f3178eeb`](https://github.com/ggml-org/llama.cpp/tree/911f6cdc8ab8a530b2bee09ee61471a6f3178eeb), using BF16 output. Verify its hash before using MTP; other assistants are not interchangeable. Direct decisions and ordinary chat do not require the assistant. [Assistant attribution and conversion provenance](docs/assistants/README.md) are included.
40
 
41
- ## Decision results
42
 
43
- The table evaluates the downloadable Q8 target, with its matching projector resident, on an RTX 5070 Ti. Direct is native candidate scoring. The experimental selective path uses raw direct maximum probability `<0.99` to trigger generated reasoning, then blends direct and augmented probabilities 50:50. It used F16 target KV, Q8 draft KV, the official BF16 assistant with MTP draft length 4, backend temperature sampling, one chat slot, and an 8,192-token context.
44
 
45
- | Panel | Direct | Selective reasoning |
46
- |---|---:|---:|
47
- | JevBench public, verified labels | 175/231 (75.8%) | 203/231 (87.9%) |
48
- | Kev v9 clean, verified labels | 728/1,046 (69.6%) | 851/1,046 (81.4%) |
49
- | Typed decisions, synthetic hard-teacher agreement | 1,237/2,000 (61.9%) | 1,365/2,000 (68.3%) |
 
 
 
 
 
 
 
50
 
51
- All three full panels were previously exposed during development, so these are regression and transfer results, not independent generalization estimates. In a separate 288-case disjoint holdout, the selective path gained only 3 agreements over the earlier transferred rule; its paired 95% interval included zero. Typed agreement is not independently verified correctness.
52
 
53
- For the three full panels in table order, mean serial API time was **0.051 → 1.368 s**, **0.023 → 1.233 s**, and **0.029 → 1.981 s** from direct to selective reasoning. These selective costs are paired-output replay estimates, not deployed-service timing or concurrent throughput. Generation length and fallback costs are included. The current integration's request-to-prompt mapping differs from the benchmark serialization, so these figures do not certify exact production-path equivalence.
54
 
55
- Image evidence is limited to a narrow 180-decision synthetic procedural panel. It does not establish photo, screenshot, or document quality. A 64K-context launch and small text/image requests were checked operationally; full-length 64K quality was not measured.
56
 
57
  ## License and attribution
58
 
59
- Gemma 4 E2B IT and the separate official E2B IT assistant are by Google DeepMind under [Apache License 2.0](https://ai.google.dev/gemma/apache_2). Winnow modifications are by EldanRing. The [license](LICENSE) and [notice](NOTICE) apply to this model package; the bundled converted assistant also has [upstream attribution](docs/assistants/README.md). The Winnow inference code is separately licensed; no inference binary or source is part of this model repository. No Google endorsement is implied.
 
13
 
14
  # Winnow-E2B
15
 
16
+ Winnow-E2B is an EldanRing decision fine-tune of [Gemma 4 E2B IT](https://huggingface.co/google/gemma-4-E2B-it). It scores supplied answer options through Winnow's `/v1/systemone` API and supports ordinary text chat. Image input uses the matching vision projector.
17
 
18
+ The Q8 and BF16 GGUFs contain the same merged fine-tune. No separate adapter or base-model download is needed.
19
 
20
+ [Inference code](https://github.com/EldanRing/winnow-inference) · [Evaluation details](docs/EVALUATION.md) · [Downloads](#downloads)
21
 
22
+ ## Decision results
23
 
24
+ Direct decisions score the supplied candidates. Adaptive reasoning generates context for selected questions, then blends direct and augmented probabilities. The results below use the Q8 target on an RTX 5070 Ti: F16 target K/V cache, an 8K text context, MTP4, a raw maximum-probability gate below 0.99, and a 50:50 blend.
 
 
 
 
 
25
 
26
+ | Panel | Direct | Adaptive reasoning |
27
+ |---|---:|---:|
28
+ | JevBench public, verified labels | 175/231 (75.76%) | 202/231 (87.45%) |
29
+ | Kev v9 clean, verified labels | 728/1,046 (69.60%) | 851/1,046 (81.36%) |
30
+ | Typed, synthetic teacher agreement | 1,237/2,000 (61.85%) | 1,366/2,000 (68.30%) |
31
 
32
+ All 3,277 cases remain in the totals, including two Kev context-limit attempts that fell back to direct scoring. These frozen panels were used during development. Jev and Kev use verified labels; Typed measures agreement with synthetic teacher labels.
33
 
34
+ ![E2B and E4B direct and adaptive reasoning scores on the same frozen cases. Lighter extensions show gains; E4B Typed declines by 0.35 percentage points.](assets/e2b-e4b-adaptive-portrait.png)
 
 
35
 
36
+ The comparison uses Q8 targets, F16 target K/V cache, and MTP4. Reasoning prompts, policies, samplers, and batch sizes differ by model. E4B's seven-agreement Typed decline is shown to scale with a downward marker. See [evaluation methods](docs/EVALUATION.md#comparison-with-e4b) for exact settings.
37
 
38
+ Mean serial adaptive reasoning time was **1.356 s** on Jev, **1.212 s** on Kev, and **1.951 s** on Typed. Typed timing excludes 16 cases affected by GPU overlap; all quality results remain included. See [timing definitions](docs/EVALUATION.md#latency) for the measurement boundaries.
39
 
40
+ ## Downloads
41
 
42
+ | File | Role | Size |
43
+ |---|---|---:|
44
+ | [Winnow-E2B-Q8_0.gguf](https://huggingface.co/EldanRing/Winnow-E2B/resolve/main/gguf/Winnow-E2B-Q8_0.gguf?download=true) | Q8 target used for the reported results | 4.95 GB |
45
+ | [Winnow-E2B-BF16.gguf](https://huggingface.co/EldanRing/Winnow-E2B/resolve/main/gguf/Winnow-E2B-BF16.gguf?download=true) | Higher-precision target | 9.31 GB |
46
+ | [mmproj-Winnow-E2B.gguf](https://huggingface.co/EldanRing/Winnow-E2B/resolve/main/gguf/mmproj-Winnow-E2B.gguf?download=true) | F16 vision projector | 986 MB |
47
+ | [Gemma-4-E2B-IT-Assistant-BF16.gguf](https://huggingface.co/EldanRing/Winnow-E2B/resolve/main/gguf/Gemma-4-E2B-IT-Assistant-BF16.gguf?download=true) | Optional MTP draft assistant | 170 MB |
48
+
49
+ Text-only use needs one target; images also need the matching E2B projector. Download the exact filenames and verify them against [SHA256SUMS](SHA256SUMS).
50
+
51
+ ## Running the model
52
+
53
+ The [quickstart](https://github.com/EldanRing/winnow-inference/blob/main/docs/QUICKSTART.md) covers installation and launch commands. Use `/v1/systemone` for decisions and `/v1/chat/completions` for ordinary chat and images.
54
 
55
+ **F16 is the default and recommended target K/V cache.** Cache precision is separate from the GGUF weight format; use `--cache q8_0` for an explicit Q8 override. The pinned MTP assistant uses the target's shared K/V cache.
56
 
57
+ MTP uses the bundled BF16 conversion of Google's official E2B IT assistant. Follow the [assistant download instructions](docs/ASSISTANT.md); no conversion is needed. The assistant is a separate draft model, distinct from the BF16 target. Direct decisions and ordinary chat need only the target. Exact provenance is in the [release manifest](release-manifest.json).
58
 
59
+ Candidate probabilities are normalized over the supplied options. Confidence describes concentration among those options, not a guarantee of correctness.
60
 
61
  ## License and attribution
62
 
63
+ Gemma 4 E2B IT and the official E2B IT assistant are by Google DeepMind under [Apache License 2.0](https://ai.google.dev/gemma/apache_2). Winnow modifications are by EldanRing. See [LICENSE](LICENSE) and [NOTICE](NOTICE). The bundled assistant includes [upstream attribution](docs/assistants/README.md). Inference code is separately licensed and is not included in this model package. No Google endorsement is implied.
assets/e2b-e4b-adaptive-portrait.png ADDED

Git LFS Details

  • SHA256: c00a7e268d50e8fa1aa6b059f18431b9d7d355a61ec7bea40f17dbbab8363716
  • Pointer size: 131 Bytes
  • Size of remote file: 227 kB
assets/e2b-e4b-adaptive-portrait.svg ADDED
docs/ASSISTANT.md ADDED
@@ -0,0 +1,29 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Optional E2B MTP assistant
2
+
3
+ Direct decisions and ordinary chat need only the Winnow-E2B target. MTP additionally uses the bundled BF16 GGUF conversion of Google's official [Gemma 4 E2B IT assistant](https://huggingface.co/google/gemma-4-E2B-it-assistant/tree/2d874ef7d29f9a30599a1e4b3c1cbc9595f005df). No conversion is required to use this download.
4
+
5
+ ## Download the evaluated assistant
6
+
7
+ Download the Q8 target and its matching assistant from the model repository:
8
+
9
+ ```sh
10
+ huggingface-cli download EldanRing/Winnow-E2B \
11
+ gguf/Winnow-E2B-Q8_0.gguf \
12
+ gguf/Gemma-4-E2B-IT-Assistant-BF16.gguf \
13
+ --local-dir Winnow-E2B
14
+ ```
15
+
16
+ Verify both files against [SHA256SUMS](../SHA256SUMS). An E4B or 12B assistant is not interchangeable with the E2B assistant. Add `gguf/mmproj-Winnow-E2B.gguf` for image input.
17
+
18
+ The assistant's BF16 precision is distinct from the optional `Winnow-E2B-BF16.gguf` target. Reported MTP4 results use Q8 target weights with BF16 assistant weights and shared F16 target K/V cache.
19
+
20
+ ## Provenance
21
+
22
+ The bundled assistant is the evaluated CPU conversion from upstream revision `2d874ef7d29f9a30599a1e4b3c1cbc9595f005df`, made with `convert_hf_to_gguf.py` at pinned [llama.cpp revision](https://github.com/ggml-org/llama.cpp/tree/911f6cdc8ab8a530b2bee09ee61471a6f3178eeb) and BF16 output. Conversion changes serialization and performs no training. A different conversion does not inherit these measurements.
23
+
24
+ | Artifact | Bytes | SHA-256 |
25
+ |---|---:|---|
26
+ | Upstream `model.safetensors` | 157,565,344 | `93682eb1c97639d18f007704dc880bd74cbe530adaf7b1bb561213863fdad2a6` |
27
+ | Evaluated BF16 GGUF | 170,194,016 | `0772dcc50761a47a2da87151d3105ad2b82994d9dd1dae3ad3c5c0cd26086b16` |
28
+
29
+ The assistant is by Google DeepMind under Apache License 2.0. The [assistant attribution directory](assistants/README.md) includes the license, notice, upstream model card, and conversion provenance. Target, projector, base revision, and assistant identities are recorded in the [release manifest](../release-manifest.json).
docs/EVALUATION.md ADDED
@@ -0,0 +1,69 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Winnow-E2B evaluation
2
+
3
+ ## Scope
4
+
5
+ The completed October 6, 2026 evaluation uses the Q8 target and official BF16 MTP assistant identified in the [release manifest](../release-manifest.json), on an RTX 5070 Ti. It evaluates the adaptive reasoning path at source revision `ee6bd37d34ae35d2e69ebb0c4b0957a10c727357`.
6
+
7
+ JevBench and Kev use verified source labels. Typed measures synthetic hard-teacher agreement: 2,000 decisions from 400 five-question groups. All three panels were previously exposed during development. They are regression and transfer evidence, not independent estimates of generalization. JevBench means public-subset accuracy, not its official composite leaderboard score. No tuning on these labels occurred in this campaign.
8
+
9
+ ## E2B method
10
+
11
+ The frozen adaptive rule routes when raw direct maximum probability is below 0.99, scores at temperature 1, and combines direct and augmented probabilities 50:50. Only completed natural-EOS generations are rescored. A failed reasoning attempt retains direct probabilities and stays in the denominator.
12
+
13
+ The text-only run used the `e2b-q8-text8k-mtp` preset: an 8,192-position context, F16 target KV shared with the pinned MTP assistant, MTP draft length 4, uncapped generation at temperature 0, and seed 314159. The retained draft-cache launch flags were Q8_0, but the assistant attention cache shares the target's F16 K/V tensors; this is not a separately measured Q8 attention-cache pool. The reasoning prompt serializes options with native labels while preserving input escaping and structured-state ownership. This is an actual integrated-path run, not the earlier selective replay.
14
+
15
+ All 3,277 cases completed with zero request errors and exact direct-vector parity against frozen references (maximum difference 0; tolerance 1e-10). Routing selected 131 Jev, 676 Kev and 1,777 Typed cases. Every selected generation completed except two Kev attempts that reached the 8K context limit; those cases retained direct scoring. Partial text was not rescored or blended.
16
+
17
+ ## Comparison with E4B
18
+
19
+ The chart uses the same frozen cases, source labels and answer order. Shared settings are Q8 targets, F16 target KV shared with the pinned MTP assistant, 8K context, one chat slot and MTP4. Model-specific prompts, samplers, augmentation paths and adaptive policies differ. Winnow-12B's published campaigns remain separate and are not included in this matched comparison.
20
+
21
+ | Panel | E2B direct → adaptive reasoning | E2B change | E4B direct → adaptive reasoning | E4B change |
22
+ |---|---:|---:|---:|---:|
23
+ | JevBench, 231 | 175 → 202 | +11.69 pp | 183 → 203 | +8.66 pp |
24
+ | Kev v9 clean, 1,046 | 728 → 851 | +11.76 pp | 762 → 842 | +7.65 pp |
25
+ | Typed teacher agreement, 2,000 | 1,237 → 1,366 | +6.45 pp | 1,446 → 1,439 | −0.35 pp |
26
+
27
+ Changes are adaptive minus the same model's direct result, in percentage points. Chart percentages are computed from exact counts, with two decimal places. The card and graph both round percentages to two decimals.
28
+
29
+ E4B used the live `e4b-calibrated75-g95-v1` path: raw maxP below 0.95, direct/augmented temperatures 1.2041180007310734 / 3.4209273427377678, and 25:75 direct/augmented weighting. Its brief reasoning prompt, released sampler, 75-second deadline, and batch 2048 differ from E2B's native-label prompt, backend temperature sampler, 180-second deadline and batch 1024. Both use ubatch 1024. F16 KV was outside E4B's measured release policy profile; calibration was held fixed, not revalidated here. The comparison does not isolate model size or the effect of reasoning alone.
30
+
31
+ E4B corrected/regressed 22/2 Jev answers, 114/34 Kev answers, and 129/136 Typed agreements. Its Typed NLL, Brier, and rating MAE also worsened slightly. The graph uses a downward cap for the seven-agreement loss. These fresh E4B counts differ from its older public Q8-KV comparisons; the historical release tables remain separate.
32
+
33
+ ## Latency
34
+
35
+ | Panel | Clean cases | Mean | Median | p95 | p99 |
36
+ |---|---:|---:|---:|---:|---:|
37
+ | JevBench | 231 | 1.356 s | 0.887 s | 4.099 s | 6.057 s |
38
+ | Kev v9 clean | 1,046 | 1.212 s | 1.029 s | 3.522 s | 6.895 s |
39
+ | Typed | 1,984 | 1.951 s | 2.095 s | 3.005 s | 3.744 s |
40
+
41
+ These are serial adaptive reasoning pipeline times, including failed reasoning attempts. Sixteen Typed cases marked by the GPU-overlap evidence are excluded from timing only; all 2,000 remain in quality scores. Scored-case pipeline cost across all cases was 5,482.2 seconds. Startup, diagnostic replays, controller overhead and idle recovery gaps are outside those times. The full launch-to-cleanup span was 169.3 minutes and must not be presented as summed benchmark latency.
42
+
43
+ The earlier E4B live comparison used different output lengths, prompts and timing boundaries. It does not establish intrinsic relative speed, equal-output MTP speedup or concurrent service throughput. Native backend allocation figures are not continuously measured peak GPU VRAM.
44
+
45
+ ## Earlier paths
46
+
47
+ | Panel | Earlier numbered options | Historical F16/backend replay | Current adaptive reasoning path |
48
+ |---|---:|---:|---:|
49
+ | JevBench, 231 | 195 | 203 | 202 |
50
+ | Kev v9 clean, 1,046 | 845 | 851 | 851 |
51
+ | Typed teacher agreement, 2,000 | 1,342 | 1,365 | 1,366 |
52
+
53
+ The current path restores most historical reasoning behavior but is not identical: it loses one Jev answer, matches Kev's winners, and gains one Typed agreement relative to the historical path. Native labels, safe escaping and state augmentation can change outcomes. These older paths are historical comparisons, not current release scores or independent calibration studies.
54
+
55
+ ## Disjoint holdout
56
+
57
+ An earlier separately frozen 288-case holdout contained 96 LogiQA2 choices, 96 PAWS Boolean decisions, and 96 HelpSteer2 consensus ratings. The 0.99/50% rule scored 158/288 versus 155/288 for the transferred 0.80/50% rule: nine fixes and six regressions. Its paired 95% change interval was −1.74 to +3.82 percentage points; NLL and Brier change intervals also included zero. Mean serial replay time rose from 1.066 to 1.677 seconds.
58
+
59
+ This holdout was disjoint from selection, but was not wholly independent external validation: 74 HelpSteer2 cases had previously been scored by other models, and public sources may have appeared in base-model pretraining. The small inconclusive gain does not establish a general benefit. It is not a holdout result for the current path.
60
+
61
+ ## Vision
62
+
63
+ This campaign is text only and adds no new vision or always-mode evidence. Vision evidence is limited to an earlier 180-decision procedural-image panel, with four context-limit fallbacks. The full text comparison above does not establish photo, screenshot, document, or image-conditioned reasoning quality.
64
+
65
+ ## Evidence identity
66
+
67
+ The [machine-readable comparison](comparison-summary.json) contains exact chart counts and receipt hashes. The retained [analysis](evidence/e2b-v3-analysis.json), [freeze](evidence/e2b-v3-freeze.json), and [completion receipt](evidence/e2b-v3-completion.json) identify the evaluated source, artifacts, fallbacks and timing exclusions. Complete case logs remain in the private review evidence bundle, separate from the model package.
68
+
69
+ The final result log SHA-256 is `db3d7cb81972ad2909a75aca659b97015d9e96a770f04620275b79b9fbdbd80d`. The chart builder verifies the analysis, freeze and completion identities alongside the retained E4B receipt. Producing these figures performed no additional inference or fitting.
docs/comparison-summary.json ADDED
@@ -0,0 +1,113 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "title": "E2B / E4B direct and adaptive comparison",
3
+ "date": "2026-10-06",
4
+ "metric": "Verified-label accuracy for Jev/Kev; synthetic teacher-label agreement for Typed",
5
+ "scope": "Previously exposed panels; same frozen cases; E2B integrated v3 and retained E4B comparison; model-specific policies and generation paths",
6
+ "common_settings": {
7
+ "target_weights": "Q8",
8
+ "target_kv": "F16",
9
+ "context": 8192,
10
+ "mtp_draft_length": 4,
11
+ "chat_slots": 1,
12
+ "draft_cache_flags": "Q8_0 (retained launch flags)",
13
+ "assistant_attention_kv": "F16 (shared target K/V tensors)"
14
+ },
15
+ "e2b": {
16
+ "gate": "raw maxP < 0.99",
17
+ "direct_augmented_weights": [
18
+ 0.5,
19
+ 0.5
20
+ ],
21
+ "sampler": "backend temperature sampling",
22
+ "batch": 1024,
23
+ "ubatch": 1024,
24
+ "deadline_seconds": 180,
25
+ "adaptive_execution": "actual integrated v3 run",
26
+ "source_commit": "ee6bd37d34ae35d2e69ebb0c4b0957a10c727357",
27
+ "policy": "e2b-raw99-blend50-v3",
28
+ "prompt": "native labels with safe input escaping and structured-state ownership"
29
+ },
30
+ "e4b": {
31
+ "policy": "e4b-calibrated75-g95-v1",
32
+ "gate": "raw maxP < 0.95",
33
+ "direct_augmented_weights": [
34
+ 0.25,
35
+ 0.75
36
+ ],
37
+ "sampler": "released sampler; no E2B backend-temperature flag",
38
+ "batch": 2048,
39
+ "ubatch": 1024,
40
+ "deadline_seconds": 75,
41
+ "adaptive_execution": "live selective client",
42
+ "calibration_scope": "F16 KV outside measured release profile; temperatures held fixed"
43
+ },
44
+ "receipt_sha256": {
45
+ "e4b-e2b-comparison-analysis.json": "b19713c86fc58ab837137009bb8796a6fb82c717465f9c84cfc63cafe5f8d19f",
46
+ "e4b-e2b-comparison-freeze.json": "6ffc48de7155b73d69659ee7757b122354813b0984b02712a84b89854c40eacd",
47
+ "e2b-v3-analysis.json": "1cd872865228fb33adaf12dd48d3656b98fd6f1f499d8adf364e04de4f1ddb27",
48
+ "e2b-v3-freeze.json": "eb163f63b0eb3a706adaaf18d20ebe3acaef4ddb0690005149eef2cbc750d2a6",
49
+ "e2b-v3-completion.json": "6e8eb1e7f38bb725321b5bc24683ec8ccfb24b34a92d5de15e94f776d30e26b0"
50
+ },
51
+ "panels": [
52
+ {
53
+ "key": "jevbench-public",
54
+ "title": "JevBench",
55
+ "n": 231,
56
+ "source_sha256": "bd304bc512d70bcbd9ff0b7d039d241b86aa0072ae5533c3c8d493c170aeb4d0",
57
+ "E2B": {
58
+ "direct": 175,
59
+ "adaptive": 202,
60
+ "direct_pct": "75.76",
61
+ "adaptive_pct": "87.45",
62
+ "change_pp": "+11.69"
63
+ },
64
+ "E4B": {
65
+ "direct": 183,
66
+ "adaptive": 203,
67
+ "direct_pct": "79.22",
68
+ "adaptive_pct": "87.88",
69
+ "change_pp": "+8.66"
70
+ }
71
+ },
72
+ {
73
+ "key": "kev-v9-clean",
74
+ "title": "Kev v9 clean",
75
+ "n": 1046,
76
+ "source_sha256": "4ca4b28a171d6130f00e078316e9cdfcdedd0933a1b5fb967517c2e20244ad90",
77
+ "E2B": {
78
+ "direct": 728,
79
+ "adaptive": 851,
80
+ "direct_pct": "69.60",
81
+ "adaptive_pct": "81.36",
82
+ "change_pp": "+11.76"
83
+ },
84
+ "E4B": {
85
+ "direct": 762,
86
+ "adaptive": 842,
87
+ "direct_pct": "72.85",
88
+ "adaptive_pct": "80.50",
89
+ "change_pp": "+7.65"
90
+ }
91
+ },
92
+ {
93
+ "key": "typed-decisions",
94
+ "title": "Typed",
95
+ "n": 2000,
96
+ "source_sha256": "143541319ddc4445ba67098c0af0480646e8124d682365d649b8e5947298cc0d",
97
+ "E2B": {
98
+ "direct": 1237,
99
+ "adaptive": 1366,
100
+ "direct_pct": "61.85",
101
+ "adaptive_pct": "68.30",
102
+ "change_pp": "+6.45"
103
+ },
104
+ "E4B": {
105
+ "direct": 1446,
106
+ "adaptive": 1439,
107
+ "direct_pct": "72.30",
108
+ "adaptive_pct": "71.95",
109
+ "change_pp": "\u22120.35"
110
+ }
111
+ }
112
+ ]
113
+ }
docs/evidence/e2b-v3-analysis.json ADDED
@@ -0,0 +1,1534 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "results_sha256": "db3d7cb81972ad2909a75aca659b97015d9e96a770f04620275b79b9fbdbd80d",
3
+ "source_commit": "ee6bd37d34ae35d2e69ebb0c4b0957a10c727357",
4
+ "freeze_sha256": "eb163f63b0eb3a706adaaf18d20ebe3acaef4ddb0690005149eef2cbc750d2a6",
5
+ "observed": 3277,
6
+ "full_panel_complete": true,
7
+ "timing_exclusions": {
8
+ "latency_exclusion_only": true,
9
+ "quality_rows_retained": true,
10
+ "prior_uncertain_ranges": {
11
+ "1491": [
12
+ 1496,
13
+ "first foreign-GPU watchdog stop; per-case timestamps unavailable"
14
+ ],
15
+ "2651": [
16
+ 2660,
17
+ "RustDesk watchdog stop; per-case timestamps unavailable"
18
+ ]
19
+ },
20
+ "new_overlap_samples": 0,
21
+ "excluded_cases": [
22
+ {
23
+ "index": 1491,
24
+ "panel": "typed-decisions",
25
+ "id": "agent_trace_observability_000042/risk",
26
+ "reasons": [
27
+ "first foreign-GPU watchdog stop; per-case timestamps unavailable"
28
+ ]
29
+ },
30
+ {
31
+ "index": 1492,
32
+ "panel": "typed-decisions",
33
+ "id": "agent_trace_observability_000042/urgency",
34
+ "reasons": [
35
+ "first foreign-GPU watchdog stop; per-case timestamps unavailable"
36
+ ]
37
+ },
38
+ {
39
+ "index": 1493,
40
+ "panel": "typed-decisions",
41
+ "id": "agent_trace_observability_000043/action",
42
+ "reasons": [
43
+ "first foreign-GPU watchdog stop; per-case timestamps unavailable"
44
+ ]
45
+ },
46
+ {
47
+ "index": 1494,
48
+ "panel": "typed-decisions",
49
+ "id": "agent_trace_observability_000043/needs_review",
50
+ "reasons": [
51
+ "first foreign-GPU watchdog stop; per-case timestamps unavailable"
52
+ ]
53
+ },
54
+ {
55
+ "index": 1495,
56
+ "panel": "typed-decisions",
57
+ "id": "agent_trace_observability_000043/outcome",
58
+ "reasons": [
59
+ "first foreign-GPU watchdog stop; per-case timestamps unavailable"
60
+ ]
61
+ },
62
+ {
63
+ "index": 1496,
64
+ "panel": "typed-decisions",
65
+ "id": "agent_trace_observability_000043/risk",
66
+ "reasons": [
67
+ "first foreign-GPU watchdog stop; per-case timestamps unavailable"
68
+ ]
69
+ },
70
+ {
71
+ "index": 2651,
72
+ "panel": "typed-decisions",
73
+ "id": "invoice_processing_000074/matches_order",
74
+ "reasons": [
75
+ "RustDesk watchdog stop; per-case timestamps unavailable"
76
+ ]
77
+ },
78
+ {
79
+ "index": 2652,
80
+ "panel": "typed-decisions",
81
+ "id": "invoice_processing_000074/urgency",
82
+ "reasons": [
83
+ "RustDesk watchdog stop; per-case timestamps unavailable"
84
+ ]
85
+ },
86
+ {
87
+ "index": 2653,
88
+ "panel": "typed-decisions",
89
+ "id": "invoice_processing_000075/discrepancy_severity",
90
+ "reasons": [
91
+ "RustDesk watchdog stop; per-case timestamps unavailable"
92
+ ]
93
+ },
94
+ {
95
+ "index": 2654,
96
+ "panel": "typed-decisions",
97
+ "id": "invoice_processing_000075/disposition",
98
+ "reasons": [
99
+ "RustDesk watchdog stop; per-case timestamps unavailable"
100
+ ]
101
+ },
102
+ {
103
+ "index": 2655,
104
+ "panel": "typed-decisions",
105
+ "id": "invoice_processing_000075/duplicate",
106
+ "reasons": [
107
+ "RustDesk watchdog stop; per-case timestamps unavailable"
108
+ ]
109
+ },
110
+ {
111
+ "index": 2656,
112
+ "panel": "typed-decisions",
113
+ "id": "invoice_processing_000075/matches_order",
114
+ "reasons": [
115
+ "RustDesk watchdog stop; per-case timestamps unavailable"
116
+ ]
117
+ },
118
+ {
119
+ "index": 2657,
120
+ "panel": "typed-decisions",
121
+ "id": "invoice_processing_000075/urgency",
122
+ "reasons": [
123
+ "RustDesk watchdog stop; per-case timestamps unavailable"
124
+ ]
125
+ },
126
+ {
127
+ "index": 2658,
128
+ "panel": "typed-decisions",
129
+ "id": "invoice_processing_000076/discrepancy_severity",
130
+ "reasons": [
131
+ "RustDesk watchdog stop; per-case timestamps unavailable"
132
+ ]
133
+ },
134
+ {
135
+ "index": 2659,
136
+ "panel": "typed-decisions",
137
+ "id": "invoice_processing_000076/disposition",
138
+ "reasons": [
139
+ "RustDesk watchdog stop; per-case timestamps unavailable"
140
+ ]
141
+ },
142
+ {
143
+ "index": 2660,
144
+ "panel": "typed-decisions",
145
+ "id": "invoice_processing_000076/duplicate",
146
+ "reasons": [
147
+ "RustDesk watchdog stop; per-case timestamps unavailable"
148
+ ]
149
+ }
150
+ ]
151
+ },
152
+ "panels": {
153
+ "jevbench-public": {
154
+ "observed": 231,
155
+ "expected": 231,
156
+ "complete": true,
157
+ "valid": 231,
158
+ "semantics": "verified label accuracy",
159
+ "errors": [],
160
+ "fallbacks": [],
161
+ "direct": {
162
+ "n": 231,
163
+ "top1": 175,
164
+ "nll": 0.6548821957788397,
165
+ "brier": 0.3384088069425174,
166
+ "score_expected_level_mae": 0.18163427398620083
167
+ },
168
+ "v2": {
169
+ "n": 231,
170
+ "top1": 195,
171
+ "nll": 0.5102332515038959,
172
+ "brier": 0.24715639752392515,
173
+ "score_expected_level_mae": 0.1169276614338399
174
+ },
175
+ "v3": {
176
+ "n": 231,
177
+ "top1": 202,
178
+ "nll": 0.4439350676441406,
179
+ "brier": 0.21173286180768633,
180
+ "score_expected_level_mae": 0.08733389277501005
181
+ },
182
+ "historical": {
183
+ "n": 231,
184
+ "top1": 203,
185
+ "nll": 0.4433104727261114,
186
+ "brier": 0.211353951167008,
187
+ "score_expected_level_mae": 0.08733389277501005
188
+ },
189
+ "direct_to_v3": {
190
+ "n": 231,
191
+ "fixes": 35,
192
+ "regressions": 8,
193
+ "winner_changes": 54,
194
+ "max_probability_delta": 0.4938297051869622
195
+ },
196
+ "v2_to_v3": {
197
+ "n": 231,
198
+ "fixes": 15,
199
+ "regressions": 8,
200
+ "winner_changes": 26,
201
+ "max_probability_delta": 0.49954774239515043
202
+ },
203
+ "historical_to_v3": {
204
+ "n": 231,
205
+ "fixes": 0,
206
+ "regressions": 1,
207
+ "winner_changes": 1,
208
+ "max_probability_delta": 0.07599484280024105
209
+ },
210
+ "routed": 131,
211
+ "completed_blends": 131,
212
+ "max_direct_delta": 0.0,
213
+ "historical_text_equal": 126,
214
+ "historical_text_compared": 131,
215
+ "generation_output_tokens": {
216
+ "sum": 123713,
217
+ "mean": 944.3740458015267,
218
+ "max": 2952,
219
+ "distribution": {
220
+ "n": 131,
221
+ "sum": 123713,
222
+ "mean": 944.3740458015267,
223
+ "median": 856,
224
+ "p95": 1936.5,
225
+ "p99": 2826.3999999999987,
226
+ "max": 2952
227
+ }
228
+ },
229
+ "pipeline_seconds_all_cost": {
230
+ "n": 231,
231
+ "sum": 313.31913769757375,
232
+ "mean": 1.3563599034527003,
233
+ "median": 0.8873364869505167,
234
+ "p95": 4.099002962524537,
235
+ "p99": 6.056786250858566,
236
+ "max": 7.182334668934345
237
+ },
238
+ "pipeline_seconds_clean": {
239
+ "n": 231,
240
+ "sum": 313.31913769757375,
241
+ "mean": 1.3563599034527003,
242
+ "median": 0.8873364869505167,
243
+ "p95": 4.099002962524537,
244
+ "p99": 6.056786250858566,
245
+ "max": 7.182334668934345
246
+ },
247
+ "timing_excluded_count": 0,
248
+ "v2_selective_seconds_clean_matched": {
249
+ "n": 231,
250
+ "sum": 314.0389325831784,
251
+ "mean": 1.3594758986284778,
252
+ "median": 0.8249348889803514,
253
+ "p95": 4.52698246500222,
254
+ "p99": 5.536772494891189,
255
+ "max": 5.958748230012134
256
+ },
257
+ "max_native_backend_allocated_bytes": 6601572352,
258
+ "natural_eos_output_tokens": {
259
+ "n": 131,
260
+ "sum": 123713,
261
+ "mean": 944.3740458015267,
262
+ "median": 856,
263
+ "p95": 1936.5,
264
+ "p99": 2826.3999999999987,
265
+ "max": 2952
266
+ },
267
+ "clean_phase_seconds": {
268
+ "/v1/chat/completions": {
269
+ "n": 231,
270
+ "sum": 283.491011307342,
271
+ "mean": 1.2272338151832987,
272
+ "median": 0.8251801070291549,
273
+ "p95": 3.6716601609950885,
274
+ "p99": 5.589920414669901,
275
+ "max": 6.879954903037287
276
+ },
277
+ "/v1/systemone": {
278
+ "n": 231,
279
+ "sum": 28.06992222147528,
280
+ "mean": 0.1215148148115813,
281
+ "median": 0.05774561199359596,
282
+ "p95": 0.4264363810652867,
283
+ "p99": 0.5270983364898711,
284
+ "max": 0.5750212969724089
285
+ },
286
+ "/v1/winnow/inspect": {
287
+ "n": 231,
288
+ "sum": 1.6974414952564985,
289
+ "mean": 0.007348231581196963,
290
+ "median": 0.004081002902239561,
291
+ "p95": 0.02844280900899321,
292
+ "p99": 0.0330876518972218,
293
+ "max": 0.039218239951878786
294
+ }
295
+ },
296
+ "clean_chat_metrics": {
297
+ "n": 131,
298
+ "output_tokens": 123713,
299
+ "decode_seconds": 272.717043,
300
+ "aggregate_decode_tokens_per_second": 453.63134859158765,
301
+ "draft_proposed": 152252,
302
+ "draft_accepted": 85751,
303
+ "draft_acceptance": 0.563217560360455,
304
+ "prefill_ms": {
305
+ "n": 131,
306
+ "sum": 10241.078,
307
+ "mean": 78.1761679389313,
308
+ "median": 46.572,
309
+ "p95": 210.508,
310
+ "p99": 227.93679999999998,
311
+ "max": 248.51
312
+ },
313
+ "decode_ms": {
314
+ "n": 131,
315
+ "sum": 272717.043,
316
+ "mean": 2081.8094885496184,
317
+ "median": 1831.607,
318
+ "p95": 4055.5175,
319
+ "p99": 5768.822199999997,
320
+ "max": 6808.545
321
+ }
322
+ },
323
+ "clean_chat_metrics_natural_eos": {
324
+ "n": 131,
325
+ "output_tokens": 123713,
326
+ "decode_seconds": 272.717043,
327
+ "aggregate_decode_tokens_per_second": 453.63134859158765,
328
+ "draft_proposed": 152252,
329
+ "draft_accepted": 85751,
330
+ "draft_acceptance": 0.563217560360455,
331
+ "prefill_ms": {
332
+ "n": 131,
333
+ "sum": 10241.078,
334
+ "mean": 78.1761679389313,
335
+ "median": 46.572,
336
+ "p95": 210.508,
337
+ "p99": 227.93679999999998,
338
+ "max": 248.51
339
+ },
340
+ "decode_ms": {
341
+ "n": 131,
342
+ "sum": 272717.043,
343
+ "mean": 2081.8094885496184,
344
+ "median": 1831.607,
345
+ "p95": 4055.5175,
346
+ "p99": 5768.822199999997,
347
+ "max": 6808.545
348
+ }
349
+ },
350
+ "by_state": {
351
+ "dict": {
352
+ "direct": {
353
+ "n": 35,
354
+ "top1": 24,
355
+ "nll": 0.8098638610329222,
356
+ "brier": 0.4492148897620551,
357
+ "score_expected_level_mae": null
358
+ },
359
+ "v2": {
360
+ "n": 35,
361
+ "top1": 30,
362
+ "nll": 0.5730443227047298,
363
+ "brier": 0.2778388271622036,
364
+ "score_expected_level_mae": null
365
+ },
366
+ "v3": {
367
+ "n": 35,
368
+ "top1": 28,
369
+ "nll": 0.5986073814983259,
370
+ "brier": 0.2874113653226452,
371
+ "score_expected_level_mae": null
372
+ },
373
+ "historical": {
374
+ "n": 35,
375
+ "top1": 29,
376
+ "nll": 0.5952986408668535,
377
+ "brier": 0.2851932314045603,
378
+ "score_expected_level_mae": null
379
+ },
380
+ "direct_to_v3": {
381
+ "n": 35,
382
+ "fixes": 7,
383
+ "regressions": 3,
384
+ "winner_changes": 12,
385
+ "max_probability_delta": 0.4847284009193595
386
+ },
387
+ "v2_to_v3": {
388
+ "n": 35,
389
+ "fixes": 1,
390
+ "regressions": 3,
391
+ "winner_changes": 5,
392
+ "max_probability_delta": 0.4927516672740554
393
+ },
394
+ "historical_to_v3": {
395
+ "n": 35,
396
+ "fixes": 0,
397
+ "regressions": 1,
398
+ "winner_changes": 1,
399
+ "max_probability_delta": 0.07599484280024105
400
+ }
401
+ },
402
+ "str": {
403
+ "direct": {
404
+ "n": 196,
405
+ "top1": 151,
406
+ "nll": 0.6272068984120392,
407
+ "brier": 0.31862200643902855,
408
+ "score_expected_level_mae": 0.18163427398620083
409
+ },
410
+ "v2": {
411
+ "n": 196,
412
+ "top1": 165,
413
+ "nll": 0.49901698878946127,
414
+ "brier": 0.24167739223137544,
415
+ "score_expected_level_mae": 0.1169276614338399
416
+ },
417
+ "v3": {
418
+ "n": 196,
419
+ "top1": 174,
420
+ "nll": 0.41631501159875034,
421
+ "brier": 0.19821884332287223,
422
+ "score_expected_level_mae": 0.08733389277501005
423
+ },
424
+ "historical": {
425
+ "n": 196,
426
+ "top1": 174,
427
+ "nll": 0.4161697284152646,
428
+ "brier": 0.19816836541030225,
429
+ "score_expected_level_mae": 0.08733389277501005
430
+ },
431
+ "direct_to_v3": {
432
+ "n": 196,
433
+ "fixes": 28,
434
+ "regressions": 5,
435
+ "winner_changes": 42,
436
+ "max_probability_delta": 0.4938297051869622
437
+ },
438
+ "v2_to_v3": {
439
+ "n": 196,
440
+ "fixes": 14,
441
+ "regressions": 5,
442
+ "winner_changes": 21,
443
+ "max_probability_delta": 0.49954774239515043
444
+ },
445
+ "historical_to_v3": {
446
+ "n": 196,
447
+ "fixes": 0,
448
+ "regressions": 0,
449
+ "winner_changes": 0,
450
+ "max_probability_delta": 0.010968767130365098
451
+ }
452
+ }
453
+ },
454
+ "by_kind": {
455
+ "choice": {
456
+ "direct": {
457
+ "n": 139,
458
+ "top1": 100,
459
+ "nll": 0.8075315805861729,
460
+ "brier": 0.39522492072947685,
461
+ "score_expected_level_mae": null
462
+ },
463
+ "v2": {
464
+ "n": 139,
465
+ "top1": 115,
466
+ "nll": 0.6044044303940057,
467
+ "brier": 0.2790112001195667,
468
+ "score_expected_level_mae": null
469
+ },
470
+ "v3": {
471
+ "n": 139,
472
+ "top1": 119,
473
+ "nll": 0.5397349946166345,
474
+ "brier": 0.24901431746883707,
475
+ "score_expected_level_mae": null
476
+ },
477
+ "historical": {
478
+ "n": 139,
479
+ "top1": 119,
480
+ "nll": 0.5396015160770918,
481
+ "brier": 0.24895584375003005,
482
+ "score_expected_level_mae": null
483
+ },
484
+ "direct_to_v3": {
485
+ "n": 139,
486
+ "fixes": 24,
487
+ "regressions": 5,
488
+ "winner_changes": 39,
489
+ "max_probability_delta": 0.4938297051869622
490
+ },
491
+ "v2_to_v3": {
492
+ "n": 139,
493
+ "fixes": 9,
494
+ "regressions": 5,
495
+ "winner_changes": 17,
496
+ "max_probability_delta": 0.49954774239515043
497
+ },
498
+ "historical_to_v3": {
499
+ "n": 139,
500
+ "fixes": 0,
501
+ "regressions": 0,
502
+ "winner_changes": 0,
503
+ "max_probability_delta": 0.010968767130365098
504
+ }
505
+ },
506
+ "noul": {
507
+ "direct": {
508
+ "n": 74,
509
+ "top1": 61,
510
+ "nll": 0.43519757844670776,
511
+ "brier": 0.26491784130498336,
512
+ "score_expected_level_mae": null
513
+ },
514
+ "v2": {
515
+ "n": 74,
516
+ "top1": 64,
517
+ "nll": 0.3773254061123309,
518
+ "brier": 0.207808720831387,
519
+ "score_expected_level_mae": null
520
+ },
521
+ "v3": {
522
+ "n": 74,
523
+ "top1": 66,
524
+ "nll": 0.3184973321031578,
525
+ "brier": 0.16806665325527254,
526
+ "score_expected_level_mae": null
527
+ },
528
+ "historical": {
529
+ "n": 74,
530
+ "top1": 67,
531
+ "nll": 0.31679830630493755,
532
+ "brier": 0.1669936733757791,
533
+ "score_expected_level_mae": null
534
+ },
535
+ "direct_to_v3": {
536
+ "n": 74,
537
+ "fixes": 8,
538
+ "regressions": 3,
539
+ "winner_changes": 11,
540
+ "max_probability_delta": 0.4496872566776117
541
+ },
542
+ "v2_to_v3": {
543
+ "n": 74,
544
+ "fixes": 5,
545
+ "regressions": 3,
546
+ "winner_changes": 8,
547
+ "max_probability_delta": 0.4990190336545404
548
+ },
549
+ "historical_to_v3": {
550
+ "n": 74,
551
+ "fixes": 0,
552
+ "regressions": 1,
553
+ "winner_changes": 1,
554
+ "max_probability_delta": 0.07599484280024105
555
+ }
556
+ },
557
+ "score": {
558
+ "direct": {
559
+ "n": 18,
560
+ "top1": 14,
561
+ "nll": 0.3792375954654199,
562
+ "brier": 0.20179167587530403,
563
+ "score_expected_level_mae": 0.18163427398620083
564
+ },
565
+ "v2": {
566
+ "n": 18,
567
+ "top1": 16,
568
+ "nll": 0.32942140112892654,
569
+ "brier": 0.16292920388246132,
570
+ "score_expected_level_mae": 0.1169276614338399
571
+ },
572
+ "v3": {
573
+ "n": 18,
574
+ "top1": 17,
575
+ "nll": 0.21983521102503334,
576
+ "brier": 0.10335381158427896,
577
+ "score_expected_level_mae": 0.08733389277501005
578
+ },
579
+ "historical": {
580
+ "n": 18,
581
+ "top1": 17,
582
+ "nll": 0.21983521102503334,
583
+ "brier": 0.10335381158427896,
584
+ "score_expected_level_mae": 0.08733389277501005
585
+ },
586
+ "direct_to_v3": {
587
+ "n": 18,
588
+ "fixes": 3,
589
+ "regressions": 0,
590
+ "winner_changes": 4,
591
+ "max_probability_delta": 0.4173095607519327
592
+ },
593
+ "v2_to_v3": {
594
+ "n": 18,
595
+ "fixes": 1,
596
+ "regressions": 0,
597
+ "winner_changes": 1,
598
+ "max_probability_delta": 0.49626072038671637
599
+ },
600
+ "historical_to_v3": {
601
+ "n": 18,
602
+ "fixes": 0,
603
+ "regressions": 0,
604
+ "winner_changes": 0,
605
+ "max_probability_delta": 0.0
606
+ }
607
+ }
608
+ }
609
+ },
610
+ "kev-v9-clean": {
611
+ "observed": 1046,
612
+ "expected": 1046,
613
+ "complete": true,
614
+ "valid": 1046,
615
+ "semantics": "verified label accuracy",
616
+ "errors": [],
617
+ "fallbacks": [
618
+ {
619
+ "id": "mmlu_pro/test/9247/answer",
620
+ "reason": "generation_context_limit_or_unverified"
621
+ },
622
+ {
623
+ "id": "unknowable_control/warranty_claim/v5-test-20260920-warranty_claim-0004/a/decision",
624
+ "reason": "generation_context_limit_or_unverified"
625
+ }
626
+ ],
627
+ "direct": {
628
+ "n": 1046,
629
+ "top1": 728,
630
+ "nll": 0.9841937846927966,
631
+ "brier": 0.44164586585817817,
632
+ "score_expected_level_mae": 0.29737852879492405
633
+ },
634
+ "v2": {
635
+ "n": 1046,
636
+ "top1": 845,
637
+ "nll": 0.7616620243944804,
638
+ "brier": 0.31723865346147595,
639
+ "score_expected_level_mae": 0.16328748801957052
640
+ },
641
+ "v3": {
642
+ "n": 1046,
643
+ "top1": 851,
644
+ "nll": 0.743563618944474,
645
+ "brier": 0.3110523200923067,
646
+ "score_expected_level_mae": 0.16352370646939757
647
+ },
648
+ "historical": {
649
+ "n": 1046,
650
+ "top1": 851,
651
+ "nll": 0.7439440288222531,
652
+ "brier": 0.3113648971845476,
653
+ "score_expected_level_mae": 0.1636530071709598
654
+ },
655
+ "direct_to_v3": {
656
+ "n": 1046,
657
+ "fixes": 173,
658
+ "regressions": 50,
659
+ "winner_changes": 264,
660
+ "max_probability_delta": 0.49877661297158915
661
+ },
662
+ "v2_to_v3": {
663
+ "n": 1046,
664
+ "fixes": 38,
665
+ "regressions": 32,
666
+ "winner_changes": 100,
667
+ "max_probability_delta": 0.4998410532110379
668
+ },
669
+ "historical_to_v3": {
670
+ "n": 1046,
671
+ "fixes": 0,
672
+ "regressions": 0,
673
+ "winner_changes": 2,
674
+ "max_probability_delta": 0.49223936331112117
675
+ },
676
+ "routed": 676,
677
+ "completed_blends": 674,
678
+ "max_direct_delta": 0.0,
679
+ "historical_text_equal": 674,
680
+ "historical_text_compared": 676,
681
+ "generation_output_tokens": {
682
+ "sum": 508239,
683
+ "mean": 751.8328402366864,
684
+ "max": 8012,
685
+ "distribution": {
686
+ "n": 676,
687
+ "sum": 508239,
688
+ "mean": 751.8328402366864,
689
+ "median": 565.5,
690
+ "p95": 1889.75,
691
+ "p99": 3618.25,
692
+ "max": 8012
693
+ }
694
+ },
695
+ "pipeline_seconds_all_cost": {
696
+ "n": 1046,
697
+ "sum": 1267.75761802739,
698
+ "mean": 1.212005370963088,
699
+ "median": 1.028721651993692,
700
+ "p95": 3.5224499475152697,
701
+ "p99": 6.895163247355956,
702
+ "max": 12.624659975990653
703
+ },
704
+ "pipeline_seconds_clean": {
705
+ "n": 1046,
706
+ "sum": 1267.75761802739,
707
+ "mean": 1.212005370963088,
708
+ "median": 1.028721651993692,
709
+ "p95": 3.5224499475152697,
710
+ "p99": 6.895163247355956,
711
+ "max": 12.624659975990653
712
+ },
713
+ "timing_excluded_count": 0,
714
+ "v2_selective_seconds_clean_matched": {
715
+ "n": 1046,
716
+ "sum": 1300.9493182514561,
717
+ "mean": 1.2437373979459427,
718
+ "median": 0.9990504610468633,
719
+ "p95": 3.6775082352978643,
720
+ "p99": 6.746468279679535,
721
+ "max": 9.355459655052982
722
+ },
723
+ "max_native_backend_allocated_bytes": 6603538432,
724
+ "natural_eos_output_tokens": {
725
+ "n": 674,
726
+ "sum": 492665,
727
+ "mean": 730.9569732937686,
728
+ "median": 563.5,
729
+ "p95": 1851.4500000000003,
730
+ "p99": 3412.4799999999996,
731
+ "max": 5334
732
+ },
733
+ "clean_phase_seconds": {
734
+ "/v1/chat/completions": {
735
+ "n": 1046,
736
+ "sum": 1196.6142937389668,
737
+ "mean": 1.1439907205917466,
738
+ "median": 0.9535917384782806,
739
+ "p95": 3.3786256242892705,
740
+ "p99": 6.638901050179261,
741
+ "max": 12.581271308939904
742
+ },
743
+ "/v1/systemone": {
744
+ "n": 1046,
745
+ "sum": 66.71202236169484,
746
+ "mean": 0.06377822405515758,
747
+ "median": 0.06213975552236661,
748
+ "p95": 0.13457796376314946,
749
+ "p99": 0.23082962960470413,
750
+ "max": 0.39296685194130987
751
+ },
752
+ "/v1/winnow/inspect": {
753
+ "n": 1046,
754
+ "sum": 4.167626591864973,
755
+ "mean": 0.003984346646142422,
756
+ "median": 0.004179205046966672,
757
+ "p95": 0.008609114302089438,
758
+ "p99": 0.010857289913110432,
759
+ "max": 0.01358645805157721
760
+ }
761
+ },
762
+ "clean_chat_metrics": {
763
+ "n": 676,
764
+ "output_tokens": 508239,
765
+ "decode_seconds": 1172.2648219999999,
766
+ "aggregate_decode_tokens_per_second": 433.55305939565255,
767
+ "draft_proposed": 666445,
768
+ "draft_accepted": 342220,
769
+ "draft_acceptance": 0.513500738995716,
770
+ "prefill_ms": {
771
+ "n": 676,
772
+ "sum": 22825.388,
773
+ "mean": 33.76536686390533,
774
+ "median": 32.602500000000006,
775
+ "p95": 41.06625,
776
+ "p99": 48.360749999999996,
777
+ "max": 80.819
778
+ },
779
+ "decode_ms": {
780
+ "n": 676,
781
+ "sum": 1172264.822,
782
+ "mean": 1734.1195591715975,
783
+ "median": 1317.3195,
784
+ "p95": 4071.9035000000003,
785
+ "p99": 8340.009,
786
+ "max": 12524.258
787
+ }
788
+ },
789
+ "clean_chat_metrics_natural_eos": {
790
+ "n": 674,
791
+ "output_tokens": 492665,
792
+ "decode_seconds": 1147.630398,
793
+ "aggregate_decode_tokens_per_second": 429.28890769935845,
794
+ "draft_proposed": 653084,
795
+ "draft_accepted": 329989,
796
+ "draft_acceptance": 0.5052780346785406,
797
+ "prefill_ms": {
798
+ "n": 674,
799
+ "sum": 22739.867,
800
+ "mean": 33.73867507418397,
801
+ "median": 32.602500000000006,
802
+ "p95": 41.03915,
803
+ "p99": 46.812639999999995,
804
+ "max": 80.819
805
+ },
806
+ "decode_ms": {
807
+ "n": 674,
808
+ "sum": 1147630.398,
809
+ "mean": 1702.7157240356082,
810
+ "median": 1315.1734999999999,
811
+ "p95": 4007.4803000000006,
812
+ "p99": 7377.133189999987,
813
+ "max": 11452.211
814
+ }
815
+ },
816
+ "by_state": {
817
+ "dict": {
818
+ "direct": {
819
+ "n": 757,
820
+ "top1": 506,
821
+ "nll": 1.0263358490810022,
822
+ "brier": 0.4660603526231764,
823
+ "score_expected_level_mae": 0.29737852879492405
824
+ },
825
+ "v2": {
826
+ "n": 757,
827
+ "top1": 621,
828
+ "nll": 0.7171781019173493,
829
+ "brier": 0.3007661792699979,
830
+ "score_expected_level_mae": 0.16328748801957052
831
+ },
832
+ "v3": {
833
+ "n": 757,
834
+ "top1": 619,
835
+ "nll": 0.7081713179511532,
836
+ "brier": 0.30210716310944896,
837
+ "score_expected_level_mae": 0.16352370646939757
838
+ },
839
+ "historical": {
840
+ "n": 757,
841
+ "top1": 619,
842
+ "nll": 0.7087038104127751,
843
+ "brier": 0.3025430916237622,
844
+ "score_expected_level_mae": 0.1636530071709598
845
+ },
846
+ "direct_to_v3": {
847
+ "n": 757,
848
+ "fixes": 150,
849
+ "regressions": 37,
850
+ "winner_changes": 225,
851
+ "max_probability_delta": 0.49877661297158915
852
+ },
853
+ "v2_to_v3": {
854
+ "n": 757,
855
+ "fixes": 24,
856
+ "regressions": 26,
857
+ "winner_changes": 78,
858
+ "max_probability_delta": 0.4998410532110379
859
+ },
860
+ "historical_to_v3": {
861
+ "n": 757,
862
+ "fixes": 0,
863
+ "regressions": 0,
864
+ "winner_changes": 2,
865
+ "max_probability_delta": 0.49223936331112117
866
+ }
867
+ },
868
+ "list": {
869
+ "direct": {
870
+ "n": 11,
871
+ "top1": 9,
872
+ "nll": 0.45089530046520027,
873
+ "brier": 0.2974764690463965,
874
+ "score_expected_level_mae": null
875
+ },
876
+ "v2": {
877
+ "n": 11,
878
+ "top1": 10,
879
+ "nll": 0.1983185854747978,
880
+ "brier": 0.12111648754883751,
881
+ "score_expected_level_mae": null
882
+ },
883
+ "v3": {
884
+ "n": 11,
885
+ "top1": 11,
886
+ "nll": 0.1368364542171407,
887
+ "brier": 0.07533518108769578,
888
+ "score_expected_level_mae": null
889
+ },
890
+ "historical": {
891
+ "n": 11,
892
+ "top1": 11,
893
+ "nll": 0.13636481228161182,
894
+ "brier": 0.07505861319212954,
895
+ "score_expected_level_mae": null
896
+ },
897
+ "direct_to_v3": {
898
+ "n": 11,
899
+ "fixes": 2,
900
+ "regressions": 0,
901
+ "winner_changes": 2,
902
+ "max_probability_delta": 0.4603535862018596
903
+ },
904
+ "v2_to_v3": {
905
+ "n": 11,
906
+ "fixes": 1,
907
+ "regressions": 0,
908
+ "winner_changes": 1,
909
+ "max_probability_delta": 0.48803378909777917
910
+ },
911
+ "historical_to_v3": {
912
+ "n": 11,
913
+ "fixes": 0,
914
+ "regressions": 0,
915
+ "winner_changes": 0,
916
+ "max_probability_delta": 0.0013950736901103267
917
+ }
918
+ },
919
+ "str": {
920
+ "direct": {
921
+ "n": 278,
922
+ "top1": 213,
923
+ "nll": 0.8905417724072999,
924
+ "brier": 0.38086923594388294,
925
+ "score_expected_level_mae": null
926
+ },
927
+ "v2": {
928
+ "n": 278,
929
+ "top1": 214,
930
+ "nll": 0.9050832731114038,
931
+ "brier": 0.3698537857923678,
932
+ "score_expected_level_mae": null
933
+ },
934
+ "v3": {
935
+ "n": 278,
936
+ "top1": 221,
937
+ "nll": 0.8639448083831234,
938
+ "brier": 0.34473711277242924,
939
+ "score_expected_level_mae": null
940
+ },
941
+ "historical": {
942
+ "n": 278,
943
+ "top1": 221,
944
+ "nll": 0.8639448083831234,
945
+ "brier": 0.34473711277242924,
946
+ "score_expected_level_mae": null
947
+ },
948
+ "direct_to_v3": {
949
+ "n": 278,
950
+ "fixes": 21,
951
+ "regressions": 13,
952
+ "winner_changes": 37,
953
+ "max_probability_delta": 0.4906821240628119
954
+ },
955
+ "v2_to_v3": {
956
+ "n": 278,
957
+ "fixes": 13,
958
+ "regressions": 6,
959
+ "winner_changes": 21,
960
+ "max_probability_delta": 0.498794986167734
961
+ },
962
+ "historical_to_v3": {
963
+ "n": 278,
964
+ "fixes": 0,
965
+ "regressions": 0,
966
+ "winner_changes": 0,
967
+ "max_probability_delta": 0.0
968
+ }
969
+ }
970
+ },
971
+ "by_kind": {
972
+ "choice": {
973
+ "direct": {
974
+ "n": 576,
975
+ "top1": 343,
976
+ "nll": 1.3449669496263348,
977
+ "brier": 0.5692832548094673,
978
+ "score_expected_level_mae": null
979
+ },
980
+ "v2": {
981
+ "n": 576,
982
+ "top1": 430,
983
+ "nll": 1.0454033035132229,
984
+ "brier": 0.4124806597366923,
985
+ "score_expected_level_mae": null
986
+ },
987
+ "v3": {
988
+ "n": 576,
989
+ "top1": 433,
990
+ "nll": 1.0170622984634021,
991
+ "brier": 0.40463624557307726,
992
+ "score_expected_level_mae": null
993
+ },
994
+ "historical": {
995
+ "n": 576,
996
+ "top1": 433,
997
+ "nll": 1.018041954435723,
998
+ "brier": 0.405391147268223,
999
+ "score_expected_level_mae": null
1000
+ },
1001
+ "direct_to_v3": {
1002
+ "n": 576,
1003
+ "fixes": 120,
1004
+ "regressions": 30,
1005
+ "winner_changes": 191,
1006
+ "max_probability_delta": 0.49877661297158915
1007
+ },
1008
+ "v2_to_v3": {
1009
+ "n": 576,
1010
+ "fixes": 27,
1011
+ "regressions": 24,
1012
+ "winner_changes": 81,
1013
+ "max_probability_delta": 0.4998410532110379
1014
+ },
1015
+ "historical_to_v3": {
1016
+ "n": 576,
1017
+ "fixes": 0,
1018
+ "regressions": 0,
1019
+ "winner_changes": 2,
1020
+ "max_probability_delta": 0.49223936331112117
1021
+ }
1022
+ },
1023
+ "noul": {
1024
+ "direct": {
1025
+ "n": 370,
1026
+ "top1": 311,
1027
+ "nll": 0.5029575433062489,
1028
+ "brier": 0.2629504138741171,
1029
+ "score_expected_level_mae": null
1030
+ },
1031
+ "v2": {
1032
+ "n": 370,
1033
+ "top1": 318,
1034
+ "nll": 0.45439529360118935,
1035
+ "brier": 0.22222817979439785,
1036
+ "score_expected_level_mae": null
1037
+ },
1038
+ "v3": {
1039
+ "n": 370,
1040
+ "top1": 321,
1041
+ "nll": 0.4465732399814046,
1042
+ "brier": 0.2164619524787992,
1043
+ "score_expected_level_mae": null
1044
+ },
1045
+ "historical": {
1046
+ "n": 370,
1047
+ "top1": 321,
1048
+ "nll": 0.44608433300775163,
1049
+ "brier": 0.21615934481882615,
1050
+ "score_expected_level_mae": null
1051
+ },
1052
+ "direct_to_v3": {
1053
+ "n": 370,
1054
+ "fixes": 28,
1055
+ "regressions": 18,
1056
+ "winner_changes": 46,
1057
+ "max_probability_delta": 0.4901768441417294
1058
+ },
1059
+ "v2_to_v3": {
1060
+ "n": 370,
1061
+ "fixes": 11,
1062
+ "regressions": 8,
1063
+ "winner_changes": 19,
1064
+ "max_probability_delta": 0.49871298353542987
1065
+ },
1066
+ "historical_to_v3": {
1067
+ "n": 370,
1068
+ "fixes": 0,
1069
+ "regressions": 0,
1070
+ "winner_changes": 0,
1071
+ "max_probability_delta": 0.028342231557091452
1072
+ }
1073
+ },
1074
+ "score": {
1075
+ "direct": {
1076
+ "n": 100,
1077
+ "top1": 74,
1078
+ "nll": 0.6867144478058433,
1079
+ "brier": 0.36762767783977907,
1080
+ "score_expected_level_mae": 0.29737852879492405
1081
+ },
1082
+ "v2": {
1083
+ "n": 100,
1084
+ "top1": 97,
1085
+ "nll": 0.2641991606057006,
1086
+ "brier": 0.12018344988441894,
1087
+ "score_expected_level_mae": 0.16328748801957052
1088
+ },
1089
+ "v3": {
1090
+ "n": 100,
1091
+ "top1": 97,
1092
+ "nll": 0.26707562707880517,
1093
+ "brier": 0.12199326949304623,
1094
+ "score_expected_level_mae": 0.16352370646939757
1095
+ },
1096
+ "historical": {
1097
+ "n": 100,
1098
+ "top1": 97,
1099
+ "nll": 0.2672208518023219,
1100
+ "brier": 0.12203424045574633,
1101
+ "score_expected_level_mae": 0.1636530071709598
1102
+ },
1103
+ "direct_to_v3": {
1104
+ "n": 100,
1105
+ "fixes": 25,
1106
+ "regressions": 2,
1107
+ "winner_changes": 27,
1108
+ "max_probability_delta": 0.4908280003843208
1109
+ },
1110
+ "v2_to_v3": {
1111
+ "n": 100,
1112
+ "fixes": 0,
1113
+ "regressions": 0,
1114
+ "winner_changes": 0,
1115
+ "max_probability_delta": 0.2012603587773023
1116
+ },
1117
+ "historical_to_v3": {
1118
+ "n": 100,
1119
+ "fixes": 0,
1120
+ "regressions": 0,
1121
+ "winner_changes": 0,
1122
+ "max_probability_delta": 0.0027665416307858237
1123
+ }
1124
+ }
1125
+ }
1126
+ },
1127
+ "typed-decisions": {
1128
+ "observed": 2000,
1129
+ "expected": 2000,
1130
+ "complete": true,
1131
+ "valid": 2000,
1132
+ "semantics": "synthetic teacher agreement",
1133
+ "errors": [],
1134
+ "fallbacks": [],
1135
+ "direct": {
1136
+ "n": 2000,
1137
+ "top1": 1237,
1138
+ "nll": 0.98012896186682,
1139
+ "brier": 0.5366594455578602,
1140
+ "score_expected_level_mae": 0.5718142114825938
1141
+ },
1142
+ "v2": {
1143
+ "n": 2000,
1144
+ "top1": 1342,
1145
+ "nll": 0.9213980450857948,
1146
+ "brier": 0.48476748880056186,
1147
+ "score_expected_level_mae": 0.5018408163629516
1148
+ },
1149
+ "v3": {
1150
+ "n": 2000,
1151
+ "top1": 1366,
1152
+ "nll": 0.8821571197982664,
1153
+ "brier": 0.4700707747544247,
1154
+ "score_expected_level_mae": 0.48623367185770394
1155
+ },
1156
+ "historical": {
1157
+ "n": 2000,
1158
+ "top1": 1365,
1159
+ "nll": 0.8821582687252656,
1160
+ "brier": 0.4699970802405581,
1161
+ "score_expected_level_mae": 0.48607930491573853
1162
+ },
1163
+ "direct_to_v3": {
1164
+ "n": 2000,
1165
+ "fixes": 334,
1166
+ "regressions": 205,
1167
+ "winner_changes": 632,
1168
+ "max_probability_delta": 0.49734083274590535
1169
+ },
1170
+ "v2_to_v3": {
1171
+ "n": 2000,
1172
+ "fixes": 173,
1173
+ "regressions": 149,
1174
+ "winner_changes": 388,
1175
+ "max_probability_delta": 0.49990235636218977
1176
+ },
1177
+ "historical_to_v3": {
1178
+ "n": 2000,
1179
+ "fixes": 1,
1180
+ "regressions": 0,
1181
+ "winner_changes": 1,
1182
+ "max_probability_delta": 0.08040670116881293
1183
+ },
1184
+ "routed": 1777,
1185
+ "completed_blends": 1777,
1186
+ "max_direct_delta": 0.0,
1187
+ "historical_text_equal": 1777,
1188
+ "historical_text_compared": 1777,
1189
+ "generation_output_tokens": {
1190
+ "sum": 1367542,
1191
+ "mean": 769.5790658413056,
1192
+ "max": 2712,
1193
+ "distribution": {
1194
+ "n": 1777,
1195
+ "sum": 1367542,
1196
+ "mean": 769.5790658413056,
1197
+ "median": 762,
1198
+ "p95": 1051.1999999999998,
1199
+ "p99": 1254.3600000000001,
1200
+ "max": 2712
1201
+ }
1202
+ },
1203
+ "pipeline_seconds_all_cost": {
1204
+ "n": 2000,
1205
+ "sum": 3901.123203223222,
1206
+ "mean": 1.950561601611611,
1207
+ "median": 2.093935512995813,
1208
+ "p95": 3.0059627732262015,
1209
+ "p99": 3.7421011833066586,
1210
+ "max": 8.422256662975997
1211
+ },
1212
+ "pipeline_seconds_clean": {
1213
+ "n": 1984,
1214
+ "sum": 3871.4878648194717,
1215
+ "mean": 1.9513547705743306,
1216
+ "median": 2.0951888249837793,
1217
+ "p95": 3.0047559435595756,
1218
+ "p99": 3.743849764257903,
1219
+ "max": 8.422256662975997
1220
+ },
1221
+ "timing_excluded_count": 16,
1222
+ "v2_selective_seconds_clean_matched": {
1223
+ "n": 1984,
1224
+ "sum": 4021.9517563040135,
1225
+ "mean": 2.0271934255564585,
1226
+ "median": 2.178467755962629,
1227
+ "p95": 3.0571437860082367,
1228
+ "p99": 4.1093565620703165,
1229
+ "max": 8.320580558967777
1230
+ },
1231
+ "max_native_backend_allocated_bytes": 6753550336,
1232
+ "natural_eos_output_tokens": {
1233
+ "n": 1777,
1234
+ "sum": 1367542,
1235
+ "mean": 769.5790658413056,
1236
+ "median": 762,
1237
+ "p95": 1051.1999999999998,
1238
+ "p99": 1254.3600000000001,
1239
+ "max": 2712
1240
+ },
1241
+ "clean_phase_seconds": {
1242
+ "/v1/chat/completions": {
1243
+ "n": 1984,
1244
+ "sum": 3670.5011939010583,
1245
+ "mean": 1.8500510049904528,
1246
+ "median": 1.9909222425194457,
1247
+ "p95": 2.8783525532402563,
1248
+ "p99": 3.5913489963649803,
1249
+ "max": 8.18496916803997
1250
+ },
1251
+ "/v1/systemone": {
1252
+ "n": 1984,
1253
+ "sum": 187.3238072542008,
1254
+ "mean": 0.09441724155957702,
1255
+ "median": 0.09884031204273924,
1256
+ "p95": 0.12675707904854788,
1257
+ "p99": 0.14250179579830732,
1258
+ "max": 0.2218943398911506
1259
+ },
1260
+ "/v1/winnow/inspect": {
1261
+ "n": 1984,
1262
+ "sum": 12.908032948966138,
1263
+ "mean": 0.006506064994438577,
1264
+ "median": 0.006854985433164984,
1265
+ "p95": 0.008897835278185084,
1266
+ "p99": 0.01095620173728094,
1267
+ "max": 0.016994947916828096
1268
+ }
1269
+ },
1270
+ "clean_chat_metrics": {
1271
+ "n": 1763,
1272
+ "output_tokens": 1356905,
1273
+ "decode_seconds": 3597.4189079999996,
1274
+ "aggregate_decode_tokens_per_second": 377.18848838607374,
1275
+ "draft_proposed": 2043980,
1276
+ "draft_accepted": 847677,
1277
+ "draft_acceptance": 0.4147188328652922,
1278
+ "prefill_ms": {
1279
+ "n": 1763,
1280
+ "sum": 68639.242,
1281
+ "mean": 38.933205899035734,
1282
+ "median": 39.707,
1283
+ "p95": 45.585,
1284
+ "p99": 51.141939999999984,
1285
+ "max": 55.325
1286
+ },
1287
+ "decode_ms": {
1288
+ "n": 1763,
1289
+ "sum": 3597418.908,
1290
+ "mean": 2040.5098740782757,
1291
+ "median": 2028.45,
1292
+ "p95": 2874.6335999999988,
1293
+ "p99": 3605.586639999997,
1294
+ "max": 8142.273
1295
+ }
1296
+ },
1297
+ "clean_chat_metrics_natural_eos": {
1298
+ "n": 1763,
1299
+ "output_tokens": 1356905,
1300
+ "decode_seconds": 3597.4189079999996,
1301
+ "aggregate_decode_tokens_per_second": 377.18848838607374,
1302
+ "draft_proposed": 2043980,
1303
+ "draft_accepted": 847677,
1304
+ "draft_acceptance": 0.4147188328652922,
1305
+ "prefill_ms": {
1306
+ "n": 1763,
1307
+ "sum": 68639.242,
1308
+ "mean": 38.933205899035734,
1309
+ "median": 39.707,
1310
+ "p95": 45.585,
1311
+ "p99": 51.141939999999984,
1312
+ "max": 55.325
1313
+ },
1314
+ "decode_ms": {
1315
+ "n": 1763,
1316
+ "sum": 3597418.908,
1317
+ "mean": 2040.5098740782757,
1318
+ "median": 2028.45,
1319
+ "p95": 2874.6335999999988,
1320
+ "p99": 3605.586639999997,
1321
+ "max": 8142.273
1322
+ }
1323
+ },
1324
+ "by_state": {
1325
+ "dict": {
1326
+ "direct": {
1327
+ "n": 2000,
1328
+ "top1": 1237,
1329
+ "nll": 0.98012896186682,
1330
+ "brier": 0.5366594455578602,
1331
+ "score_expected_level_mae": 0.5718142114825938
1332
+ },
1333
+ "v2": {
1334
+ "n": 2000,
1335
+ "top1": 1342,
1336
+ "nll": 0.9213980450857948,
1337
+ "brier": 0.48476748880056186,
1338
+ "score_expected_level_mae": 0.5018408163629516
1339
+ },
1340
+ "v3": {
1341
+ "n": 2000,
1342
+ "top1": 1366,
1343
+ "nll": 0.8821571197982664,
1344
+ "brier": 0.4700707747544247,
1345
+ "score_expected_level_mae": 0.48623367185770394
1346
+ },
1347
+ "historical": {
1348
+ "n": 2000,
1349
+ "top1": 1365,
1350
+ "nll": 0.8821582687252656,
1351
+ "brier": 0.4699970802405581,
1352
+ "score_expected_level_mae": 0.48607930491573853
1353
+ },
1354
+ "direct_to_v3": {
1355
+ "n": 2000,
1356
+ "fixes": 334,
1357
+ "regressions": 205,
1358
+ "winner_changes": 632,
1359
+ "max_probability_delta": 0.49734083274590535
1360
+ },
1361
+ "v2_to_v3": {
1362
+ "n": 2000,
1363
+ "fixes": 173,
1364
+ "regressions": 149,
1365
+ "winner_changes": 388,
1366
+ "max_probability_delta": 0.49990235636218977
1367
+ },
1368
+ "historical_to_v3": {
1369
+ "n": 2000,
1370
+ "fixes": 1,
1371
+ "regressions": 0,
1372
+ "winner_changes": 1,
1373
+ "max_probability_delta": 0.08040670116881293
1374
+ }
1375
+ }
1376
+ },
1377
+ "by_kind": {
1378
+ "choice": {
1379
+ "direct": {
1380
+ "n": 600,
1381
+ "top1": 358,
1382
+ "nll": 1.0383454926500515,
1383
+ "brier": 0.5461734111343014,
1384
+ "score_expected_level_mae": null
1385
+ },
1386
+ "v2": {
1387
+ "n": 600,
1388
+ "top1": 390,
1389
+ "nll": 0.9958149854968003,
1390
+ "brier": 0.5063973319925225,
1391
+ "score_expected_level_mae": null
1392
+ },
1393
+ "v3": {
1394
+ "n": 600,
1395
+ "top1": 393,
1396
+ "nll": 0.9563012200742745,
1397
+ "brier": 0.4976482188673111,
1398
+ "score_expected_level_mae": null
1399
+ },
1400
+ "historical": {
1401
+ "n": 600,
1402
+ "top1": 393,
1403
+ "nll": 0.9558431778730259,
1404
+ "brier": 0.4973719628971465,
1405
+ "score_expected_level_mae": null
1406
+ },
1407
+ "direct_to_v3": {
1408
+ "n": 600,
1409
+ "fixes": 93,
1410
+ "regressions": 58,
1411
+ "winner_changes": 190,
1412
+ "max_probability_delta": 0.49734083274590535
1413
+ },
1414
+ "v2_to_v3": {
1415
+ "n": 600,
1416
+ "fixes": 50,
1417
+ "regressions": 47,
1418
+ "winner_changes": 123,
1419
+ "max_probability_delta": 0.49990235636218977
1420
+ },
1421
+ "historical_to_v3": {
1422
+ "n": 600,
1423
+ "fixes": 0,
1424
+ "regressions": 0,
1425
+ "winner_changes": 0,
1426
+ "max_probability_delta": 0.02832484742590488
1427
+ }
1428
+ },
1429
+ "noul": {
1430
+ "direct": {
1431
+ "n": 600,
1432
+ "top1": 438,
1433
+ "nll": 0.781763055105711,
1434
+ "brier": 0.4553453509151658,
1435
+ "score_expected_level_mae": null
1436
+ },
1437
+ "v2": {
1438
+ "n": 600,
1439
+ "top1": 470,
1440
+ "nll": 0.7073079185383874,
1441
+ "brier": 0.3751118454262236,
1442
+ "score_expected_level_mae": null
1443
+ },
1444
+ "v3": {
1445
+ "n": 600,
1446
+ "top1": 477,
1447
+ "nll": 0.6776685742578799,
1448
+ "brier": 0.36021695730742676,
1449
+ "score_expected_level_mae": null
1450
+ },
1451
+ "historical": {
1452
+ "n": 600,
1453
+ "top1": 476,
1454
+ "nll": 0.6782093598018603,
1455
+ "brier": 0.360367007607702,
1456
+ "score_expected_level_mae": null
1457
+ },
1458
+ "direct_to_v3": {
1459
+ "n": 600,
1460
+ "fixes": 73,
1461
+ "regressions": 34,
1462
+ "winner_changes": 107,
1463
+ "max_probability_delta": 0.49051290118046414
1464
+ },
1465
+ "v2_to_v3": {
1466
+ "n": 600,
1467
+ "fixes": 31,
1468
+ "regressions": 24,
1469
+ "winner_changes": 55,
1470
+ "max_probability_delta": 0.49906662495041354
1471
+ },
1472
+ "historical_to_v3": {
1473
+ "n": 600,
1474
+ "fixes": 1,
1475
+ "regressions": 0,
1476
+ "winner_changes": 1,
1477
+ "max_probability_delta": 0.010967217955198949
1478
+ }
1479
+ },
1480
+ "score": {
1481
+ "direct": {
1482
+ "n": 800,
1483
+ "top1": 441,
1484
+ "nll": 1.085240993850228,
1485
+ "brier": 0.59050954235755,
1486
+ "score_expected_level_mae": 0.5718142114825938
1487
+ },
1488
+ "v2": {
1489
+ "n": 800,
1490
+ "top1": 482,
1491
+ "nll": 1.0261529346880962,
1492
+ "brier": 0.550786838937345,
1493
+ "score_expected_level_mae": 0.5018408163629516
1494
+ },
1495
+ "v3": {
1496
+ "n": 800,
1497
+ "top1": 496,
1498
+ "nll": 0.9799154537465503,
1499
+ "brier": 0.5317780547550084,
1500
+ "score_expected_level_mae": 0.48623367185770394
1501
+ },
1502
+ "historical": {
1503
+ "n": 800,
1504
+ "top1": 496,
1505
+ "nll": 0.9798562685569995,
1506
+ "brier": 0.5316884727227588,
1507
+ "score_expected_level_mae": 0.48607930491573853
1508
+ },
1509
+ "direct_to_v3": {
1510
+ "n": 800,
1511
+ "fixes": 168,
1512
+ "regressions": 113,
1513
+ "winner_changes": 335,
1514
+ "max_probability_delta": 0.49632280685398744
1515
+ },
1516
+ "v2_to_v3": {
1517
+ "n": 800,
1518
+ "fixes": 92,
1519
+ "regressions": 78,
1520
+ "winner_changes": 210,
1521
+ "max_probability_delta": 0.49985733751079203
1522
+ },
1523
+ "historical_to_v3": {
1524
+ "n": 800,
1525
+ "fixes": 0,
1526
+ "regressions": 0,
1527
+ "winner_changes": 0,
1528
+ "max_probability_delta": 0.08040670116881293
1529
+ }
1530
+ }
1531
+ }
1532
+ }
1533
+ }
1534
+ }
docs/evidence/e2b-v3-completion.json ADDED
@@ -0,0 +1,24 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "done": 3277,
3
+ "total": 3277,
4
+ "panels": {
5
+ "jevbench-public": 231,
6
+ "kev-v9-clean": 1046,
7
+ "typed-decisions": 2000
8
+ },
9
+ "routed": 2584,
10
+ "completed_blends": 2582,
11
+ "fallbacks": 2,
12
+ "errors": 0,
13
+ "max_direct_delta": 0.0,
14
+ "recovered_diagnosed_cases": [
15
+ 1001,
16
+ 1236
17
+ ],
18
+ "recovery_elapsed_seconds": 1326.1061786840437,
19
+ "time_utc": "2026-10-06T16:38:00Z",
20
+ "complete": true,
21
+ "source_commit": "ee6bd37d34ae35d2e69ebb0c4b0957a10c727357",
22
+ "freeze_sha256": "eb163f63b0eb3a706adaaf18d20ebe3acaef4ddb0690005149eef2cbc750d2a6",
23
+ "output_sha256": "db3d7cb81972ad2909a75aca659b97015d9e96a770f04620275b79b9fbdbd80d"
24
+ }
docs/evidence/e2b-v3-freeze.json ADDED
@@ -0,0 +1,142 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "source_commit": "ee6bd37d34ae35d2e69ebb0c4b0957a10c727357",
3
+ "parent_commit": "947bf3ad350e180e5fb126ee80c921fab19465b5",
4
+ "source_archive_sha256": "999179036fa60ab7524706bf78050cda8f70d130fbfd756f6ab41d5325f1ed71",
5
+ "policy_id": "e2b-raw99-blend50-v3",
6
+ "runtime_profile": "e2b-q8-text8k-mtp",
7
+ "panels": {
8
+ "jevbench-public": 231,
9
+ "kev-v9-clean": 1046,
10
+ "typed-decisions": 2000
11
+ },
12
+ "baseline_files": {
13
+ "jevbench-public": "8384afe4d943ce1ddfb9e86fde9599d773a0023229094c9c2ef5126e56a00de9",
14
+ "kev-v9-clean": "0702129a098dfb2dcc33dc8c64fa67180c829532c5fa470bc2e279d42ec8bf86",
15
+ "typed-decisions": "5b02533afcd7ff0845bc755d80d70e35584fce3748b4cd5c787f9e7bca5c3aa7"
16
+ },
17
+ "direct_tolerance": 1e-10,
18
+ "wall_cap_seconds": 10800,
19
+ "files": {
20
+ "FULL-VALIDATION-PROTOCOL.md": "c249f73a9499a21a1bc65c5220691219d3f49a84b6357bc1ce2c8132e0f55b01",
21
+ "direct-reference.jsonl": "98914abd055858cc9a9dbc435397059c60c1e97497ba02dee244e28882114865",
22
+ "e2b-v3-source.tar.gz": "999179036fa60ab7524706bf78050cda8f70d130fbfd756f6ab41d5325f1ed71",
23
+ "frozen-requests.jsonl": "bedf68610bc24dbf660bca54fd6d5a47dabb9e46404d059e2439121e8025c9ca",
24
+ "full_controller.py": "28041d4dde21a34471230276dce525d1a5f3b4e6a0d1957e932f32525c336091",
25
+ "full_probe.py": "f595767c69b2a9c8df7a340cc8572c33c235ed4abd4f658d0ee0dff0c67e607e"
26
+ },
27
+ "source_files": {
28
+ ".clang-format": "96cd8537a8a322552fd6d23a9e3470e6e748233116f05093e10c334cdf4eb787",
29
+ ".dockerignore": "9712f1e0aa72bf285399c336dc20937acb3d7171a8e1eb47db6707ed0cb02466",
30
+ ".gitattributes": "0bcc1cafcbe5d86f8f6f245e0db28ae06eb9b403dba8821d72fa7c51d96d249d",
31
+ ".gitignore": "d6497c2e377dc24fa49e418c95bd3136ee22b15e2b632897e7ea2c9821dba2f4",
32
+ "CMakeLists.txt": "5b797d4ca694e875803b02b7979cfa89c3d8fc64cccb2a0d61ccbdb2466ed01d",
33
+ "Dockerfile": "9da5f2a7b799c8ea28d9963c61e52ff4b0e7937631d2363775e632d026340034",
34
+ "LICENSE": "6cc43947bc50233b132fe3cf248c86aa70ed38c7d7b5253e01ea5b3f65fe88e0",
35
+ "README.md": "914df5cb8e8b010ad6d5850085ce813ca4fcc56e3c9dc9f66585020712e0d64b",
36
+ "docs/ADAPTIVE.md": "d84abd5a17ccaf3c4457a8edc3cd4b70c2a65a5b3015c99383a6a648a0c52ae1",
37
+ "docs/API.md": "aa93d4be062cdceb42e8ad9a0618cf742e7986bb5eff2d82e23491778ea2ecda",
38
+ "docs/BENCHMARKS.md": "319118495adde81bda28bec6d25886c3edd787a7e41d2ba83305c881efec96d5",
39
+ "docs/E2B.md": "fe0e75c6dba1730d7a948481c2b5d441c79e0301d912398ae1b97cb21aeb4715",
40
+ "docs/EVALUATION.md": "60bc73264c9373aa2526016c6f52a312210510a58a476419d33caae51f90fb29",
41
+ "docs/IMAGE-REASONING.md": "ddc2a4efa91c411c2c6a5733cbf686d25a0d7b1f3d97c6fb16bf2092ca026d99",
42
+ "docs/IMPLEMENTATION.md": "8093f0dbd9e766fdfc994016e1487889d09057910d28fc50e366ba30ea428414",
43
+ "docs/INSTALL.md": "439b55cdbbe45d4cce4b3d36aa7cdbe4e8f05ca2f1657f71b196099f27259cf1",
44
+ "docs/PRESETS-AND-ASSETS.md": "afa29722eb43cc75cfec2fc3278dd6421cb2ba90732c4a25a2ba353b52ecaba7",
45
+ "docs/QUICKSTART.md": "deedf4fb4bb8587b3d50c6d715ccbea5fb7851847bd5e973c7f75a8c2d83ac61",
46
+ "docs/REASONING-RESULTS.md": "fa12e0a4f95845aeec5b219f25a5c309e677a2c3a8fed2f3f5441c1c79edef50",
47
+ "docs/RELEASE.md": "521b0dea553bec89dd3b8f936afa28c9b48168a40489c64976079b33a7e2eb97",
48
+ "docs/VALIDATION.md": "2e4a2aeff052f4fb6ec76c3ac4673572ba070befe0c30c77e5a3a3dc56508a76",
49
+ "docs/assets/01-5070ti-capability.png": "ae66377e15a8f72f1d9013dc6287b7f8ee0183fca982075a9d2abe7393631b62",
50
+ "docs/assets/01-5070ti-capability.svg": "3edf0643e3523530b1f583bad8a0fa6cb91af4dba4f7508dee331df8c0e47040",
51
+ "docs/assets/02-decision-quality.png": "6b8fc874fbebe0dae497ebb485b373db674461f88d1544b1d465213fb1df1bc9",
52
+ "docs/assets/02-decision-quality.svg": "e618d9a18241fce10fe53bfb1bcbc3a3806993a7ad640fb2f161dd23f70968cf",
53
+ "docs/assets/03-deployment-memory.png": "59ae7f71e9fd54a8ecc5d7540812a044802d8e0ef95508ba54b9a892be3d821c",
54
+ "docs/assets/03-deployment-memory.svg": "2134d134df1a7f64f98105b5b8df8ba7fb84839e125df190e8cc194aafc34034",
55
+ "docs/assets/04-inference-speed.png": "e0620bb64850df6cc63304ddbf56aefb77cb8d5d36a640479bfe4aeb81fb7773",
56
+ "docs/assets/04-inference-speed.svg": "08699674712723922413c7b89a3480117b26024d7a91a21da616bf5a762b263e",
57
+ "docs/assets/12b-03-q8-reasoning.png": "937894a60a0fa571116e98a5250e47a3e7e3a66f6782199d8ab1263e7a490f84",
58
+ "docs/assets/12b-03-q8-reasoning.svg": "a3511ade45f22cb18dcac6648d522af6c7132b7d1b8cebf8c8e247254a923449",
59
+ "docs/assets/IBM-Plex-Mono-LICENSE.txt": "7e6b2818edbd8f6a01ae80641cc8f16a51080d08fb4e532be3a0b6f74adb07da",
60
+ "docs/assets/Instrument-Sans-LICENSE.txt": "9e27a72ed30eb49a08678f6a5d6ed98ec7ba5368f541637ee0683ec9134ef966",
61
+ "docs/assets/e4b-01-reasoning-outcomes.png": "37746e0cc052931e02d707072ee4a2424a6a7e8386121c687a042960c1610527",
62
+ "docs/assets/e4b-01-reasoning-outcomes.svg": "f86516098d157457b5f1df7e98c77f4356f01636dc40230911ae12077ba4f046",
63
+ "docs/assets/e4b-02-reasoning-latency.png": "e8ea198a76a638d41dc3eeeba1511f23434b2dcc0321f63b105e254485f3d0ea",
64
+ "docs/assets/e4b-02-reasoning-latency.svg": "565c71da5306159f455140fe7160dc20fc8513a225abc68a3d607ac9432e1501",
65
+ "docs/benchmarks.json": "84a7e835d1dce3d079c4d5206f8cd2ba8cff6dca3aac1600de20f140cd82513c",
66
+ "docs/e2b-evidence.json": "4a28f9f39f5de4015591675c6e4d5e52ebe78c9932f9367c6c2e61e7bd922e25",
67
+ "docs/reasoning-results.json": "d3a0c5c9dd2ff856eaa330b2002da11cbdd40c38aa0297ff19cf05703c3816f8",
68
+ "examples/adaptive-decision.json": "01a243f2dff4097c0a61c888706047d1951dbb539c5592e84d0b905f740abff6",
69
+ "examples/client.mjs": "e1d4d697b21f518acedf117940fbea2bc42c0489e3a1f72cad1b8444452b493a",
70
+ "examples/client.py": "a2ebcda4327947bf35703012440f43714da8c5cc4eb9bc131567d2ab4b218b24",
71
+ "examples/decisions.json": "f30cd9b930f62b12c4a4c13a94e029b33a818cd4d87a883f44e0df833b4135d3",
72
+ "manifests/adaptive-v1.json": "9a38af7981f2c035ee11f29cce02751dcef798c5149d19c376f34bfbdb125eb9",
73
+ "manifests/assistants-v1.json": "4e1e7a63b4a5676c26323574a7b8a8577cb25c986f89c4f9163ee16fcb26360a",
74
+ "manifests/models.json": "62d53fdfd1b68def79f68e2ab957e7d2e2ce5ec26d09902676248a8bc3cd72a0",
75
+ "manifests/release-assets-v1.json": "23bd549612c31ee9f373ca64b825757094b62e16801d9ead6d308cd428acfc25",
76
+ "manifests/runtime-presets-v1.json": "4fb87a1ae9ffd4e00c95e465b9440c292b1d8791fd711098f06f6ed7bf9dae79",
77
+ "native/bridge.cpp": "c8815cfc741238011612e49cca6112b79289acff7a9ba74b7abcab873fdced98",
78
+ "native/bridge.h": "6db293f54c7f9411d5e9b99e1d7b1f5ed066d515066324b772d5566f7a49af9f",
79
+ "native/engine.h": "9a27e5ad9644b69424aacba0f54a36e8c489733aed01c3918c6cf50012d629ac",
80
+ "native/images.h": "5200be571801ec4ab107c06f2364dfa136a71b177ccc1c8ec5db927fa3630a34",
81
+ "native/planner.h": "37cac384797573c55c8d95684dad3f149aa69d2341982f763332e5166822fe86",
82
+ "native/protocol.h": "bc9918e5fe0c072c5a5dc54e4e6a9c73adeb30e08f7dc7a0abdbfc1d11f3b2f5",
83
+ "patches/llama-bounded-swa-fork.patch": "79f095b43ea7e072aaf51ab5f3a3edd290d2c598fadd8450f0ab508306715561",
84
+ "patches/llama-classifier-head.patch": "0c381d753cdf5705cd55e8b72d07904241ae45e3d6005010545031774594237d",
85
+ "patches/llama-e2b-mtp-shape.patch": "47e85e9bff800e1c46934f70bc65571c372ea6d170b06a9a3fad57ac9a955bb0",
86
+ "patches/llama-gemma4-tied-embedding.patch": "caeaa51bec0c2dc5c8e48ccea2f5fce227333f60013e749f92152b87f8250b32",
87
+ "patches/llama-managed-config.patch": "ab58bd6ff5ea296dedb5871f18aa51f11f8de18be56b7bcee52b43f1c38b5a95",
88
+ "patches/llama-resident-mtp-opt-in.patch": "f3c66d57b2202c4b2e9ef3fdd366c3685c045c22f9c0a9dc5c70c376e41d4d45",
89
+ "patches/llama-vision-mtp-state-guard.patch": "3c24df4e4d234824603254ad8a0ceb7ba95ea6d2bae6e3a973aaa24c947a46c5",
90
+ "patches/llama-winnow-server.patch": "608c8642d164a5676cb817ae5d7e851148dd6bd008459cc2b3080a1652b39423",
91
+ "ruff.toml": "21ec6966cdc163503d041d2e0b844799f43b1cd9dbe915fe82263e90219536d5",
92
+ "runtime.lock.json": "9b20caf5c6032f6274350c8d20b1d90e5cbc5b7dfa468ca85c77e198e24845df",
93
+ "scripts/adaptive_policy.py": "43c8652373e949778ab9e099b6a52c338f94d71724ba10e70f6729b1f9f858c8",
94
+ "scripts/assets.py": "23eec166f908fe411835bb6acd130da16c57061a60c1667ebd39643526b4f5b0",
95
+ "scripts/bench.py": "204822b2c58370791dcf2acf090b9e35c13e32e6b0b3061d909cedc4adb42145",
96
+ "scripts/build.py": "045df273eb93bc918feeedf37286c2eecaf6a0cebc93c6a9d61b1de60db5bc75",
97
+ "scripts/chat_capacity.py": "8aac9044024a1a2278083c407f57c113ead60d4d4b270c599088b91b55ece4fb",
98
+ "scripts/check.py": "af58b3082893e68052ecac98bb3bbf4db0a7e676fda57fe11f2ae4767c196562",
99
+ "scripts/check_http.py": "f3b8a26f4176c50675f73960f24f2f934336e5c568bdbf68a085d842794bff5e",
100
+ "scripts/check_package.py": "329e6d3c9e2e64052985d87c7ff2c40c7bb9e00f71a1d8b6e16cae2778fa39ea",
101
+ "scripts/decision_client.py": "ceaff3cdb7ab505e0427fca4e96daeed041a6110918498f7282ca5ad92bc2c86",
102
+ "scripts/download.py": "c385532c1b2e5586af0ffcac68ac3f05b0f8d8d0f71105e8c7961dbb569e5e0b",
103
+ "scripts/http_client.py": "b4d267d4f21a978915c7fc62a33e794387bcd626eafcc6cb6c6415e182e8d08d",
104
+ "scripts/install_assistants.py": "fa5c75fe8292cfdee5977e67c4c94fade941a0dd729599cae20cb2edc0255256",
105
+ "scripts/launch_options.py": "77ced9e55dd028065f133e62c06427889da55803880544ab1fdf41a0570c6567",
106
+ "scripts/lifecycle.py": "701438434c3752fa9fc17df46f74860e04500d947853f3915b9dbecca60bdab9",
107
+ "scripts/package_local_candidate.py": "5d66f7f7022e54d358846ee855e0c5fc378822cd747d7ff76e93a1df2275221c",
108
+ "scripts/package_macos.py": "12050d90958e7bdabb0b0b3275cfd719f401726cc51f1ecb55f28e82788ee3c6",
109
+ "scripts/package_source.py": "f2a207e1771dc2f721b96f774bce0a1b4c89fe1bc4566d40a672f8a7cf38dee2",
110
+ "scripts/parity.py": "a943df2ca7c27bd85435992dea76d3eb3de76ea59391fc19174d4981a6b33d21",
111
+ "scripts/profiles.py": "dd4dd47c36bb9f5f5496988ad6192b5306eacf941c73324f18c88b954b319821",
112
+ "scripts/reasoning_contract.py": "ec71323a25965411f71b23adeb7c278b9942928c4471bc238771e3fb8806442d",
113
+ "scripts/release_check.py": "22f1c53ca42a53909292bb867a3dc3a3b8e51dfc003db11eca84fba002344ef2",
114
+ "scripts/serve.py": "67d748639c7691c6a350b44a29dd0ffbf7a783bdca6fadc0d1519e0f53581e73",
115
+ "scripts/setup.py": "1a963c2118e1ac14fdab1cc94976b11535be034574c7cfbdb160ac721845f84a",
116
+ "scripts/verify_model.py": "31e8d97669cadedf71c219c1e51ff1a794444c3ad91a85bc50079821ef794fe9",
117
+ "scripts/winnow.py": "2d4c540800a0bc89c8966db479d7188e0571b64a3a961a2b000852be1c5671d8",
118
+ "tests/parity-requests.json": "f0316bfba49d8bae27f7ec06cdc5dc165e4c57623363287ddb6b49f42df2e768",
119
+ "tests/test_adaptive_pipeline.py": "33ae167dc2ef2ed35a89be8b965cd17247d1177a787da3a430fae706e8d5b5da",
120
+ "tests/test_assistant_install.py": "1ef64a0176bebbf9a29911ec3008106121e064425627cf8c32f291937e211a93",
121
+ "tests/test_e2b_native_labels.py": "be93515a0235a7c7c0ae03d666e567b130205db100e59f36e709f8e8573ceda2",
122
+ "tests/test_e2b_release.py": "855e6328d9ac5816bd8e6ace07ecd8fd26e521b33041a85e8e7b4e4cde4f7ea8",
123
+ "tests/test_image_reasoning.py": "73c3e8466c570babbfa1981705fec6f88b53adaa5613ae7f1fce5c370f661daf",
124
+ "tests/test_onboarding.py": "915677b83349f2c5f01a63af4bf48b4a6c0d8dced3259029c103df534daa64c1",
125
+ "tests/test_public_reasoning_parity.py": "835c0662ef551081b1a9a129ab8852686effd0cfbdfe45e9269787edae379047",
126
+ "tests/test_reasoning_modes.py": "ecd9b40e7d5a209ccafdbc425b98acf8bf73af8fb14a799b70361992eec7e03c",
127
+ "tests/test_release_tools.py": "0c568317107231ee6ec3f01d8d02ed998333ca2d7c97df2eee79333868f9c0d1",
128
+ "tests/test_release_usability.py": "6550dee609d693a859639a273b50b9bfff49499f0c5435ddb63163547ab77182",
129
+ "tests/test_runtime_package.py": "de41b66660bbb2a45a0eab3b9930377fad6c55545b8bbfada266d07441b58de3",
130
+ "tests/test_runtime_patch_lock.py": "af9d28cd7ff4e62960b431f90875ed59a6ce27f674d1215f5849fb79616f4d35",
131
+ "tests/test_runtime_presets.py": "e91aee90d2983277d84f4b8c6c67ac524fe02755fd0f83df556d6e95e136bf80",
132
+ "tests/test_transport_cancellation.py": "1204773b4940ee856b577cdc4f244ebe84c49a30057cadad7251bb2f49081780",
133
+ "tests/unit.cpp": "1c8d8e3023f9ff7d4c4e9e5f2e200e7410f193f2ece7efdf7518a49201240ec4",
134
+ "third_party/README.md": "52d1407fd8944dd816d0f2ec078242882b3f6958edd1ed609785fa88ba0f6b8c",
135
+ "third_party/cpp-httplib-LICENSE.txt": "4b45cbe16d7b71b89ae6127e26e0d90a029198ca5e958ad8e3d0b8bbed364d8b",
136
+ "third_party/gemma-assistants-LICENSE-APACHE-2.0.txt": "cfc7749b96f63bd31c3c42b5c471bf756814053e847c10f3eb003417bc523d30",
137
+ "third_party/gemma-assistants-NOTICE.txt": "f54573f51292300e9aa87df0d2351260e78c2bd220f9249fcc9eded1e448b7df",
138
+ "third_party/llama.cpp-LICENSE": "94f29bbed6a22c35b992c5c6ebf0e7c92f13b836b90f36f461c9cf2f0f1d010d",
139
+ "third_party/nlohmann-json-LICENSE.txt": "c0d068392ea65358b798b8c165103560f06e9e3b38c4ab4e2d8810a7b931af86",
140
+ "third_party/rotate-bits-LICENSE.txt": "ca0b44aec101afb6b46f4c39b23369fa06a64fb9c0add37c2689d702156fee55"
141
+ }
142
+ }
release-manifest.json CHANGED
@@ -22,7 +22,10 @@
22
  "converter_sha256": "e9a1da876330bbce9687541ab31736542a01b4ac43c6686126514a50f122fb7f",
23
  "merge_receipt_sha256": "36e83f26a69bc797cef332c2b4c93aadb41dcd3860dbac8446410da872ad05a3",
24
  "gguf_file_type": 32,
25
- "tensor_types": {"bf16": 318, "f32": 283},
 
 
 
26
  "tensor_layout_sha256": "05a32f8c9728812a06d6c7075be4e5daaf8b678bc62c1af92136b618987d3227",
27
  "evaluation_scope": "Not evaluated or calibrated in the reported Q8 tests."
28
  },
@@ -57,6 +60,7 @@
57
  "raw_max_probability_gate": 0.99,
58
  "direct_probability_weight": 0.5,
59
  "augmented_probability_weight": 0.5,
60
- "evidence_scope": "Previously exposed full text panels; current v2 integration prompt mapping differs from historical benchmark serialization."
 
61
  }
62
  }
 
22
  "converter_sha256": "e9a1da876330bbce9687541ab31736542a01b4ac43c6686126514a50f122fb7f",
23
  "merge_receipt_sha256": "36e83f26a69bc797cef332c2b4c93aadb41dcd3860dbac8446410da872ad05a3",
24
  "gguf_file_type": 32,
25
+ "tensor_types": {
26
+ "bf16": 318,
27
+ "f32": 283
28
+ },
29
  "tensor_layout_sha256": "05a32f8c9728812a06d6c7075be4e5daaf8b678bc62c1af92136b618987d3227",
30
  "evaluation_scope": "Not evaluated or calibrated in the reported Q8 tests."
31
  },
 
60
  "raw_max_probability_gate": 0.99,
61
  "direct_probability_weight": 0.5,
62
  "augmented_probability_weight": 0.5,
63
+ "evidence_scope": "Previously exposed full text panels; current v2 integration prompt mapping differs from historical benchmark serialization.",
64
+ "draft_cache_note": "The draft_kv value records retained launch flags. The pinned Gemma 4 assistant shares target K/V tensors and uses F16 attention cache in this evaluated profile; it has no separately measured Q8 attention-cache pool."
65
  }
66
  }