yqi0 commited on
Commit
af1f7fa
·
verified ·
1 Parent(s): 3379124

Publish petitgpt research-v1 native alpha075 model and scoped documentation

Browse files
DOCUMENTATION_LICENSE.md ADDED
@@ -0,0 +1,9 @@
 
 
 
 
 
 
 
 
 
 
1
+ # Licence scope
2
+
3
+ Copyright 2026 Yang Qi. Owner-controlled code and the selected model/tokenizer are licensed under the standard Apache License 2.0 in LICENSE, only for rights Yang Qi is entitled to grant.
4
+
5
+ As an explicit exception to the root code licence, author-written reports and documentation (including README, native run guide and versioned report) are licensed under Creative Commons Attribution 4.0 International (CC BY 4.0): https://creativecommons.org/licenses/by/4.0/ and https://creativecommons.org/licenses/by/4.0/legalcode.en . Attribute Yang Qi and petitgpt, link the licence, and indicate changes. Existing third-party content and notices retain their applicable terms; they are not relicensed.
6
+
7
+ Research describes intended use; it adds no noncommercial or research-only restriction to Apache-licensed artifacts. This grant was approved by the owner through the explicit research-release execution instruction. Earlier PENDING_OWNER_DECISION records remain historical evidence; this does not claim an earlier licence choice.
8
+
9
+ Source metadata is not rights clearance. Weights are not the raw corpus; neither automatic inheritance nor automatic non-application of all dataset terms is asserted. No infringement guarantee or legal certification is given. A disclaimer does not replace applicable permission.
LICENSE ADDED
@@ -0,0 +1,202 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+
2
+ Apache License
3
+ Version 2.0, January 2004
4
+ http://www.apache.org/licenses/
5
+
6
+ TERMS AND CONDITIONS FOR USE, REPRODUCTION, AND DISTRIBUTION
7
+
8
+ 1. Definitions.
9
+
10
+ "License" shall mean the terms and conditions for use, reproduction,
11
+ and distribution as defined by Sections 1 through 9 of this document.
12
+
13
+ "Licensor" shall mean the copyright owner or entity authorized by
14
+ the copyright owner that is granting the License.
15
+
16
+ "Legal Entity" shall mean the union of the acting entity and all
17
+ other entities that control, are controlled by, or are under common
18
+ control with that entity. For the purposes of this definition,
19
+ "control" means (i) the power, direct or indirect, to cause the
20
+ direction or management of such entity, whether by contract or
21
+ otherwise, or (ii) ownership of fifty percent (50%) or more of the
22
+ outstanding shares, or (iii) beneficial ownership of such entity.
23
+
24
+ "You" (or "Your") shall mean an individual or Legal Entity
25
+ exercising permissions granted by this License.
26
+
27
+ "Source" form shall mean the preferred form for making modifications,
28
+ including but not limited to software source code, documentation
29
+ source, and configuration files.
30
+
31
+ "Object" form shall mean any form resulting from mechanical
32
+ transformation or translation of a Source form, including but
33
+ not limited to compiled object code, generated documentation,
34
+ and conversions to other media types.
35
+
36
+ "Work" shall mean the work of authorship, whether in Source or
37
+ Object form, made available under the License, as indicated by a
38
+ copyright notice that is included in or attached to the work
39
+ (an example is provided in the Appendix below).
40
+
41
+ "Derivative Works" shall mean any work, whether in Source or Object
42
+ form, that is based on (or derived from) the Work and for which the
43
+ editorial revisions, annotations, elaborations, or other modifications
44
+ represent, as a whole, an original work of authorship. For the purposes
45
+ of this License, Derivative Works shall not include works that remain
46
+ separable from, or merely link (or bind by name) to the interfaces of,
47
+ the Work and Derivative Works thereof.
48
+
49
+ "Contribution" shall mean any work of authorship, including
50
+ the original version of the Work and any modifications or additions
51
+ to that Work or Derivative Works thereof, that is intentionally
52
+ submitted to Licensor for inclusion in the Work by the copyright owner
53
+ or by an individual or Legal Entity authorized to submit on behalf of
54
+ the copyright owner. For the purposes of this definition, "submitted"
55
+ means any form of electronic, verbal, or written communication sent
56
+ to the Licensor or its representatives, including but not limited to
57
+ communication on electronic mailing lists, source code control systems,
58
+ and issue tracking systems that are managed by, or on behalf of, the
59
+ Licensor for the purpose of discussing and improving the Work, but
60
+ excluding communication that is conspicuously marked or otherwise
61
+ designated in writing by the copyright owner as "Not a Contribution."
62
+
63
+ "Contributor" shall mean Licensor and any individual or Legal Entity
64
+ on behalf of whom a Contribution has been received by Licensor and
65
+ subsequently incorporated within the Work.
66
+
67
+ 2. Grant of Copyright License. Subject to the terms and conditions of
68
+ this License, each Contributor hereby grants to You a perpetual,
69
+ worldwide, non-exclusive, no-charge, royalty-free, irrevocable
70
+ copyright license to reproduce, prepare Derivative Works of,
71
+ publicly display, publicly perform, sublicense, and distribute the
72
+ Work and such Derivative Works in Source or Object form.
73
+
74
+ 3. Grant of Patent License. Subject to the terms and conditions of
75
+ this License, each Contributor hereby grants to You a perpetual,
76
+ worldwide, non-exclusive, no-charge, royalty-free, irrevocable
77
+ (except as stated in this section) patent license to make, have made,
78
+ use, offer to sell, sell, import, and otherwise transfer the Work,
79
+ where such license applies only to those patent claims licensable
80
+ by such Contributor that are necessarily infringed by their
81
+ Contribution(s) alone or by combination of their Contribution(s)
82
+ with the Work to which such Contribution(s) was submitted. If You
83
+ institute patent litigation against any entity (including a
84
+ cross-claim or counterclaim in a lawsuit) alleging that the Work
85
+ or a Contribution incorporated within the Work constitutes direct
86
+ or contributory patent infringement, then any patent licenses
87
+ granted to You under this License for that Work shall terminate
88
+ as of the date such litigation is filed.
89
+
90
+ 4. Redistribution. You may reproduce and distribute copies of the
91
+ Work or Derivative Works thereof in any medium, with or without
92
+ modifications, and in Source or Object form, provided that You
93
+ meet the following conditions:
94
+
95
+ (a) You must give any other recipients of the Work or
96
+ Derivative Works a copy of this License; and
97
+
98
+ (b) You must cause any modified files to carry prominent notices
99
+ stating that You changed the files; and
100
+
101
+ (c) You must retain, in the Source form of any Derivative Works
102
+ that You distribute, all copyright, patent, trademark, and
103
+ attribution notices from the Source form of the Work,
104
+ excluding those notices that do not pertain to any part of
105
+ the Derivative Works; and
106
+
107
+ (d) If the Work includes a "NOTICE" text file as part of its
108
+ distribution, then any Derivative Works that You distribute must
109
+ include a readable copy of the attribution notices contained
110
+ within such NOTICE file, excluding those notices that do not
111
+ pertain to any part of the Derivative Works, in at least one
112
+ of the following places: within a NOTICE text file distributed
113
+ as part of the Derivative Works; within the Source form or
114
+ documentation, if provided along with the Derivative Works; or,
115
+ within a display generated by the Derivative Works, if and
116
+ wherever such third-party notices normally appear. The contents
117
+ of the NOTICE file are for informational purposes only and
118
+ do not modify the License. You may add Your own attribution
119
+ notices within Derivative Works that You distribute, alongside
120
+ or as an addendum to the NOTICE text from the Work, provided
121
+ that such additional attribution notices cannot be construed
122
+ as modifying the License.
123
+
124
+ You may add Your own copyright statement to Your modifications and
125
+ may provide additional or different license terms and conditions
126
+ for use, reproduction, or distribution of Your modifications, or
127
+ for any such Derivative Works as a whole, provided Your use,
128
+ reproduction, and distribution of the Work otherwise complies with
129
+ the conditions stated in this License.
130
+
131
+ 5. Submission of Contributions. Unless You explicitly state otherwise,
132
+ any Contribution intentionally submitted for inclusion in the Work
133
+ by You to the Licensor shall be under the terms and conditions of
134
+ this License, without any additional terms or conditions.
135
+ Notwithstanding the above, nothing herein shall supersede or modify
136
+ the terms of any separate license agreement you may have executed
137
+ with Licensor regarding such Contributions.
138
+
139
+ 6. Trademarks. This License does not grant permission to use the trade
140
+ names, trademarks, service marks, or product names of the Licensor,
141
+ except as required for reasonable and customary use in describing the
142
+ origin of the Work and reproducing the content of the NOTICE file.
143
+
144
+ 7. Disclaimer of Warranty. Unless required by applicable law or
145
+ agreed to in writing, Licensor provides the Work (and each
146
+ Contributor provides its Contributions) on an "AS IS" BASIS,
147
+ WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or
148
+ implied, including, without limitation, any warranties or conditions
149
+ of TITLE, NON-INFRINGEMENT, MERCHANTABILITY, or FITNESS FOR A
150
+ PARTICULAR PURPOSE. You are solely responsible for determining the
151
+ appropriateness of using or redistributing the Work and assume any
152
+ risks associated with Your exercise of permissions under this License.
153
+
154
+ 8. Limitation of Liability. In no event and under no legal theory,
155
+ whether in tort (including negligence), contract, or otherwise,
156
+ unless required by applicable law (such as deliberate and grossly
157
+ negligent acts) or agreed to in writing, shall any Contributor be
158
+ liable to You for damages, including any direct, indirect, special,
159
+ incidental, or consequential damages of any character arising as a
160
+ result of this License or out of the use or inability to use the
161
+ Work (including but not limited to damages for loss of goodwill,
162
+ work stoppage, computer failure or malfunction, or any and all
163
+ other commercial damages or losses), even if such Contributor
164
+ has been advised of the possibility of such damages.
165
+
166
+ 9. Accepting Warranty or Additional Liability. While redistributing
167
+ the Work or Derivative Works thereof, You may choose to offer,
168
+ and charge a fee for, acceptance of support, warranty, indemnity,
169
+ or other liability obligations and/or rights consistent with this
170
+ License. However, in accepting such obligations, You may act only
171
+ on Your own behalf and on Your sole responsibility, not on behalf
172
+ of any other Contributor, and only if You agree to indemnify,
173
+ defend, and hold each Contributor harmless for any liability
174
+ incurred by, or claims asserted against, such Contributor by reason
175
+ of your accepting any such warranty or additional liability.
176
+
177
+ END OF TERMS AND CONDITIONS
178
+
179
+ APPENDIX: How to apply the Apache License to your work.
180
+
181
+ To apply the Apache License to your work, attach the following
182
+ boilerplate notice, with the fields enclosed by brackets "[]"
183
+ replaced with your own identifying information. (Don't include
184
+ the brackets!) The text should be enclosed in the appropriate
185
+ comment syntax for the file format. We also recommend that a
186
+ file or class name and description of purpose be included on the
187
+ same "printed page" as the copyright notice for easier
188
+ identification within third-party archives.
189
+
190
+ Copyright [yyyy] [name of copyright owner]
191
+
192
+ Licensed under the Apache License, Version 2.0 (the "License");
193
+ you may not use this file except in compliance with the License.
194
+ You may obtain a copy of the License at
195
+
196
+ http://www.apache.org/licenses/LICENSE-2.0
197
+
198
+ Unless required by applicable law or agreed to in writing, software
199
+ distributed under the License is distributed on an "AS IS" BASIS,
200
+ WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
201
+ See the License for the specific language governing permissions and
202
+ limitations under the License.
MODEL_PROVENANCE.json ADDED
@@ -0,0 +1,26 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "project": "petitgpt",
3
+ "author": "Yang Qi",
4
+ "checkpoint": "alpha075",
5
+ "lineage": [
6
+ "Base",
7
+ "P2 step750",
8
+ "P3 step320",
9
+ "interpolation alpha=0.75"
10
+ ],
11
+ "excluded_updates": [
12
+ "DeepSeek-response-KD",
13
+ "unified Base-SFT",
14
+ "DPO",
15
+ "soft-KD",
16
+ "LoRA"
17
+ ],
18
+ "source_checkpoint_sha256": "1da85cc329d55e92dacf51c36623779558c4c6c9a39d78a34a064f61fcddbe97",
19
+ "model_safetensors_sha256": "4396efb7a52b047e7fdf513e46d1b401dfc70582d3aca1f9cb5a07e97d426ef1",
20
+ "tokenizer_sha256": "d8f84df58928023edebd809e152b3b38a0dac53b9f887bd2455f427661e9b9ce",
21
+ "accepted_archive_sha256": "0c844962c6ebfcc1cd6b17cdb19c05d88936fa280f0173c080ff864bb54dca1e",
22
+ "unique_parameters": 124635456,
23
+ "release_model_operations": 0,
24
+ "license": "Apache-2.0 only for author-controlled rights; see notices",
25
+ "tokenizer_source_linkage": "Digest-linked to six pinned corpus releases; PES2O and StackExchange not in tokenizer corpus"
26
+ }
README.md ADDED
@@ -0,0 +1,123 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ tags:
4
+ - petitgpt
5
+ - native-pytorch
6
+ - research
7
+ ---
8
+
9
+ # petitgpt
10
+
11
+ Author: Yang Qi. Selected checkpoint: alpha075.
12
+
13
+ ## Identity
14
+
15
+ | Field | Value |
16
+ |---|---|
17
+ | Status | accepted native research inference artifact |
18
+ | Unique parameters | 124,635,456 (124.6M) |
19
+ | Layers / width / FFN | 30 / 576 / 1536 |
20
+ | Attention | 9 query heads, 3 key/value heads (GQA), head dim 64 |
21
+ | Vocabulary / context | 32,000 / 2,048 |
22
+ | Embeddings | tied input/output |
23
+ | Normalization | RMSNorm, epsilon 1e-6 |
24
+ | Positions | RoPE, theta 10000, full head rotation |
25
+ | Dropout | 0.0 |
26
+ | Stored weights | FP32 |
27
+
28
+ Complete checkpoint-derived settings ship in the bundle's `config.json`. The three core modules (model, chat template, token contract) are byte-identical to the project originals. Checkpoint and archive hashes are in MODEL_PROVENANCE.json.
29
+
30
+ ## Provenance
31
+
32
+ The selected weights are a **parameter interpolation**, not a training step:
33
+
34
+ > `theta = theta_P2_step750 + 0.75 · (theta_P3_step320 − theta_P2_step750)`
35
+
36
+ Ancestry: tokenizer release → Stage A pretraining (steps 0–38,146) → Stage B continued pretraining (steps 38,146–49,590, exact full-state resume, accepted as Base) → P2 concise-instruction SFT (750 updates, weights-only initialization from Base) → P3 basic-instruction adaptation (parent B is **step 320**, not the step-640 endpoint) → this interpolation, executed with **zero optimizer updates and zero backward passes** → a numerically unchanged FP32 export.
37
+
38
+ This checkpoint **does not contain** later DeepSeek-response-KD, unified Base-SFT, DPO, soft-KD or LoRA branch updates. Several later branches were initialized *from* it (a one-pass behaviour mix, a DPO pilot, a chosen-answer CE control, and two response-distillation runs); others were not — a unified SFT curve started from the accepted pretrained Base, a loss-allocation A/B split from that curve's step 403, and a shared-tokenizer soft-KD lab ran entirely on external models. A preference-data build used this model's generations but produced no checkpoint. Branch exposures must not be summed into this model's training history.
39
+
40
+ ## Training data
41
+
42
+ Pretraining consumed 13,000,005,634 retained packed tokens over 13,755,731 documents, of which the optimizer stepped over 12,999,720,960 model-input positions, one exposure per block, with no replay of the executed token-position traversal.
43
+
44
+ **Stage A (10,000,003,234 selected serialized tokens, 4 sources):** FineWeb-Edu dedup 71.11%, DCLM-Edu 20.32%, Wikipedia (FineWiki EN) 5.08%, Python-Edu 3.50%.
45
+
46
+ **Stage B (3,000,004,240 selected serialized tokens, 7 sources):** FineWeb-Edu dedup 40.10%, DCLM-Edu 22.92%, structured tutorial content 11.46%, Python-Edu 8.33%, Wikipedia 5.73%, PES2O 5.73%, StackExchange 5.73%.
47
+
48
+ Upstream datasets, pinned revisions and the licence string recorded at each pinned revision are in `tables/PRETRAIN_SOURCE_MIXTURE.csv`. Those recorded strings are evidence of what the builder captured at that revision; they are not a legal determination, are not asserted to be today's terms, and do not by themselves determine the licence of trained weights.
49
+
50
+ Post-training used seven instruction subsets. These are values of a row-level `source` column inside one pinned collection, `HuggingFaceTB/smol-smoltalk` at revision `f73fe857d519ff6ac5af2ea67c4d3834da7b8bcc`, config `default`, train split, established by digest joins through the project's own census and cleanup records. The publisher's card at that pinned revision carries a flat Apache-2.0 badge; its parent collection limits that grant to four newly generated subsets and refers readers to the original dataset for each incorporated public dataset. Four of the seven labels correspond to the newly generated subsets. Of the three incorporated components, one declares `apache-2.0`, one declares `odc-by`, and one declares no licence in its card metadata. Component notices were read from current publisher pages, not from revisions contemporaneous with this training run, and **no component revision is established**. No licence determination is made here.
51
+
52
+ ## Intended use, and use it is not intended for
53
+
54
+ **Intended:** research and engineering study of a small from-scratch language model — reproducing the recorded measurements, inspecting the pipeline, and analysing failure modes.
55
+
56
+ **Not intended:** a general assistant, a production system, anything safety- or correctness-certified, or a source of factual answers. Do not execute code it generates without independent review.
57
+
58
+ **Not evaluated at all:** long-context work, multilingual behaviour, tool use, extended multi-turn dialogue, safety and refusal behaviour, factual currency, retrieval, and any public generative benchmark.
59
+
60
+ ## Evaluation — public multiple-choice likelihood
61
+
62
+ Frozen FP32 results, copied byte-identically from the accepted measurement and **not recomputed**:
63
+
64
+ | Dataset | Split / documents | acc | acc_norm |
65
+ |---|---|---|---|
66
+ | ARC-Easy | test / 2,376 | 1372/2376 = 0.5774410774410774 | 1244/2376 = 0.5235690235690236 |
67
+ | PIQA | validation / 1,838 | 1167/1838 = 0.6349292709466812 | 1145/1838 = 0.6229597388465724 |
68
+
69
+ Protocol: zero-shot raw `Question: …\nAnswer:` completion scored by candidate-answer likelihood. No chat template, no role tokens, no few-shot examples, no BOS insertion, no scored EOS, no generation, no cleanup. FP32 parameters and forward with autocast disabled, TF32 off for matmul and cuDNN, MATH SDPA, batch size 1 unpadded, no KV cache, no compile. `acc` is the first argmax of summed continuation log-likelihood; `acc_norm` divides by `len()` of the **original** answer text in Unicode characters, not tokenizer length; ties take the first index.
70
+
71
+ Comparators measured under the identical protocol on the same rows: SmolLM-135M-Instruct 0.4924 / 0.6708 and SmolLM2-135M-Instruct 0.5400 / 0.6670 (acc, ARC-Easy / PIQA).
72
+
73
+ **Qualifications.** The evaluator is a native protocol-compatible implementation pinned to a specific lm-evaluation-harness commit; the harness package was not installed and a full installed-harness run is not claimed. Both datasets are prior project diagnostics with no contamination audit — they are **not untouched final tests**. Training data, compute, tokenizers and architectures are unmatched across the three models. This is a protocol-bounded descriptive comparison; no significance test was run. Multiple-choice accuracy does not establish free-generation reliability.
74
+
75
+ ## Evaluation — historical full-answer assistant review
76
+
77
+ A separate evaluation family scored generated text across 189 prompts per model (567 answers, 565 distinct prompt/output/contract units) over three models. **Two named versions exist and must not be combined in one table:** `assistant_review_fable_v1` (the original 567 final records) and `assistant_owner_clarification_4_v1` (four explicit final-content decisions, every other axis preserved). The version shown below is `assistant_owner_clarification_4_v1`.
78
+
79
+ | Slice | Content / joint (true / false / unknown) | Other axes |
80
+ |---|---|---|
81
+ | old_qa41 | 6 / 33 / 2 | format 0/1 |
82
+ | old_practical38 | 10 / 26 / 2 | explicit format 20/21 |
83
+ | old_python14 | 0 / 14 / 0 | interface 11/14; finite execution `not_recorded` |
84
+ | new_natural64 | 3 / 61 / 0 | format 7/18 |
85
+ | new_python32 | 0 / 32 / 0 | interface 31/32 |
86
+
87
+ Practical joint bounds: **26/81 .. 10/27** (= 52/162 .. 60/162). This denominator is **27 equally weighted dialogue groups over 38 scored turns** — rows are averaged inside a group first — not an ordinary row average. These are exact unknown-retention bounds, **not confidence intervals**.
88
+
89
+ **Qualifications.** Judgments are model-assisted, not human adjudication. The chronology was: an initial pass over the 565 units with model metadata masked and its own recorded limitations; then a pass with the mapping visible that produced 17 consistency edits; then four owner clarifications forming a separate version. It was therefore neither strictly blinded throughout nor fully label-visible throughout. Development sets were reused across many runs. `not_recorded` means the field is unavailable in this imported view — it does not establish that no historical function test was ever run. Unsupported-builtin results stay unknown and are never converted into demonstrated failures.
90
+
91
+ In these specific Python diagnostics the model produced a correct function **interface** in 42 of 46 prompts and a correct **whole answer** in 0 of 46. This does not establish that it can never write correct code.
92
+
93
+ Ordinary QA and complete natural-task generation remained limited across the evaluated configurations. Individual results differ by suite and label version and are reported with their source rather than reduced to a cross-version maximum; see the technical report §10.1.
94
+
95
+ ## Export parity
96
+
97
+ Eight frozen fixture pairs (four `bf16_native`, four `fp32_math`) matched exactly on prompt IDs, boundaries, full-shape logits, greedy output IDs and stop reason, with a maximum absolute logit difference of **0 within each profile**. All 213 named state entries and 60 non-persistent rotary buffers matched after strict load and safetensors reload.
98
+
99
+ This is a **numerical parity check between source and export under the same profile**. It is not a quality test, not a semantic evaluation, and it does not assert that the two profiles agree with each other. Local import closure was demonstrated once in a fresh isolated process on the measured environment; that is **not** a clean-machine, CPU, cross-hardware or fresh-installation test.
100
+
101
+ ## Format support
102
+
103
+ Native PyTorch CUDA inference only. **No** Transformers `AutoModel`, GGUF, ONNX, vLLM or llama.cpp compatibility is implemented or tested. `special_tokens_map.json` is descriptive native metadata. A CUDA GPU is required.
104
+
105
+ ## Licence and distribution
106
+
107
+ Copyright 2026 Yang Qi. Owner-controlled code and the selected model/tokenizer are licensed under the standard Apache License 2.0 in LICENSE, only for rights Yang Qi is entitled to grant.
108
+
109
+ As an explicit exception to the root code licence, author-written reports and documentation (including README, native run guide and versioned report) are licensed under Creative Commons Attribution 4.0 International (CC BY 4.0): https://creativecommons.org/licenses/by/4.0/ and https://creativecommons.org/licenses/by/4.0/legalcode.en . Attribute Yang Qi and petitgpt, link the licence, and indicate changes. Existing third-party content and notices retain their applicable terms; they are not relicensed.
110
+
111
+ Research describes intended use; it adds no noncommercial or research-only restriction to Apache-licensed artifacts. This grant was approved by the owner through the explicit research-release execution instruction. Earlier PENDING_OWNER_DECISION records remain historical evidence; this does not claim an earlier licence choice.
112
+
113
+ Source metadata is not rights clearance. Weights are not the raw corpus; neither automatic inheritance nor automatic non-application of all dataset terms is asserted. No infringement guarantee or legal certification is given. A disclaimer does not replace applicable permission.
114
+
115
+ ## Recorded runtime
116
+
117
+ Python 3.10.12, torch 2.11.0+cu126, numpy 2.2.6, tokenizers 0.22.2, safetensors 0.8.0, NVIDIA GeForce RTX 4090. Matching versions do not guarantee bit-identical results on other hardware or untested software; recorded driver versions differ across project phases, which is a recorded difference rather than a resolved equivalence.
118
+
119
+ ## Files and report
120
+
121
+ See [native run guide](RUN_GUIDE.md), [source notice](SOURCE_NOTICE.md), [third-party notices](THIRD_PARTY_NOTICES.md) and [file manifest](SHA256SUMS).
122
+
123
+ GitHub project: https://github.com/yangqi0/petitgpt . The complete research-v1 report is staged for `docs/petitgpt-v1/TECHNICAL_REPORT.md` there; publication is blocked by missing GitHub authentication. That report is not yet publicly available, and no live report link is claimed.
RUN_GUIDE.md ADDED
@@ -0,0 +1,125 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # petitgpt native run guide
2
+
3
+ Author: Yang Qi. Documentation: CC BY 4.0. Download the loose model repository files together, preserving the src/ directory. The exact release file list and hashes are in SHA256SUMS. Model weights, tokenizer, config and executable code retain the accepted export bytes.
4
+
5
+ ## 2. Requirements
6
+
7
+ A **CUDA GPU is required** by this CLI. The tested configuration is Python 3.10.12 with:
8
+
9
+ ```
10
+ torch==2.11.0+cu126
11
+ numpy==2.2.6
12
+ tokenizers==0.22.2
13
+ safetensors==0.8.0
14
+ ```
15
+
16
+ on an NVIDIA GeForce RTX 4090 (driver 580.178.04). Prepare the environment separately, from locally supplied wheels:
17
+
18
+ ```sh
19
+ python -m pip install --no-index --find-links /path/to/local/wheelhouse \
20
+ -r /path/to/bundle/requirements-inference-tested.txt
21
+ ```
22
+
23
+ No packages were installed during the export itself. Matching versions do not guarantee bit-identical results on arbitrary hardware or untested software.
24
+
25
+ ## 3. Command line
26
+
27
+ Replace `/path/to/bundle` with the real extracted location.
28
+
29
+ ```sh
30
+ python /path/to/bundle/inference.py \
31
+ --model-directory /path/to/bundle \
32
+ --prompt "Say hello in one sentence." \
33
+ --profile bf16_native \
34
+ --max-new-tokens 32
35
+
36
+ python /path/to/bundle/inference.py \
37
+ --model-directory /path/to/bundle \
38
+ --messages-json /path/to/messages.json \
39
+ --profile fp32_math \
40
+ --max-new-tokens 32
41
+ ```
42
+
43
+ A synthetic `messages.json`:
44
+
45
+ ```json
46
+ [{"role":"system","content":"Use short sentences."},
47
+ {"role":"user","content":"My name is Lin."},
48
+ {"role":"assistant","content":"Hello, Lin."},
49
+ {"role":"user","content":"What name did I give you?"}]
50
+ ```
51
+
52
+ > **Illustrative and unexecuted.** The greeting above is an authored example showing command syntax only. It was written for this guide, was not run, and is not a demonstration of model quality. Do not put frozen evaluation prompts or their outputs into a public demo.
53
+
54
+ ## 4. Python API
55
+
56
+ With the extracted bundle directory on `sys.path`:
57
+
58
+ ```python
59
+ from inference import load_bundle, generate
60
+
61
+ model, tokenizer = load_bundle("/path/to/bundle")
62
+ model = model.to("cuda").eval()
63
+ result = generate(
64
+ model, tokenizer,
65
+ [{"role": "user", "content": "Say hello."}],
66
+ cap=32,
67
+ profile="fp32_math",
68
+ )
69
+ print(result["output_text_raw_including_terminal_eos"])
70
+ ```
71
+
72
+ ## 5. Input contract
73
+
74
+ `messages` must be a JSON array of objects with **exactly** the fields `role` and `content`. A conversation is an optional initial system turn followed by alternating user/assistant turns, **ending in user**. A plain `--prompt` becomes one user message.
75
+
76
+ Rejected, by design, with an explicit error rather than a repair:
77
+
78
+ | Case | Recorded rejection |
79
+ |---|---|
80
+ | Missing `content` | `Each message must contain only role and content` |
81
+ | Unsupported role (e.g. `tool`) | `message 0 has invalid role 'tool'` |
82
+ | Any extra message field | `Each message must contain only role and content` |
83
+ | Conversation ending in `assistant` | `chat must end with a non-empty user turn; got 'assistant'` |
84
+ | Prompt + budget over 2,048 | `Context overflow: prompt plus token budget exceeds 2048` |
85
+ | `max_new_tokens` outside 1..384 | `max_new_tokens must be 1..384` |
86
+
87
+ `default_system=None`: a supplied system turn and full history are retained, with **no injected default and no text normalization**. Literal special-token spellings inside content — for example the literal spelling `[EOS]` — are encoded as ordinary text and cannot inject control IDs.
88
+
89
+ Native token structure:
90
+
91
+ ```
92
+ [BOS] <|system|> system <|user|> user <|assistant|> assistant [EOS] … <|user|> user <|assistant|>
93
+ ```
94
+
95
+ The system segment is omitted when absent. `[BOS]` occurs once; `[EOS]` closes completed assistant turns; no duplicate assistant prefix is added. IDs are fixed: `[PAD]=0`, `[UNK]=1`, `[BOS]=2`, `[EOS]=3`, `<|system|>=4`, `<|user|>=5`, `<|assistant|>=6`.
96
+
97
+ ## 6. Precision profiles
98
+
99
+ Both profiles store FP32 parameters. They differ only in the forward numerical path, and **they are not asserted to agree with each other**.
100
+
101
+ | | `bf16_native` | `fp32_math` |
102
+ |---|---|---|
103
+ | Forward | CUDA BF16 autocast | FP32, autocast off |
104
+ | matmul TF32 | off | off |
105
+ | cuDNN TF32 | **on** | off |
106
+ | `float32_matmul_precision` | highest | highest |
107
+ | SDPA backend | native backends enabled | MATH |
108
+
109
+ ## 7. Decoding and stopping
110
+
111
+ Greedy only: `temperature=0`, `top_k=0`, `top_p=1`, `EOS=3`. Default `max_new_tokens` is **384**; an explicit lower integer in `1..384` is supported. Prompt length plus budget must fit 2,048 — there is no cropping.
112
+
113
+ Output JSON preserves the generated token IDs, the raw text **including the terminal EOS**, and a stop reason of `eos` or `max_new_tokens`. There is no answer cleanup, no fact fixing, no best-of-N, and no retry. A token cap can truncate an answer mid-sentence; that is the recorded behaviour, not a defect.
114
+
115
+ ## 8. Format support and limits
116
+
117
+ Native PyTorch CUDA inference only. There is **no implemented or tested** Transformers `AutoModel`, GGUF, ONNX, vLLM or llama.cpp path. `special_tokens_map.json` is descriptive native metadata, not a loader contract.
118
+
119
+ Local import closure was demonstrated once, in a fresh process with a temporary working directory and an empty `PYTHONPATH`, under an audit hook that denied network access and out-of-bundle repository access; no denied access occurred. That establishes closure **on the measured environment only**. It is not a clean-machine test, not a CPU test, not a cross-hardware test, and not a fresh dependency-installation test.
120
+
121
+ The recorded export parity is likewise **profile-specific**: within each of the two profiles the export reproduced its source exactly, and no claim is made that the two profiles agree with each other, nor that any Transformers, GGUF, ONNX, vLLM or llama.cpp path exists.
122
+
123
+ ## 9. Before you use output
124
+
125
+ This is a research artifact, not a safety- or correctness-certified assistant. In the recorded Python diagnostics the model produced a valid function interface far more often than a correct whole answer. **Do not execute generated code without independent review.** Backend details and the bounded parity evidence live in a separate private evidence archive and are not part of this bundle.
SHA256SUMS ADDED
@@ -0,0 +1,21 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ 74322f799773266d15acf47c40165d2c7e5bf9df8e5ca7a076dd98e51ee41656 DOCUMENTATION_LICENSE.md
2
+ cfc7749b96f63bd31c3c42b5c471bf756814053e847c10f3eb003417bc523d30 LICENSE
3
+ 936b54b619edeb38bc5ff13e01e4bb1682f7d5c6f8fdf78756cc4c4c9f953aaf MODEL_PROVENANCE.json
4
+ d6ce84427abcf0f0a3d1d45475c5282c6e5bd9b5b88d9055ac502698cee66605 README.md
5
+ b6e7c48e4c287517c1c4e0013ed0f8b33f49cae479d185fa7af734e057ab28a4 RUN_GUIDE.md
6
+ 46dc9f7c632aa5b7ffeb5678d76dbe03a9f78b70a8d721f8a3d96dda4a0f5876 SOURCE_NOTICE.md
7
+ aeb343a8a11ef0bffa736b360e73273d5a05f48429ae8346a97517155b75c3ef THIRD_PARTY_NOTICES.md
8
+ fe6cd464a09a0605f6ead38867c0497edc71de31fa45446f790e3f9a4f937369 config.json
9
+ 17e307631b8e5fd8be8c9d7e992632a6dfbc26708287318fca1259f6d1dd548d inference.py
10
+ 4396efb7a52b047e7fdf513e46d1b401dfc70582d3aca1f9cb5a07e97d426ef1 model.safetensors
11
+ dd6b7d954f0153cc6898a9aa0f45901e3727906172e0045d9bc776daf2f7163e requirements-inference-tested.txt
12
+ 79bc49c15e40cfc59779b5559d452c058a06f684a2867d97c3c0190933d8801f special_tokens_map.json
13
+ e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855 src/__init__.py
14
+ 5c4214f0d2a985ee7b24da431ce1bf521c96ae6f7dfa816ef800b505eb68013c src/accepted_generate.py
15
+ c25068dcb7d3a907dba7b7405d53a2962d91886abcaa25b0d5832a82e8bc821b src/chat_template.py
16
+ 2bc9fa8ae16636837c4a2937301a2419d0ac92faa2cc27560dacbd29a5144dc2 src/model.py
17
+ f767b864d7c8e0cb5e2c166c4f019f3f14666dcbf2b944c7909db050e4cf1e96 src/special_tokens.py
18
+ 4dd0077babf854389d09623c154844b46f977ca64753f06ecf50662297a2caf5 tables/ASSISTANT_RESULTS_VERSIONED.csv
19
+ b452952e2762f6fe07484e995a7d590f8659c5d5cf60d13e1785cd975a48ace9 tables/PRETRAIN_SOURCE_MIXTURE.csv
20
+ 5350ff2d0b30e67645d3b088974f893a9943d9c5f7407f45548917f1c54da297 tables/PUBLIC_BENCHMARK_RESULTS.csv
21
+ d8f84df58928023edebd809e152b3b38a0dac53b9f887bd2455f427661e9b9ce tokenizer.json
SOURCE_NOTICE.md ADDED
@@ -0,0 +1,39 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Source notice
2
+
3
+ petitgpt by Yang Qi; selected checkpoint alpha075. Parameter lineage: Base -> P2 step750 -> P3 step320 -> interpolation (0.75 toward P3). Later DeepSeek-response-KD, unified Base-SFT, DPO, soft-KD and LoRA updates are absent.
4
+
5
+ Pretraining source and tokenizer linkage have been established by the existing digest records. The tokenizer corpus uses six of the eight frozen releases; PES2O and StackExchange occur in pretraining, not tokenizer training. Selected source counts are not per-source consumed-token measurements. No new corpus audit is claimed.
6
+
7
+ ## Pretraining and tokenizer sources
8
+
9
+ The following terms are builder-recorded metadata at pinned revisions, not independent verification of every licensor's rights. Dataset authors and contributors retain their rights. Exact mixture and transport revisions are preserved in the companion PRETRAIN_SOURCE_MIXTURE.csv.
10
+
11
+ | Dataset / config | Pinned revision | Recorded terms |
12
+ |---|---|---|
13
+ | HuggingFaceTB/smollm-corpus / fineweb-edu-dedup | `3ba9d605774198c5868892d7a8deda78031a781f` | odc-by-1.0 |
14
+ | HuggingFaceTB/dclm-edu / default | `dbad8ad71224482740cd9c9d353591adbf62fe04` | cc-by-4.0 |
15
+ | HuggingFaceFW/finewiki / en | `8bd13e72e6a002407649b3e898535f42ceb1aeb9` | cc-by-sa-4.0 |
16
+ | common-pile/stackv2_edu_filtered / default | `c354dbe88469a1153e97c6a63ac50591849654de` | per-record metadata.license (Software Heritage permissive subset) |
17
+ | HuggingFaceTB/smollm-corpus + HuggingFaceFW/finephrase / cosmopedia-v2 + tutorial | `3ba9d605774198c5868892d7a8deda78031a781f + 78cf4a5ed0099214979c094c963e699c19163838` | odc-by-1.0 (both) |
18
+ | allenai/dolmino-mix-1124 / pes2o | `a319f19eef1e257417b11ea8c30da266ae175557` | odc-by-1.0 |
19
+ | allenai/dolmino-mix-1124 / stackexchange | `a319f19eef1e257417b11ea8c30da266ae175557` | cc-by-sa |
20
+
21
+ Publisher datasets: https://huggingface.co/datasets/HuggingFaceTB/smollm-corpus ; https://huggingface.co/datasets/HuggingFaceTB/dclm-edu ; https://huggingface.co/datasets/HuggingFaceFW/finewiki ; https://huggingface.co/datasets/common-pile/stackv2_edu_filtered ; https://huggingface.co/datasets/HuggingFaceFW/finephrase ; https://huggingface.co/datasets/allenai/dolmino-mix-1124 .
22
+
23
+ ## Instruction components
24
+
25
+ The collection is HuggingFaceTB/smol-smoltalk, revision `f73fe857d519ff6ac5af2ea67c4d3834da7b8bcc`, default/train. Its pinned card has an Apache-2.0 badge. P2's seven labels are source-column values, not separate repositories. Parent/component correspondence is documentary, not a per-row join; exact component revisions remain unestablished. Parent and component notices below are CURRENT_ONLY observations retrieved on 2026-09-10, not terms proven contemporaneous with training.
26
+
27
+ - **openhermes-50k**: teknium/OpenHermes-2.5. Recorded metadata/grant: no licence in component-card metadata. component dataset card metadata at head b82037821055c377bed0d495e72e46de3bc72e84 (retrieved 2026-09-10T17:53:40Z)
28
+ - **smol-contraints**: Smol-contraints. Recorded metadata/grant: Apache-2.0. parent collection card at head 5feaf2fd3ffca7c237fc38d1861bc30365d48ffa (retrieved 2026-09-10T17:52:48Z): "All the new datasets (Smol-Magpie-Ultra, Smol-contraints, Smol-rewrite, Smol-summarize) are licensed under Apache 2.0."
29
+ - **smollm-rewrite-30k**: Smol-rewrite. Recorded metadata/grant: Apache-2.0. parent collection card at head 5feaf2fd3ffca7c237fc38d1861bc30365d48ffa (retrieved 2026-09-10T17:52:48Z): "All the new datasets (Smol-Magpie-Ultra, Smol-contraints, Smol-rewrite, Smol-summarize) are licensed under Apache 2.0."
30
+ - **smol-magpie-ultra-short**: Smol-Magpie-Ultra. Recorded metadata/grant: Apache-2.0. parent collection card at head 5feaf2fd3ffca7c237fc38d1861bc30365d48ffa (retrieved 2026-09-10T17:52:48Z): "All the new datasets (Smol-Magpie-Ultra, Smol-contraints, Smol-rewrite, Smol-summarize) are licensed under Apache 2.0."
31
+ - **self-oss-instruct**: bigcode/self-oss-instruct-sc2-exec-filter-50k. Recorded metadata/grant: odc-by. component dataset card metadata at head 356bb069eee815daa6e23e9a282eeefe1490ad44 (retrieved 2026-09-10T17:53:40Z)
32
+ - **smol-summarize-20k**: Smol-summarize. Recorded metadata/grant: Apache-2.0. parent collection card at head 5feaf2fd3ffca7c237fc38d1861bc30365d48ffa (retrieved 2026-09-10T17:52:48Z): "All the new datasets (Smol-Magpie-Ultra, Smol-contraints, Smol-rewrite, Smol-summarize) are licensed under Apache 2.0."
33
+ - **everyday-conversations**: HuggingFaceTB/everyday-conversations-llama3.1-2k. Recorded metadata/grant: apache-2.0. component dataset card metadata at head 14f543216b9ba42b6b951dc5bd199460d193b162 (retrieved 2026-09-10T17:53:41Z)
34
+
35
+ The parent publisher limits its Apache grant to its newly generated subsets and points to component-specific terms for incorporated datasets. The self-oss-instruct ODC-By metadata differs from the collection badge; a declaration difference is not proof of legal incompatibility. OpenHermes-2.5 refers to component-specific licences; its empty card licence metadata does not resolve those terms, and exact upstream-component rows for this subset remain unresolved. Ordinary historical unknowns have not been newly resolved.
36
+
37
+ Sources: https://huggingface.co/datasets/HuggingFaceTB/smol-smoltalk/tree/f73fe857d519ff6ac5af2ea67c4d3834da7b8bcc ; https://huggingface.co/datasets/HuggingFaceTB/smoltalk ; https://huggingface.co/datasets/teknium/OpenHermes-2.5 ; https://huggingface.co/datasets/bigcode/self-oss-instruct-sc2-exec-filter-50k ; https://huggingface.co/datasets/HuggingFaceTB/everyday-conversations-llama3.1-2k .
38
+
39
+ No raw corpus, frozen evaluation prompts or model answers are distributed. Author licensing does not relicense upstream datasets or clear third-party rights. See DOCUMENTATION_LICENSE.md for the scoped grant.
THIRD_PARTY_NOTICES.md ADDED
@@ -0,0 +1,5 @@
 
 
 
 
 
 
1
+ # Third-party notices
2
+
3
+ Existing third-party code, attribution headers and notices retain their own applicable terms; the root Apache-2.0 licence grants only rights controlled by Yang Qi. The native code is copied byte-for-byte from the accepted export, with its existing comments preserved. Dependencies are declared, not vendored: PyTorch, NumPy, Hugging Face Tokenizers and Safetensors retain their upstream licences and notices. This release makes no blanket claim about all existing repository code.
4
+
5
+ Dataset authors and contributors are acknowledged in SOURCE_NOTICE.md, including ODC-By components, Wikipedia/FineWiki and StackExchange share-alike metadata, per-record Python licence metadata, and unresolved OpenHermes component terms. Dataset terms are not replaced by the model's Apache-2.0 declaration.
config.json ADDED
@@ -0,0 +1,13 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "vocab_size": 32000,
3
+ "n_layers": 30,
4
+ "d_model": 576,
5
+ "n_heads": 9,
6
+ "n_kv_heads": 3,
7
+ "d_ff": 1536,
8
+ "max_seq_len": 2048,
9
+ "dropout": 0.0,
10
+ "tie_embeddings": true,
11
+ "rope_theta": 10000.0,
12
+ "rope_pct": 1.0
13
+ }
inference.py ADDED
@@ -0,0 +1,63 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Local native alpha075 loading and precision profiles."""
2
+ import argparse,json
3
+ from pathlib import Path
4
+ from contextlib import contextmanager
5
+ import torch
6
+ from torch.nn.attention import sdpa_kernel,SDPBackend
7
+ from safetensors.torch import load_model
8
+ from src.model import GPT,gpt_config_from_checkpoint_dict,audit_gpt_parameter_count
9
+ from src.chat_template import load_chat_tokenizer,encode_prompt
10
+ from src.accepted_generate import generate as accepted_generate
11
+
12
+ @contextmanager
13
+ def precision(profile):
14
+ if profile not in ('bf16_native','fp32_math'): raise ValueError('Unsupported precision profile')
15
+ torch.set_float32_matmul_precision('highest')
16
+ torch.backends.cuda.matmul.allow_tf32=False
17
+ torch.backends.cudnn.allow_tf32=profile=='bf16_native'
18
+ with sdpa_kernel([SDPBackend.MATH] if profile=='fp32_math' else [SDPBackend.FLASH_ATTENTION,SDPBackend.EFFICIENT_ATTENTION,SDPBackend.MATH,SDPBackend.CUDNN_ATTENTION]):
19
+ with torch.autocast('cuda',dtype=torch.bfloat16,enabled=profile=='bf16_native'):
20
+ yield
21
+
22
+ def load_bundle(model_dir):
23
+ root=Path(model_dir).resolve()
24
+ cfg=gpt_config_from_checkpoint_dict(json.loads((root/'config.json').read_text()))
25
+ model=GPT(cfg).eval()
26
+ load_model(model,str(root/'model.safetensors'),strict=True,device='cpu')
27
+ audit=audit_gpt_parameter_count(model,cfg)
28
+ if audit['actual_total']!=124635456: raise ValueError('Parameter identity guard failed')
29
+ if model.tok_emb.weight is not model.lm_head.weight: raise ValueError('Embedding tie lost')
30
+ if model.tok_emb.weight.data_ptr()!=model.lm_head.weight.data_ptr(): raise ValueError('Storage tie lost')
31
+ if any(p.dtype!=torch.float32 for p in model.parameters()): raise ValueError('Expected FP32 weights')
32
+ return model,load_chat_tokenizer(str(root/'tokenizer.json'))
33
+
34
+ def validate_input(tok,messages,cap):
35
+ if not isinstance(messages,list) or not messages: raise ValueError('Messages must be a non-empty JSON array')
36
+ if any(not isinstance(m,dict) or set(m)!={'role','content'} for m in messages): raise ValueError('Each message must contain only role and content')
37
+ if isinstance(cap,bool) or not isinstance(cap,int) or not 1<=cap<=384: raise ValueError('max_new_tokens must be 1..384')
38
+ ids=encode_prompt(tok,messages,default_system=None,mode='full_context')
39
+ if len(ids)+cap>2048: raise ValueError('Context overflow: prompt plus token budget exceeds 2048')
40
+ return ids
41
+
42
+ @torch.inference_mode()
43
+ def generate(model,tok,messages,cap=384,profile='bf16_native'):
44
+ validate_input(tok,messages,cap)
45
+ model.eval()
46
+ with precision(profile):
47
+ return accepted_generate(model,tok,messages,cap,autocast_enabled=profile=='bf16_native')
48
+
49
+ def main():
50
+ p=argparse.ArgumentParser(description=__doc__)
51
+ p.add_argument('--model-directory',required=True,type=Path)
52
+ g=p.add_mutually_exclusive_group(required=True)
53
+ g.add_argument('--messages-json',type=Path);g.add_argument('--prompt')
54
+ p.add_argument('--profile',choices=['bf16_native','fp32_math'],default='bf16_native')
55
+ p.add_argument('--max-new-tokens',type=int,default=384)
56
+ a=p.parse_args()
57
+ messages=json.loads(a.messages_json.read_text()) if a.messages_json else [{'role':'user','content':a.prompt}]
58
+ tok=load_chat_tokenizer(str(a.model_directory/'tokenizer.json'))
59
+ validate_input(tok,messages,a.max_new_tokens)
60
+ model,tok=load_bundle(a.model_directory)
61
+ model=model.to('cuda').eval()
62
+ print(json.dumps(generate(model,tok,messages,a.max_new_tokens,a.profile),ensure_ascii=False))
63
+ if __name__=='__main__': main()
model.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:4396efb7a52b047e7fdf513e46d1b401dfc70582d3aca1f9cb5a07e97d426ef1
3
+ size 498562192
requirements-inference-tested.txt ADDED
@@ -0,0 +1,5 @@
 
 
 
 
 
 
1
+ # Python 3.10.12; CUDA 12.6 PyTorch build
2
+ torch==2.11.0+cu126
3
+ numpy==2.2.6
4
+ tokenizers==0.22.2
5
+ safetensors==0.8.0
special_tokens_map.json ADDED
@@ -0,0 +1,12 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "format": "descriptive native metadata; no HF AutoModel claim",
3
+ "tokens": {
4
+ "[PAD]": 0,
5
+ "[UNK]": 1,
6
+ "[BOS]": 2,
7
+ "[EOS]": 3,
8
+ "<|system|>": 4,
9
+ "<|user|>": 5,
10
+ "<|assistant|>": 6
11
+ }
12
+ }
src/__init__.py ADDED
File without changes
src/accepted_generate.py ADDED
@@ -0,0 +1,18 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ import torch
2
+ from src.chat_template import encode_prompt
3
+
4
+ @torch.inference_mode()
5
+ def generate(model,tok,messages,cap,autocast_enabled=True):
6
+ ids=encode_prompt(tok,messages,default_system=None,mode='full_context')
7
+ assert len(ids)+cap<=2048
8
+ with torch.autocast('cuda',dtype=torch.bfloat16,enabled=autocast_enabled):
9
+ gen=model.generate(torch.tensor([ids],device='cuda'),max_new_tokens=cap,temperature=0,top_k=0,top_p=1,eos_id=3)
10
+ new=gen[0,len(ids):].tolist();eos=bool(new and new[-1]==3);before=new[:-1] if eos else new
11
+ text=tok.decode(before,skip_special_tokens=False)
12
+ grams=[tuple(before[i:i+4]) for i in range(max(0,len(before)-3))]
13
+ return {'prompt_token_ids':ids,'generated_token_ids':new,'output_text_for_scoring':text,
14
+ 'output_text_raw_including_terminal_eos':tok.decode(new,skip_special_tokens=False),
15
+ 'stop_reason':'eos' if eos else 'max_new_tokens','generated_tokens_including_eos':len(new),
16
+ 'empty_output':not text,'repeated_4gram_fraction':1-len(set(grams))/len(grams) if grams else 0.,
17
+ 'unexpected_control_token_count':sum(v in {0,2,4,5,6} for v in before),
18
+ 'answer_cleanup':False,'constrained_decoding':False,'max_new_tokens':cap}
src/chat_template.py ADDED
@@ -0,0 +1,507 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Single source of truth for the chat format, shared by SFT / distill / DPO /
2
+ GRPO training AND their sampling code (previously seven duplicated copies).
3
+
4
+ Token-level template — role boundaries are special tokens, not plain text:
5
+
6
+ [BOS] <|system|> {system} <|user|> {q1} <|assistant|> {a1} [EOS] <|user|> {q2} ...
7
+
8
+ Design rules:
9
+ - Role tokens delimit turns. BPE can never merge across a special token, so a
10
+ conversation encodes to the same ids whether it is built turn-by-turn during
11
+ training or as a generation prompt at inference (the old plain-text
12
+ "User: ...\\n\\n" template drifted at every segment boundary).
13
+ - [EOS] appears ONLY after assistant turns: it means "assistant finished,
14
+ stop generating" — the same stop semantics as document ends in pretraining.
15
+ System/user turns need no terminator; the next role token is the boundary.
16
+ - The supervised span is exactly each assistant turn's content tokens plus its
17
+ trailing [EOS] (so the model is explicitly taught to stop).
18
+ - Content is encoded with `tokenizer.encode_special_tokens = True`
19
+ (see `load_chat_tokenizer`), so literal "[EOS]" / "<|user|>" strings inside
20
+ user or corpus text are tokenized as plain text and can never inject real
21
+ control tokens.
22
+ """
23
+
24
+ from __future__ import annotations
25
+
26
+ from typing import Any
27
+
28
+ from tokenizers import Tokenizer
29
+
30
+ from src.special_tokens import (
31
+ ASSISTANT_ID,
32
+ BOS_ID,
33
+ EOS_ID,
34
+ PAD_ID,
35
+ SYSTEM_ID,
36
+ USER_ID,
37
+ assert_tokenizer_contract,
38
+ )
39
+
40
+ DEFAULT_SYSTEM = "You are a helpful assistant."
41
+
42
+ IGNORE_INDEX = -100
43
+
44
+
45
+ # -------------------------
46
+ # Tokenizer loading
47
+ # -------------------------
48
+ def configure_chat_tokenizer(tok: Tokenizer) -> Tokenizer:
49
+ """Make `tok.encode` treat special-token strings in raw text as plain text.
50
+
51
+ All special tokens in this pipeline are inserted by ID by the code below,
52
+ never parsed out of content — this closes the prompt-injection hole where a
53
+ document containing the literal string "[EOS]" would encode to the real
54
+ EOS id.
55
+ """
56
+ tok.encode_special_tokens = True
57
+ return tok
58
+
59
+
60
+ def load_chat_tokenizer(tokenizer_path: str) -> Tokenizer:
61
+ """Load tokenizer.json, assert the hardcoded special-token IDs, and disable
62
+ special-token matching in raw text. Every chat-stage script should load its
63
+ tokenizer through this."""
64
+ assert_tokenizer_contract(tokenizer_path)
65
+ return configure_chat_tokenizer(Tokenizer.from_file(tokenizer_path))
66
+
67
+
68
+ # -------------------------
69
+ # Text cleaning
70
+ # -------------------------
71
+ def norm_newlines(s: str) -> str:
72
+ return (s or "").replace("\r\n", "\n").replace("\r", "\n")
73
+
74
+
75
+ def clean_text(s: str) -> str:
76
+ # for system/user text: strip leading/trailing whitespace
77
+ return norm_newlines(s).strip()
78
+
79
+
80
+ def clean_text_assistant(s: str) -> str:
81
+ # IMPORTANT: do not strip assistant text (keeps code indentation / markdown
82
+ # formatting); only trailing whitespace goes (EOS follows immediately).
83
+ return norm_newlines(s).rstrip()
84
+
85
+
86
+ def _normalized_messages(
87
+ messages: list[dict[str, str]],
88
+ default_system: str | None,
89
+ *,
90
+ expected_end: str | None,
91
+ ) -> list[dict[str, str]]:
92
+ """Clean and validate the canonical chat state machine.
93
+
94
+ Normally a chat has an initial system turn, followed by alternating user
95
+ and assistant turns. Explicit default_system=None preserves source content
96
+ verbatim and permits a missing system turn without injecting one. ``expected_end`` is ``"user"`` for
97
+ a generation prompt and ``"assistant"`` for a complete training sample.
98
+ No malformed turn is silently skipped.
99
+ """
100
+ out: list[dict[str, str]] = []
101
+ for index, message in enumerate(messages or []):
102
+ if not isinstance(message, dict):
103
+ raise ValueError(f"message {index} must be an object")
104
+ role_value = message.get("role")
105
+ role = role_value.strip().lower() if isinstance(role_value, str) else ""
106
+ if role not in ("system", "user", "assistant"):
107
+ raise ValueError(f"message {index} has invalid role {role_value!r}")
108
+ raw = message.get("content")
109
+ if not isinstance(raw, str):
110
+ raise ValueError(f"message {index} content must be a string")
111
+ text = raw if default_system is None else (
112
+ clean_text_assistant(raw) if role == "assistant" else clean_text(raw)
113
+ )
114
+ if not text.strip():
115
+ if index == 0 and role == "system" and default_system is not None:
116
+ fallback = clean_text(default_system)
117
+ if not fallback:
118
+ raise ValueError(
119
+ "chat requires a non-empty initial system turn or non-empty default_system"
120
+ )
121
+ text = fallback
122
+ else:
123
+ raise ValueError(f"message {index} ({role}) content must be non-empty")
124
+ out.append({"role": role, "content": text})
125
+
126
+ if (not out or out[0]["role"] != "system") and default_system is not None:
127
+ fallback = clean_text(default_system)
128
+ if not fallback:
129
+ raise ValueError(
130
+ "chat requires a non-empty initial system turn or non-empty default_system"
131
+ )
132
+ out.insert(0, {"role": "system", "content": fallback})
133
+
134
+ has_system = bool(out) and out[0]["role"] == "system"
135
+ if not out or (has_system and len(out) == 1):
136
+ raise ValueError("chat requires at least one non-empty user turn")
137
+ for index, message in enumerate(out):
138
+ turn_index = index - int(has_system)
139
+ expected = "system" if has_system and index == 0 else (
140
+ "user" if turn_index % 2 == 0 else "assistant"
141
+ )
142
+ if message["role"] != expected:
143
+ raise ValueError(
144
+ "invalid chat role order: initial system must be followed by "
145
+ f"alternating user/assistant turns (index {index}: expected "
146
+ f"{expected!r}, got {message['role']!r})"
147
+ )
148
+ if expected_end is not None and out[-1]["role"] != expected_end:
149
+ raise ValueError(
150
+ f"chat must end with a non-empty {expected_end} turn; got {out[-1]['role']!r}"
151
+ )
152
+ return out
153
+
154
+
155
+ def prepare_prompt_messages(
156
+ messages: list[dict[str, str]],
157
+ default_system: str | None = DEFAULT_SYSTEM,
158
+ ) -> list[dict[str, str]]:
159
+ """Explicitly turn a valid conversation/example into a USER-ending prompt.
160
+
161
+ Callers sampling from a complete SFT example must opt in to removing its
162
+ final assistant answer. ``encode_prompt`` itself never drops turns.
163
+ """
164
+ normalized = _normalized_messages(messages, default_system, expected_end=None)
165
+ if normalized[-1]["role"] == "assistant":
166
+ normalized = normalized[:-1]
167
+ if normalized[-1]["role"] != "user":
168
+ raise ValueError("generation prompt context must end with a user turn")
169
+ return normalized
170
+
171
+
172
+ # -------------------------
173
+ # Core encoding
174
+ # -------------------------
175
+ _ROLE_TOKEN_ID = {"system": SYSTEM_ID, "user": USER_ID, "assistant": ASSISTANT_ID}
176
+
177
+
178
+ def _encode_normalized_chat(
179
+ tok: Tokenizer, messages: list[dict[str, str]]
180
+ ) -> tuple[list[int], list[int]]:
181
+ ids: list[int] = [BOS_ID]
182
+ labels: list[int] = [IGNORE_INDEX]
183
+ for index, message in enumerate(messages):
184
+ role = message["role"]
185
+ content_ids = tok.encode(message["content"]).ids
186
+ if not content_ids:
187
+ raise ValueError(f"message {index} ({role}) must encode to at least one token")
188
+ ids.append(_ROLE_TOKEN_ID[role])
189
+ labels.append(IGNORE_INDEX)
190
+ if role == "assistant":
191
+ ids.extend(content_ids)
192
+ labels.extend(content_ids)
193
+ ids.append(EOS_ID)
194
+ labels.append(EOS_ID)
195
+ else:
196
+ ids.extend(content_ids)
197
+ labels.extend([IGNORE_INDEX] * len(content_ids))
198
+ return ids, labels
199
+
200
+
201
+ def encode_chat(
202
+ tok: Tokenizer,
203
+ messages: list[dict[str, str]],
204
+ default_system: str | None = DEFAULT_SYSTEM,
205
+ ) -> tuple[list[int], list[int]]:
206
+ """Encode a full conversation for training.
207
+
208
+ Returns (ids, labels), same length. labels[i] == ids[i] on supervised
209
+ positions (every assistant turn's content + its trailing EOS) and
210
+ IGNORE_INDEX everywhere else (BOS, role tokens, system/user content).
211
+ """
212
+ msgs = _normalized_messages(messages, default_system, expected_end="assistant")
213
+ return _encode_normalized_chat(tok, msgs)
214
+
215
+
216
+ def encode_prompt(
217
+ tok: Tokenizer,
218
+ messages: list[dict[str, str]],
219
+ default_system: str | None = DEFAULT_SYSTEM,
220
+ mode: str = "full_context",
221
+ ) -> list[int]:
222
+ """Encode a generation prompt: context ending in the ``<|assistant|>`` cue.
223
+
224
+ mode:
225
+ - "full_context": the entire validated USER-ending context (earlier
226
+ assistant turns keep their [EOS]).
227
+ - "last_user": system turn + last user turn only.
228
+
229
+ The returned ids start with BOS and end with ASSISTANT_ID, exactly matching
230
+ the training-time prefix for an assistant turn.
231
+ """
232
+ msgs = _normalized_messages(messages, default_system, expected_end="user")
233
+
234
+ if mode == "last_user":
235
+ msgs = [msgs[0], msgs[-1]] if msgs[0]["role"] == "system" else [msgs[-1]]
236
+ elif mode != "full_context":
237
+ raise ValueError(f"unknown prompt mode: {mode}")
238
+
239
+ ids, _ = _encode_normalized_chat(tok, msgs)
240
+ ids.append(ASSISTANT_ID)
241
+ return ids
242
+
243
+
244
+ def encode_completion(
245
+ tok: Tokenizer,
246
+ messages: list[dict[str, str]],
247
+ completion: str,
248
+ default_system: str | None = DEFAULT_SYSTEM,
249
+ ) -> tuple[list[int], list[int]]:
250
+ """Encode (prompt context + one assistant completion) for DPO-style scoring.
251
+
252
+ `messages` is the shared context (must end with a user turn); `completion`
253
+ is a plain assistant string. Returns (ids, labels) where the supervised
254
+ span is the completion tokens + trailing EOS — logps therefore include the
255
+ stop decision.
256
+ """
257
+ ids = encode_prompt(tok, messages, default_system, mode="full_context")
258
+ labels = [IGNORE_INDEX] * len(ids)
259
+ if not isinstance(completion, str) or not completion.strip():
260
+ raise ValueError("DPO completion must be a non-empty, non-whitespace string")
261
+ comp_ids = tok.encode(completion if default_system is None else clean_text_assistant(completion)).ids
262
+ if not comp_ids:
263
+ raise ValueError("DPO completion must encode to at least one token")
264
+ ids.extend(comp_ids)
265
+ labels.extend(comp_ids)
266
+ ids.append(EOS_ID)
267
+ labels.append(EOS_ID)
268
+ return ids, labels
269
+
270
+
271
+ def truncate_chat_sequence(
272
+ ids: list[int],
273
+ labels: list[int] | None,
274
+ max_len: int,
275
+ ) -> tuple[list[int], list[int] | None]:
276
+ """Validate and truncate only at complete user-turn boundaries.
277
+
278
+ The prefix is BOS, plus the initial SYSTEM turn when present. It is always
279
+ retained; source-preserving conversations may begin BOS + USER. The remainder is the largest recent suffix that begins at a real
280
+ ``USER_ID`` marker and runs through the original sequence end. A training
281
+ sequence must end with an assistant EOS; a prompt must end with the final
282
+ assistant cue. Literal special-token text cannot create these boundaries
283
+ because chat tokenizers disable special-token matching for raw content.
284
+
285
+ If the system prefix plus the latest complete user-led suffix cannot fit,
286
+ this function raises instead of cutting content, role markers, assistant
287
+ completions, or EOS targets.
288
+ """
289
+ if max_len <= 0:
290
+ raise ValueError("max_len must be positive")
291
+ if labels is not None and len(labels) != len(ids):
292
+ raise ValueError("ids and labels must have identical lengths")
293
+ if len(ids) < 3 or ids[0] != BOS_ID or ids[1] not in {SYSTEM_ID, USER_ID}:
294
+ raise ValueError("chat sequence must start with BOS_ID and a SYSTEM or USER turn")
295
+
296
+ forbidden_content_ids = {PAD_ID, BOS_ID, EOS_ID, SYSTEM_ID, USER_ID, ASSISTANT_ID}
297
+ role_ids = {SYSTEM_ID, USER_ID, ASSISTANT_ID}
298
+ first_role = 1
299
+ if ids[1] == SYSTEM_ID:
300
+ first_role = next((index for index in range(2, len(ids)) if ids[index] in role_ids), len(ids))
301
+ if first_role == 2:
302
+ raise ValueError("initial system content must contain at least one token")
303
+ if any(token_id in forbidden_content_ids for token_id in ids[2:first_role]):
304
+ raise ValueError("system content contains a structural special-token ID")
305
+ if first_role == len(ids) or ids[first_role] != USER_ID:
306
+ raise ValueError("chat sequence must contain a user turn after the initial system turn")
307
+
308
+ user_starts: list[int] = []
309
+ cursor = first_role
310
+ while cursor < len(ids):
311
+ if ids[cursor] != USER_ID:
312
+ raise ValueError("chat role sequence must alternate USER and ASSISTANT")
313
+ user_starts.append(cursor)
314
+ user_content_start = cursor + 1
315
+ assistant_pos = next(
316
+ (
317
+ index
318
+ for index in range(user_content_start, len(ids))
319
+ if ids[index] in forbidden_content_ids
320
+ ),
321
+ len(ids),
322
+ )
323
+ if assistant_pos == user_content_start:
324
+ raise ValueError("user content must contain at least one token")
325
+ if assistant_pos == len(ids) or ids[assistant_pos] != ASSISTANT_ID:
326
+ raise ValueError("each user turn must be followed by an assistant turn")
327
+
328
+ assistant_content_start = assistant_pos + 1
329
+ if assistant_content_start == len(ids):
330
+ if labels is not None:
331
+ raise ValueError("training chat must end with assistant content and EOS")
332
+ cursor = len(ids)
333
+ break
334
+
335
+ eos_pos = next(
336
+ (
337
+ index
338
+ for index in range(assistant_content_start, len(ids))
339
+ if ids[index] in forbidden_content_ids
340
+ ),
341
+ len(ids),
342
+ )
343
+ if eos_pos == assistant_content_start:
344
+ raise ValueError("assistant content must contain at least one token")
345
+ if eos_pos == len(ids) or ids[eos_pos] != EOS_ID:
346
+ raise ValueError("each assistant completion must end with EOS")
347
+ cursor = eos_pos + 1
348
+ if cursor == len(ids):
349
+ if labels is None:
350
+ raise ValueError("generation prompt must end with an assistant cue")
351
+ break
352
+
353
+ if labels is not None:
354
+ if ids[-1] != EOS_ID:
355
+ raise ValueError("training chat must end with a supervised assistant EOS")
356
+ if labels[-1] != EOS_ID:
357
+ raise ValueError("final assistant EOS must be supervised")
358
+
359
+ prefix_end = first_role
360
+ chosen_start = next(
361
+ (start for start in user_starts if prefix_end + (len(ids) - start) <= max_len),
362
+ None,
363
+ )
364
+ if chosen_start is None:
365
+ minimum = prefix_end + (len(ids) - user_starts[-1])
366
+ raise ValueError(
367
+ "chat sequence does not fit without cutting the system prefix or latest "
368
+ f"user-led suffix (requires at least {minimum} tokens, max_len={max_len})"
369
+ )
370
+
371
+ if chosen_start == prefix_end:
372
+ return list(ids), list(labels) if labels is not None else None
373
+
374
+ kept_ids = ids[:prefix_end] + ids[chosen_start:]
375
+ if labels is None:
376
+ return kept_ids, None
377
+ return kept_ids, labels[:prefix_end] + labels[chosen_start:]
378
+
379
+
380
+ def pad_or_truncate(
381
+ ids: list[int],
382
+ labels: list[int],
383
+ seq_len: int,
384
+ pad_id: int = PAD_ID,
385
+ ) -> tuple[list[int], list[int]]:
386
+ """Structure-aware fixed-length shaping plus right padding."""
387
+ ids, maybe_labels = truncate_chat_sequence(ids, labels, seq_len)
388
+ assert maybe_labels is not None
389
+ labels = maybe_labels
390
+ pad_n = seq_len - len(ids)
391
+ return ids + [pad_id] * pad_n, labels + [IGNORE_INDEX] * pad_n
392
+
393
+
394
+ def count_chat_tokens(
395
+ tok: Tokenizer,
396
+ messages: list[dict[str, str]],
397
+ default_system: str | None = DEFAULT_SYSTEM,
398
+ ) -> int:
399
+ """Exact token count of the training encoding (for mix/token budgeting)."""
400
+ ids, _ = encode_chat(tok, messages, default_system)
401
+ return len(ids)
402
+
403
+
404
+ def extract_last_user_and_ref(messages: list[dict[str, str]]) -> tuple[str, str]:
405
+ """Return (last_user_text, last_assistant_text_if_any) — for sample logs."""
406
+ last_user = ""
407
+ ref = ""
408
+ for m in reversed(messages or []):
409
+ if (m.get("role") or "").strip().lower() == "user":
410
+ last_user = clean_text(m.get("content", ""))
411
+ break
412
+ for m in reversed(messages or []):
413
+ if (m.get("role") or "").strip().lower() == "assistant":
414
+ ref = clean_text_assistant(m.get("content", ""))
415
+ break
416
+ return last_user, ref
417
+
418
+
419
+ def decode_completion(tok: Tokenizer, ids: list[int]) -> str:
420
+ """Decode generated completion ids (the tokens AFTER the prompt), dropping
421
+ a trailing EOS if present. Replaces the old fragile rfind('Assistant: ')."""
422
+ if ids and ids[-1] == EOS_ID:
423
+ ids = ids[:-1]
424
+ return tok.decode(ids).strip() if ids else ""
425
+
426
+
427
+ # -------------------------
428
+ # Refusal detection (shared by SFT/distill example weighting)
429
+ # -------------------------
430
+ def is_refusal_text(text: str, patterns: list[str]) -> bool:
431
+ """If assistant content contains any refusal-ish substring, treat as refusal."""
432
+ t = (text or "").strip().lower()
433
+ if not t:
434
+ return False
435
+ for p in patterns:
436
+ p2 = p.strip().lower()
437
+ if p2 and p2 in t:
438
+ return True
439
+ return False
440
+
441
+
442
+ def compute_example_weight_from_messages(
443
+ messages: list[dict[str, str]],
444
+ refusal_downweight: float,
445
+ refusal_patterns: list[str],
446
+ refusal_mode: str,
447
+ ) -> float:
448
+ """Scalar loss weight for a training example; downweights refusal-looking
449
+ assistant turns (refusal_mode="contains_any")."""
450
+ if refusal_downweight >= 1.0:
451
+ return 1.0
452
+ if refusal_downweight <= 0.0:
453
+ return 0.0
454
+ if refusal_mode != "contains_any":
455
+ raise ValueError(f"unknown refusal_mode: {refusal_mode}")
456
+ for m in messages or []:
457
+ if (m.get("role") or "").strip().lower() == "assistant":
458
+ if is_refusal_text(m.get("content", ""), refusal_patterns):
459
+ return refusal_downweight
460
+ return 1.0
461
+
462
+
463
+ def build_example(
464
+ ex: dict[str, Any],
465
+ tok: Tokenizer,
466
+ seq_len: int,
467
+ default_system: str | None,
468
+ refusal_downweight: float,
469
+ refusal_patterns: list[str],
470
+ refusal_mode: str,
471
+ pad_id: int = PAD_ID,
472
+ ):
473
+ """One SFT/distill training example -> (input_ids, labels, example_weight).
474
+
475
+ Tensors are torch.long of length seq_len; labels use IGNORE_INDEX outside
476
+ the supervised assistant spans. Honors meta.bucket safety exemption and
477
+ meta.weight multipliers exactly as before.
478
+ """
479
+ import torch
480
+
481
+ messages = ex.get("messages") or []
482
+ if not messages:
483
+ raise ValueError("missing messages")
484
+
485
+ meta = ex.get("meta") or {}
486
+ bucket = str(meta.get("bucket", "")).strip()
487
+ # Do NOT downweight refusals inside the safety bucket (otherwise safety
488
+ # examples get muted).
489
+ refusal_dw_eff = 1.0 if bucket in ("D_safety", "D") else refusal_downweight
490
+ ex_weight = compute_example_weight_from_messages(
491
+ messages, refusal_dw_eff, refusal_patterns, refusal_mode
492
+ )
493
+ w0 = meta.get("weight", None) if isinstance(meta, dict) else None
494
+ if isinstance(w0, (int, float)):
495
+ ex_weight *= float(w0)
496
+
497
+ ids, labels = encode_chat(tok, messages, default_system)
498
+ if default_system is None and len(ids) > seq_len:
499
+ raise ValueError("source-preserving SFT rejects overlength conversations")
500
+ ids, labels = pad_or_truncate(ids, labels, seq_len, pad_id)
501
+ if all(label == IGNORE_INDEX for label in labels):
502
+ raise ValueError("SFT example requires a retained non-empty assistant turn")
503
+ return (
504
+ torch.tensor(ids, dtype=torch.long),
505
+ torch.tensor(labels, dtype=torch.long),
506
+ float(ex_weight),
507
+ )
src/model.py ADDED
@@ -0,0 +1,470 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ from __future__ import annotations
2
+
3
+ from dataclasses import dataclass
4
+ import math
5
+
6
+ import torch
7
+ import torch.nn as nn
8
+ import torch.nn.functional as F
9
+
10
+
11
+ @dataclass
12
+ class GPTConfig:
13
+ vocab_size: int = 32000
14
+ n_layers: int = 30
15
+ d_model: int = 576
16
+ n_heads: int = 9
17
+ n_kv_heads: int = 3 # GQA KV heads; == n_heads is plain MHA (pre-GQA checkpoints)
18
+ d_ff: int = 1536 # SwiGLU: ~2.67x d_model (MobileLLM/SmolLM2-135M deep-thin shape)
19
+ max_seq_len: int = 2048
20
+ dropout: float = 0.0
21
+ tie_embeddings: bool = True
22
+
23
+ # RoPE (rotary positional embedding)
24
+ rope_theta: float = 10000.0
25
+ rope_pct: float = 1.0 # fraction of head_dim to rotate (1.0 = full head_dim)
26
+
27
+
28
+ CANONICAL_DENSE_PARAMETER_COUNT = 124_635_456
29
+ _CANONICAL_PARAMETERIZATION = {
30
+ "vocab_size": 32_000,
31
+ "n_layers": 30,
32
+ "d_model": 576,
33
+ "n_heads": 9,
34
+ "n_kv_heads": 3,
35
+ "d_ff": 1_536,
36
+ "tie_embeddings": True,
37
+ }
38
+
39
+
40
+ def expected_gpt_parameter_count(cfg: GPTConfig) -> int:
41
+ """Derive the unique parameter count for the dense bias-free GPT."""
42
+ integer_fields = {
43
+ "vocab_size": cfg.vocab_size,
44
+ "n_layers": cfg.n_layers,
45
+ "d_model": cfg.d_model,
46
+ "n_heads": cfg.n_heads,
47
+ "n_kv_heads": cfg.n_kv_heads,
48
+ "d_ff": cfg.d_ff,
49
+ }
50
+ for name, value in integer_fields.items():
51
+ if isinstance(value, bool) or not isinstance(value, int) or value <= 0:
52
+ raise ValueError(f"GPTConfig.{name} must be a positive integer")
53
+ if cfg.d_model % cfg.n_heads:
54
+ raise ValueError("GPTConfig.d_model must be divisible by n_heads")
55
+ if cfg.n_heads % cfg.n_kv_heads:
56
+ raise ValueError("GPTConfig.n_heads must be divisible by n_kv_heads")
57
+
58
+ head_dim = cfg.d_model // cfg.n_heads
59
+ kv_dim = cfg.n_kv_heads * head_dim
60
+ token_matrices = 1 if cfg.tie_embeddings else 2
61
+ embeddings = token_matrices * cfg.vocab_size * cfg.d_model
62
+ # q + output projections are d_model x d_model; k and v are d_model x kv_dim (GQA)
63
+ attention = 2 * cfg.d_model * cfg.d_model + 2 * cfg.d_model * kv_dim
64
+ swiglu = 3 * cfg.d_model * cfg.d_ff
65
+ block_norms = 2 * cfg.d_model
66
+ final_norm = cfg.d_model
67
+ return int(embeddings + cfg.n_layers * (attention + swiglu + block_norms) + final_norm)
68
+
69
+
70
+ def audit_gpt_parameter_count(model: nn.Module, cfg: GPTConfig) -> dict[str, int | bool | str]:
71
+ """Fail fast on implementation/config drift and return manifest metadata."""
72
+ expected = expected_gpt_parameter_count(cfg)
73
+ actual = int(sum(parameter.numel() for parameter in model.parameters()))
74
+ trainable = int(
75
+ sum(parameter.numel() for parameter in model.parameters() if parameter.requires_grad)
76
+ )
77
+ if actual != expected:
78
+ raise RuntimeError(
79
+ "GPT parameter count disagrees with the architecture-derived count: "
80
+ f"actual={actual:,}, expected={expected:,}"
81
+ )
82
+
83
+ canonical = all(
84
+ getattr(cfg, field) == expected_value
85
+ for field, expected_value in _CANONICAL_PARAMETERIZATION.items()
86
+ )
87
+ if canonical and actual != CANONICAL_DENSE_PARAMETER_COUNT:
88
+ raise RuntimeError(
89
+ "canonical PetitGPT parameter count mismatch: "
90
+ f"actual={actual:,}, expected={CANONICAL_DENSE_PARAMETER_COUNT:,}"
91
+ )
92
+
93
+ return {
94
+ "status": "passed",
95
+ "counting_method": "unique_parameter_objects_excluding_buffers",
96
+ "actual_total": actual,
97
+ "actual_trainable": trainable,
98
+ "derived_expected_total": expected,
99
+ "canonical_parameterization": canonical,
100
+ "canonical_expected_total": CANONICAL_DENSE_PARAMETER_COUNT,
101
+ "canonical_match": canonical and actual == CANONICAL_DENSE_PARAMETER_COUNT,
102
+ }
103
+
104
+
105
+ def gpt_config_from_checkpoint_dict(cfg_dict: dict) -> GPTConfig:
106
+ """Rebuild a GPTConfig from a checkpoint's serialized config dict.
107
+
108
+ Pre-GQA checkpoints carry no n_kv_heads; absence means plain MHA
109
+ (n_kv_heads == n_heads), whose fused-QKV weight layout is unchanged.
110
+ """
111
+ cfg_dict = dict(cfg_dict)
112
+ cfg_dict.setdefault("n_kv_heads", cfg_dict["n_heads"])
113
+ return GPTConfig(**cfg_dict)
114
+
115
+
116
+ class RMSNorm(nn.Module):
117
+ def __init__(self, dim: int, eps: float = 1e-6):
118
+ super().__init__()
119
+ self.eps = eps
120
+ self.weight = nn.Parameter(torch.ones(dim))
121
+
122
+ def forward(self, x: torch.Tensor) -> torch.Tensor:
123
+ # x: [B, T, C]
124
+ rms = x.pow(2).mean(dim=-1, keepdim=True).add(self.eps).rsqrt()
125
+ return x * rms * self.weight
126
+
127
+
128
+ def _rotate_half(x: torch.Tensor) -> torch.Tensor:
129
+ # x: [..., D]. Half-split layout (Llama/GPT-NeoX): pairs are (i, i+D/2),
130
+ # matching the `cat([freqs, freqs])` cos/sin cache below.
131
+ half = x.shape[-1] // 2
132
+ x1 = x[..., :half]
133
+ x2 = x[..., half:]
134
+ return torch.cat((-x2, x1), dim=-1)
135
+
136
+
137
+ class RotaryEmbedding(nn.Module):
138
+ """Precomputes RoPE cos/sin caches up to max_seq_len."""
139
+
140
+ def __init__(self, head_dim: int, max_seq_len: int, theta: float = 10000.0, pct: float = 1.0):
141
+ super().__init__()
142
+ if head_dim % 2 != 0:
143
+ raise ValueError(f"RoPE requires even head_dim, got {head_dim}")
144
+ self.head_dim = int(head_dim)
145
+ self.max_seq_len = int(max_seq_len)
146
+ self.theta = float(theta)
147
+ self.pct = float(pct)
148
+
149
+ rope_dim = int(self.head_dim * self.pct)
150
+ rope_dim = rope_dim - (rope_dim % 2)
151
+ rope_dim = max(0, min(rope_dim, self.head_dim))
152
+ self.rope_dim = rope_dim
153
+
154
+ if self.rope_dim > 0:
155
+ inv_freq = 1.0 / (
156
+ self.theta ** (torch.arange(0, self.rope_dim, 2).float() / self.rope_dim)
157
+ )
158
+ t = torch.arange(self.max_seq_len, dtype=torch.float32)
159
+ freqs = torch.outer(t, inv_freq) # [T, rope_dim/2]
160
+ emb = torch.cat([freqs, freqs], dim=-1) # [T, rope_dim]
161
+ cos = emb.cos()
162
+ sin = emb.sin()
163
+ else:
164
+ cos = torch.empty(self.max_seq_len, 0, dtype=torch.float32)
165
+ sin = torch.empty(self.max_seq_len, 0, dtype=torch.float32)
166
+
167
+ self.register_buffer("cos_cached", cos, persistent=False)
168
+ self.register_buffer("sin_cached", sin, persistent=False)
169
+
170
+ def forward(
171
+ self, q: torch.Tensor, k: torch.Tensor, seq_len: int, offset: int = 0
172
+ ) -> tuple[torch.Tensor, torch.Tensor]:
173
+ """Apply RoPE to q,k. q,k: [B, nH, T, Hd].
174
+
175
+ `offset` is the absolute position of the first token in q,k — nonzero
176
+ during KV-cached incremental decoding, where the new tokens sit at
177
+ positions [offset, offset+seq_len).
178
+ """
179
+ end = offset + seq_len
180
+ if end > self.max_seq_len:
181
+ raise ValueError(
182
+ f"position {end} exceeds max_seq_len={self.max_seq_len} for RoPE cache"
183
+ )
184
+ if self.rope_dim == 0:
185
+ return q, k
186
+
187
+ cos = self.cos_cached[offset:end].to(dtype=q.dtype, device=q.device) # [T, rope_dim]
188
+ sin = self.sin_cached[offset:end].to(dtype=q.dtype, device=q.device) # [T, rope_dim]
189
+ cos = cos.unsqueeze(0).unsqueeze(0) # [1,1,T,rope_dim]
190
+ sin = sin.unsqueeze(0).unsqueeze(0)
191
+
192
+ q1, q2 = q[..., : self.rope_dim], q[..., self.rope_dim :]
193
+ k1, k2 = k[..., : self.rope_dim], k[..., self.rope_dim :]
194
+
195
+ q1 = q1 * cos + _rotate_half(q1) * sin
196
+ k1 = k1 * cos + _rotate_half(k1) * sin
197
+
198
+ q = torch.cat([q1, q2], dim=-1)
199
+ k = torch.cat([k1, k2], dim=-1)
200
+ return q, k
201
+
202
+
203
+ class SwiGLU(nn.Module):
204
+ def __init__(self, cfg: GPTConfig):
205
+ super().__init__()
206
+ self.w1 = nn.Linear(cfg.d_model, cfg.d_ff, bias=False)
207
+ self.w3 = nn.Linear(cfg.d_model, cfg.d_ff, bias=False)
208
+ self.w2 = nn.Linear(cfg.d_ff, cfg.d_model, bias=False)
209
+ self.drop = nn.Dropout(cfg.dropout)
210
+
211
+ def forward(self, x: torch.Tensor) -> torch.Tensor:
212
+ x = F.silu(self.w1(x)) * self.w3(x)
213
+ x = self.w2(x)
214
+ return self.drop(x)
215
+
216
+
217
+ class CausalSelfAttention(nn.Module):
218
+ def __init__(self, cfg: GPTConfig):
219
+ super().__init__()
220
+ assert cfg.d_model % cfg.n_heads == 0
221
+ assert cfg.n_heads % cfg.n_kv_heads == 0
222
+ self.cfg = cfg
223
+ self.head_dim = cfg.d_model // cfg.n_heads
224
+ self.kv_dim = cfg.n_kv_heads * self.head_dim
225
+
226
+ # QKV fused: one matmul instead of three. K/V carry n_kv_heads (GQA);
227
+ # n_kv_heads == n_heads is plain MHA with the historical 3*d_model layout.
228
+ self.qkv = nn.Linear(cfg.d_model, cfg.d_model + 2 * self.kv_dim, bias=False)
229
+ # residual branch output projection
230
+ self.proj = nn.Linear(cfg.d_model, cfg.d_model, bias=False)
231
+ self.drop = nn.Dropout(cfg.dropout)
232
+
233
+ self.rope = RotaryEmbedding(
234
+ head_dim=self.head_dim,
235
+ max_seq_len=cfg.max_seq_len,
236
+ theta=cfg.rope_theta,
237
+ pct=cfg.rope_pct,
238
+ )
239
+
240
+ @staticmethod
241
+ def _incremental_mask(T: int, past_len: int, device: torch.device) -> torch.Tensor:
242
+ """Bottom-right causal mask [T, past_len+T] (True = attend) for decoding
243
+ T new queries against past_len cached keys plus the new keys."""
244
+ q_pos = past_len + torch.arange(T, device=device)
245
+ k_pos = torch.arange(past_len + T, device=device)
246
+ return k_pos[None, :] <= q_pos[:, None]
247
+
248
+ def forward(
249
+ self,
250
+ x: torch.Tensor,
251
+ past_kv: tuple[torch.Tensor, torch.Tensor] | None = None,
252
+ use_cache: bool = False,
253
+ ):
254
+ """x: [B, T, C]. With no cache this is byte-identical to a plain causal
255
+ forward and returns the output tensor. With `use_cache` (or a supplied
256
+ `past_kv`) it also returns the updated (k, v) for this layer."""
257
+ B, T, C = x.shape
258
+ past_len = 0 if past_kv is None else past_kv[0].size(2)
259
+ if past_len + T > self.cfg.max_seq_len:
260
+ raise ValueError(
261
+ f"cache length {past_len + T} exceeds max_seq_len={self.cfg.max_seq_len}"
262
+ )
263
+
264
+ qkv = self.qkv(x) # [B, T, C + 2*kv_dim]
265
+ q, k, v = qkv.split([C, self.kv_dim, self.kv_dim], dim=-1)
266
+
267
+ q = q.view(B, T, self.cfg.n_heads, self.head_dim).transpose(1, 2) # [B,nH,T,Hd]
268
+ k = k.view(B, T, self.cfg.n_kv_heads, self.head_dim).transpose(1, 2) # [B,nKV,T,Hd]
269
+ v = v.view(B, T, self.cfg.n_kv_heads, self.head_dim).transpose(1, 2)
270
+
271
+ # RoPE rotates only the new tokens, at their absolute positions.
272
+ q, k = self.rope(q, k, seq_len=T, offset=past_len)
273
+
274
+ # Prepend cached keys/values (already rotated when they were new). The
275
+ # cache stays un-expanded at n_kv_heads so its memory reflects GQA.
276
+ if past_kv is not None:
277
+ k = torch.cat([past_kv[0], k], dim=2)
278
+ v = torch.cat([past_kv[1], v], dim=2)
279
+ present = (k, v) if use_cache else None
280
+
281
+ # Expand grouped KV heads to the full head count for attention. KV head g
282
+ # serves query heads [g*rep, (g+1)*rep) — repeat_interleave matches SDPA's
283
+ # enable_gqa grouping (torch >= 2.5), which can replace this someday.
284
+ if self.cfg.n_kv_heads != self.cfg.n_heads:
285
+ rep = self.cfg.n_heads // self.cfg.n_kv_heads
286
+ k = k.repeat_interleave(rep, dim=1)
287
+ v = v.repeat_interleave(rep, dim=1)
288
+
289
+ dropout_p = float(self.cfg.dropout) if (self.training and self.cfg.dropout > 0) else 0.0
290
+
291
+ if q.device.type == "cuda":
292
+ if past_len == 0:
293
+ y = F.scaled_dot_product_attention(
294
+ q, k, v, attn_mask=None, dropout_p=dropout_p, is_causal=True
295
+ )
296
+ else:
297
+ y = F.scaled_dot_product_attention(
298
+ q,
299
+ k,
300
+ v,
301
+ attn_mask=self._incremental_mask(T, past_len, q.device),
302
+ dropout_p=dropout_p,
303
+ )
304
+ else:
305
+ scale = 1.0 / math.sqrt(self.head_dim)
306
+ att = torch.matmul(q * scale, k.transpose(-2, -1)) # [B,nH,T,past_len+T]
307
+ if past_len == 0:
308
+ mask = torch.triu(torch.ones((T, T), device=q.device, dtype=torch.bool), diagonal=1)
309
+ att = att.masked_fill(mask, float("-inf"))
310
+ else:
311
+ allow = self._incremental_mask(T, past_len, q.device) # [T, past_len+T]
312
+ att = att.masked_fill(~allow, float("-inf"))
313
+ att = F.softmax(att, dim=-1)
314
+ if dropout_p > 0.0:
315
+ att = F.dropout(att, p=dropout_p)
316
+ y = torch.matmul(att, v)
317
+
318
+ y = y.transpose(1, 2).contiguous().view(B, T, C)
319
+ y = self.drop(self.proj(y))
320
+ if use_cache:
321
+ return y, present
322
+ return y
323
+
324
+
325
+ class Block(nn.Module):
326
+ def __init__(self, cfg: GPTConfig):
327
+ super().__init__()
328
+ self.norm1 = RMSNorm(cfg.d_model)
329
+ self.attn = CausalSelfAttention(cfg)
330
+ self.norm2 = RMSNorm(cfg.d_model)
331
+ self.mlp = SwiGLU(cfg)
332
+
333
+ def forward(
334
+ self,
335
+ x: torch.Tensor,
336
+ past_kv: tuple[torch.Tensor, torch.Tensor] | None = None,
337
+ use_cache: bool = False,
338
+ ):
339
+ if use_cache or past_kv is not None:
340
+ attn_out, present = self.attn(self.norm1(x), past_kv=past_kv, use_cache=True)
341
+ x = x + attn_out
342
+ x = x + self.mlp(self.norm2(x))
343
+ return x, present
344
+ x = x + self.attn(self.norm1(x))
345
+ x = x + self.mlp(self.norm2(x))
346
+ return x
347
+
348
+
349
+ class GPT(nn.Module):
350
+ def __init__(self, cfg: GPTConfig):
351
+ super().__init__()
352
+ self.cfg = cfg
353
+
354
+ self.tok_emb = nn.Embedding(cfg.vocab_size, cfg.d_model)
355
+ self.drop = nn.Dropout(cfg.dropout)
356
+
357
+ self.blocks = nn.ModuleList([Block(cfg) for _ in range(cfg.n_layers)])
358
+ self.norm_f = RMSNorm(cfg.d_model)
359
+
360
+ self.lm_head = nn.Linear(cfg.d_model, cfg.vocab_size, bias=False)
361
+ if cfg.tie_embeddings:
362
+ self.lm_head.weight = self.tok_emb.weight
363
+
364
+ # Base init everywhere...
365
+ self.apply(self._init_weights)
366
+ # ...then scale init ONLY on residual-branch output projections: attn.proj and mlp.w2
367
+ self._init_residual_projections()
368
+
369
+ def _init_weights(self, m: nn.Module):
370
+ if isinstance(m, (nn.Linear, nn.Embedding)):
371
+ torch.nn.init.normal_(m.weight, mean=0.0, std=0.02)
372
+
373
+ def _init_residual_projections(self):
374
+ std = 0.02 / math.sqrt(2.0 * float(self.cfg.n_layers))
375
+ for blk in self.blocks:
376
+ torch.nn.init.normal_(blk.attn.proj.weight, mean=0.0, std=std)
377
+ torch.nn.init.normal_(blk.mlp.w2.weight, mean=0.0, std=std)
378
+
379
+ def forward(
380
+ self,
381
+ input_ids: torch.Tensor,
382
+ past_kv: list[tuple[torch.Tensor, torch.Tensor]] | None = None,
383
+ use_cache: bool = False,
384
+ ):
385
+ """Default call `model(input_ids)` returns logits [B, T, V] — unchanged.
386
+
387
+ For incremental decoding, pass `use_cache=True` to also get a per-layer
388
+ list of (k, v) tensors, and feed it back as `past_kv` with only the new
389
+ token(s) on the next call. See `generate`.
390
+ """
391
+ B, T = input_ids.shape
392
+ past_len = 0 if past_kv is None else past_kv[0][0].size(2)
393
+ if past_len + T > self.cfg.max_seq_len:
394
+ raise ValueError(
395
+ f"T={T} with cache={past_len} exceeds max_seq_len={self.cfg.max_seq_len}"
396
+ )
397
+ if T < 1:
398
+ raise ValueError("Empty sequence")
399
+ caching = use_cache or (past_kv is not None)
400
+
401
+ x = self.tok_emb(input_ids)
402
+ x = self.drop(x)
403
+ presents: list[tuple[torch.Tensor, torch.Tensor]] = []
404
+ for i, blk in enumerate(self.blocks):
405
+ layer_past = past_kv[i] if past_kv is not None else None
406
+ if caching:
407
+ x, present = blk(x, past_kv=layer_past, use_cache=True)
408
+ presents.append(present)
409
+ else:
410
+ x = blk(x)
411
+ x = self.norm_f(x)
412
+ logits = self.lm_head(x)
413
+ if caching:
414
+ return logits, presents
415
+ return logits
416
+
417
+ @staticmethod
418
+ def _sample_token(
419
+ logits: torch.Tensor, temperature: float, top_k: int, top_p: float
420
+ ) -> torch.Tensor:
421
+ """logits: [B, V] -> next token [B, 1]. temperature<=0 is greedy."""
422
+ if temperature <= 0:
423
+ return logits.argmax(dim=-1, keepdim=True)
424
+ logits = logits / temperature
425
+ if top_k and top_k > 0:
426
+ k = min(int(top_k), logits.size(-1))
427
+ thresh = torch.topk(logits, k, dim=-1).values[:, -1, None]
428
+ logits = logits.masked_fill(logits < thresh, float("-inf"))
429
+ if top_p and top_p < 1.0:
430
+ sorted_logits, sorted_idx = torch.sort(logits, descending=True, dim=-1)
431
+ cum = torch.softmax(sorted_logits, dim=-1).cumsum(dim=-1)
432
+ drop_sorted = cum > top_p
433
+ drop_sorted[..., 0] = False
434
+ drop = torch.zeros_like(drop_sorted).scatter(-1, sorted_idx, drop_sorted)
435
+ logits = logits.masked_fill(drop, float("-inf"))
436
+ probs = torch.softmax(logits, dim=-1)
437
+ return torch.multinomial(probs, num_samples=1)
438
+
439
+ @torch.no_grad()
440
+ def generate(
441
+ self,
442
+ input_ids: torch.Tensor,
443
+ max_new_tokens: int,
444
+ *,
445
+ temperature: float = 1.0,
446
+ top_k: int = 0,
447
+ top_p: float = 1.0,
448
+ eos_id: int | None = None,
449
+ ) -> torch.Tensor:
450
+ """KV-cached incremental decoding. input_ids: [B, T] -> [B, T + n].
451
+
452
+ Prefills the prompt once, then feeds one new token per step against the
453
+ cache (O(T) forwards of length 1) instead of re-running the full growing
454
+ sequence each step. Stops early if all rows emit `eos_id`.
455
+ """
456
+ was_training = self.training
457
+ self.eval()
458
+ logits, past = self.forward(input_ids, use_cache=True)
459
+ out = input_ids
460
+ for _ in range(int(max_new_tokens)):
461
+ next_tok = self._sample_token(logits[:, -1, :], temperature, top_k, top_p)
462
+ out = torch.cat([out, next_tok], dim=1)
463
+ if eos_id is not None and bool((next_tok.squeeze(1) == eos_id).all()):
464
+ break
465
+ if out.size(1) >= self.cfg.max_seq_len:
466
+ break
467
+ logits, past = self.forward(next_tok, past_kv=past, use_cache=True)
468
+ if was_training:
469
+ self.train()
470
+ return out
src/special_tokens.py ADDED
@@ -0,0 +1,138 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Canonical special-token IDs for the whole pipeline.
2
+
3
+ Every training/data script hardcodes these IDs (see CLAUDE.md):
4
+
5
+ [PAD]=0 [UNK]=1 [BOS]=2 [EOS]=3
6
+ <|system|>=4 <|user|>=5 <|assistant|>=6
7
+
8
+ The three role tokens delimit chat turns at the *token* level (see
9
+ src/chat_template.py) — BPE can never merge across a special token, so
10
+ training-time and inference-time encodings agree by construction.
11
+
12
+ `assert_special_token_ids` turns the silent assumption into a loud startup
13
+ check: if the tokenizer at --tokenizer_path is ever retrained and the IDs move,
14
+ scripts fail immediately instead of training with a misaligned loss mask or a
15
+ broken EOS stop condition.
16
+ """
17
+
18
+ from __future__ import annotations
19
+
20
+ import json
21
+ from os import PathLike
22
+
23
+ PAD_ID = 0
24
+ UNK_ID = 1
25
+ BOS_ID = 2
26
+ EOS_ID = 3
27
+ SYSTEM_ID = 4
28
+ USER_ID = 5
29
+ ASSISTANT_ID = 6
30
+ CANONICAL_VOCAB_SIZE = 32_000
31
+
32
+ PAD_TOKEN = "[PAD]"
33
+ UNK_TOKEN = "[UNK]"
34
+ BOS_TOKEN = "[BOS]"
35
+ EOS_TOKEN = "[EOS]"
36
+ SYSTEM_TOKEN = "<|system|>"
37
+ USER_TOKEN = "<|user|>"
38
+ ASSISTANT_TOKEN = "<|assistant|>"
39
+
40
+ SPECIAL_TOKEN_IDS: dict[str, int] = {
41
+ PAD_TOKEN: PAD_ID,
42
+ UNK_TOKEN: UNK_ID,
43
+ BOS_TOKEN: BOS_ID,
44
+ EOS_TOKEN: EOS_ID,
45
+ SYSTEM_TOKEN: SYSTEM_ID,
46
+ USER_TOKEN: USER_ID,
47
+ ASSISTANT_TOKEN: ASSISTANT_ID,
48
+ }
49
+
50
+ # Ordered by ID — the exact list BpeTrainer must receive so IDs come out right.
51
+ SPECIAL_TOKENS: list[str] = sorted(SPECIAL_TOKEN_IDS, key=SPECIAL_TOKEN_IDS.get)
52
+
53
+
54
+ def assert_special_token_ids(tokenizer_path: str) -> None:
55
+ """Validate the exact registered-special-token contract.
56
+
57
+ Reads the file as plain JSON (no `tokenizers` import needed) so it is cheap
58
+ to call from any script entry point. Merely finding the strings in the BPE
59
+ vocabulary is insufficient: all seven must be registered as special tokens,
60
+ and no eighth registered special token is allowed.
61
+ """
62
+ with open(tokenizer_path, encoding="utf-8") as f:
63
+ obj = json.load(f)
64
+
65
+ special_entries = [
66
+ entry
67
+ for entry in (obj.get("added_tokens") or [])
68
+ if isinstance(entry, dict) and entry.get("special") is True
69
+ ]
70
+ registered_tokens = [entry.get("content") for entry in special_entries]
71
+ if len(special_entries) != len(SPECIAL_TOKEN_IDS) or set(registered_tokens) != set(
72
+ SPECIAL_TOKEN_IDS
73
+ ):
74
+ raise ValueError(
75
+ f"{tokenizer_path} must register exactly these seven special tokens: "
76
+ f"{list(SPECIAL_TOKEN_IDS)}; got {registered_tokens!r}"
77
+ )
78
+
79
+ added = {entry["content"]: entry.get("id") for entry in special_entries}
80
+ for token, expected in SPECIAL_TOKEN_IDS.items():
81
+ got = added.get(token)
82
+ if got != expected:
83
+ raise ValueError(
84
+ f"special token {token!r} has id {got!r} in {tokenizer_path}, but the "
85
+ f"pipeline hardcodes {expected}. Retrain the tokenizer with "
86
+ f"tokenizer/tokenizer_training/train_tokenizer.py (its --strict_special_ids "
87
+ f"default enforces this layout) or reconcile src/special_tokens.py."
88
+ )
89
+
90
+
91
+ def assert_tokenizer_contract(tokenizer_path: str | PathLike[str]) -> None:
92
+ """Fail fast unless ``tokenizer.json`` is the canonical production artifact.
93
+
94
+ This is deliberately stricter than :func:`assert_special_token_ids`, which
95
+ remains useful for tiny unit-test tokenizers. Production entry points must
96
+ call this function so a valid-looking seven-token map cannot hide an old
97
+ vocabulary, normalizer, prefix-space rule, or automatic BOS/EOS processor.
98
+ """
99
+ path = str(tokenizer_path)
100
+ assert_special_token_ids(path)
101
+ with open(path, encoding="utf-8") as f:
102
+ obj = json.load(f)
103
+
104
+ model = obj.get("model") or {}
105
+ if model.get("type") != "BPE":
106
+ raise ValueError(f"{path} must use a BPE model; got {model.get('type')!r}")
107
+ if model.get("unk_token") != UNK_TOKEN:
108
+ raise ValueError(
109
+ f"{path} BPE unk_token must be {UNK_TOKEN!r}; got {model.get('unk_token')!r}"
110
+ )
111
+
112
+ model_vocab = model.get("vocab") or {}
113
+ token_ids = set(model_vocab.values())
114
+ token_ids.update(
115
+ entry.get("id")
116
+ for entry in (obj.get("added_tokens") or [])
117
+ if isinstance(entry, dict) and isinstance(entry.get("id"), int)
118
+ )
119
+ expected_ids = set(range(CANONICAL_VOCAB_SIZE))
120
+ if token_ids != expected_ids:
121
+ raise ValueError(
122
+ f"{path} runtime vocab IDs must be exactly 0..{CANONICAL_VOCAB_SIZE - 1} "
123
+ f"(vocab_size={CANONICAL_VOCAB_SIZE}); got {len(token_ids)} unique IDs"
124
+ )
125
+
126
+ if obj.get("normalizer") is not None:
127
+ raise ValueError(f"{path} must not configure a tokenizer normalizer")
128
+ if obj.get("post_processor") is not None:
129
+ raise ValueError(f"{path} must not configure an automatic BOS/EOS post-processor")
130
+
131
+ pre = obj.get("pre_tokenizer") or {}
132
+ if pre.get("type") != "ByteLevel" or pre.get("add_prefix_space") is not False:
133
+ raise ValueError(
134
+ f"{path} pre_tokenizer must be ByteLevel(add_prefix_space=False); got {pre!r}"
135
+ )
136
+ decoder = obj.get("decoder") or {}
137
+ if decoder.get("type") != "ByteLevel":
138
+ raise ValueError(f"{path} decoder must be ByteLevel; got {decoder!r}")
tables/ASSISTANT_RESULTS_VERSIONED.csv ADDED
@@ -0,0 +1,145 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ score_version,model,slice,axis,true,false,unknown,n,bound_lower,bound_upper,bound_lower_pct,bound_upper_pct,finite_execution_categories,evidence_id
2
+ assistant_review_fable_v1,alpha075,old_qa41,content,7,33,1,41,,,,,,EVID-ASSIST-V1
3
+ assistant_review_fable_v1,alpha075,old_qa41,joint,7,33,1,41,,,,,,EVID-ASSIST-V1
4
+ assistant_review_fable_v1,alpha075,old_qa41,format,0,1,0,1,,,,,,EVID-ASSIST-V1
5
+ assistant_review_fable_v1,alpha075,old_practical38,content,10,26,2,38,,,,,,EVID-ASSIST-V1
6
+ assistant_review_fable_v1,alpha075,old_practical38,joint,10,26,2,38,,,,,,EVID-ASSIST-V1
7
+ assistant_review_fable_v1,alpha075,old_practical38,format,20,1,0,21,,,,,,EVID-ASSIST-V1
8
+ assistant_review_fable_v1,alpha075,old_python14,content,0,14,0,14,,,,,,EVID-ASSIST-V1
9
+ assistant_review_fable_v1,alpha075,old_python14,joint,0,14,0,14,,,,,,EVID-ASSIST-V1
10
+ assistant_review_fable_v1,alpha075,old_python14,format,3,1,0,4,,,,,,EVID-ASSIST-V1
11
+ assistant_review_fable_v1,alpha075,old_python14,interface,11,3,0,14,,,,,,EVID-ASSIST-V1
12
+ assistant_review_fable_v1,alpha075,old_python14,historical_finite_categories,,,,,,,,,"{""not_recorded"": 14}",EVID-ASSIST-V1
13
+ assistant_review_fable_v1,alpha075,new_natural64,content,2,62,0,64,,,,,,EVID-ASSIST-V1
14
+ assistant_review_fable_v1,alpha075,new_natural64,joint,2,62,0,64,,,,,,EVID-ASSIST-V1
15
+ assistant_review_fable_v1,alpha075,new_natural64,format,7,11,0,18,,,,,,EVID-ASSIST-V1
16
+ assistant_review_fable_v1,alpha075,new_python32,content,0,32,0,32,,,,,,EVID-ASSIST-V1
17
+ assistant_review_fable_v1,alpha075,new_python32,joint,0,32,0,32,,,,,,EVID-ASSIST-V1
18
+ assistant_review_fable_v1,alpha075,new_python32,format,0,0,0,0,,,,,,EVID-ASSIST-V1
19
+ assistant_review_fable_v1,alpha075,new_python32,interface,31,1,0,32,,,,,,EVID-ASSIST-V1
20
+ assistant_review_fable_v1,alpha075,new_python32,historical_finite_categories,,,,,,,,,"{""fail"": 25, ""fail_or_unextractable"": 1, ""tool_unknown"": 6}",EVID-ASSIST-V1
21
+ assistant_review_fable_v1,alpha075,practical27_joint,joint_group_bounds,,,,,26/81,10/27,32.098765,37.037037,,EVID-ASSIST-V1
22
+ assistant_review_fable_v1,alpha075,new_natural64,category:R1DEV-CONTEXT,0,16,0,16,,,,,,EVID-ASSIST-V1
23
+ assistant_review_fable_v1,alpha075,new_natural64,category:R1DEV-REWRITE,0,16,0,16,,,,,,EVID-ASSIST-V1
24
+ assistant_review_fable_v1,alpha075,new_natural64,category:R1DEV-FOLLOWUP,1,15,0,16,,,,,,EVID-ASSIST-V1
25
+ assistant_review_fable_v1,alpha075,new_natural64,category:R1DEV-CONDITIONAL,1,15,0,16,,,,,,EVID-ASSIST-V1
26
+ assistant_review_fable_v1,EXT-A,old_qa41,content,7,34,0,41,,,,,,EVID-ASSIST-V1
27
+ assistant_review_fable_v1,EXT-A,old_qa41,joint,7,34,0,41,,,,,,EVID-ASSIST-V1
28
+ assistant_review_fable_v1,EXT-A,old_qa41,format,0,1,0,1,,,,,,EVID-ASSIST-V1
29
+ assistant_review_fable_v1,EXT-A,old_practical38,content,18,18,2,38,,,,,,EVID-ASSIST-V1
30
+ assistant_review_fable_v1,EXT-A,old_practical38,joint,18,18,2,38,,,,,,EVID-ASSIST-V1
31
+ assistant_review_fable_v1,EXT-A,old_practical38,format,18,3,0,21,,,,,,EVID-ASSIST-V1
32
+ assistant_review_fable_v1,EXT-A,old_python14,content,2,12,0,14,,,,,,EVID-ASSIST-V1
33
+ assistant_review_fable_v1,EXT-A,old_python14,joint,2,12,0,14,,,,,,EVID-ASSIST-V1
34
+ assistant_review_fable_v1,EXT-A,old_python14,format,3,1,0,4,,,,,,EVID-ASSIST-V1
35
+ assistant_review_fable_v1,EXT-A,old_python14,interface,14,0,0,14,,,,,,EVID-ASSIST-V1
36
+ assistant_review_fable_v1,EXT-A,old_python14,historical_finite_categories,,,,,,,,,"{""fail"": 7, ""fail_or_unextractable"": 2, ""pass"": 3, ""tool_unknown"": 2}",EVID-ASSIST-V1
37
+ assistant_review_fable_v1,EXT-A,new_natural64,content,4,60,0,64,,,,,,EVID-ASSIST-V1
38
+ assistant_review_fable_v1,EXT-A,new_natural64,joint,4,60,0,64,,,,,,EVID-ASSIST-V1
39
+ assistant_review_fable_v1,EXT-A,new_natural64,format,13,5,0,18,,,,,,EVID-ASSIST-V1
40
+ assistant_review_fable_v1,EXT-A,new_python32,content,4,28,0,32,,,,,,EVID-ASSIST-V1
41
+ assistant_review_fable_v1,EXT-A,new_python32,joint,4,28,0,32,,,,,,EVID-ASSIST-V1
42
+ assistant_review_fable_v1,EXT-A,new_python32,format,0,0,0,0,,,,,,EVID-ASSIST-V1
43
+ assistant_review_fable_v1,EXT-A,new_python32,interface,32,0,0,32,,,,,,EVID-ASSIST-V1
44
+ assistant_review_fable_v1,EXT-A,new_python32,historical_finite_categories,,,,,,,,,"{""fail"": 19, ""fail_or_unextractable"": 4, ""pass"": 2, ""tool_unknown"": 7}",EVID-ASSIST-V1
45
+ assistant_review_fable_v1,EXT-A,practical27_joint,joint_group_bounds,,,,,47/81,17/27,58.024691,62.962963,,EVID-ASSIST-V1
46
+ assistant_review_fable_v1,EXT-A,new_natural64,category:R1DEV-CONTEXT,0,16,0,16,,,,,,EVID-ASSIST-V1
47
+ assistant_review_fable_v1,EXT-A,new_natural64,category:R1DEV-REWRITE,2,14,0,16,,,,,,EVID-ASSIST-V1
48
+ assistant_review_fable_v1,EXT-A,new_natural64,category:R1DEV-FOLLOWUP,2,14,0,16,,,,,,EVID-ASSIST-V1
49
+ assistant_review_fable_v1,EXT-A,new_natural64,category:R1DEV-CONDITIONAL,0,16,0,16,,,,,,EVID-ASSIST-V1
50
+ assistant_review_fable_v1,EXT-B,old_qa41,content,5,35,1,41,,,,,,EVID-ASSIST-V1
51
+ assistant_review_fable_v1,EXT-B,old_qa41,joint,5,35,1,41,,,,,,EVID-ASSIST-V1
52
+ assistant_review_fable_v1,EXT-B,old_qa41,format,0,1,0,1,,,,,,EVID-ASSIST-V1
53
+ assistant_review_fable_v1,EXT-B,old_practical38,content,6,31,1,38,,,,,,EVID-ASSIST-V1
54
+ assistant_review_fable_v1,EXT-B,old_practical38,joint,6,31,1,38,,,,,,EVID-ASSIST-V1
55
+ assistant_review_fable_v1,EXT-B,old_practical38,format,1,20,0,21,,,,,,EVID-ASSIST-V1
56
+ assistant_review_fable_v1,EXT-B,old_python14,content,4,10,0,14,,,,,,EVID-ASSIST-V1
57
+ assistant_review_fable_v1,EXT-B,old_python14,joint,3,11,0,14,,,,,,EVID-ASSIST-V1
58
+ assistant_review_fable_v1,EXT-B,old_python14,format,2,2,0,4,,,,,,EVID-ASSIST-V1
59
+ assistant_review_fable_v1,EXT-B,old_python14,interface,13,1,0,14,,,,,,EVID-ASSIST-V1
60
+ assistant_review_fable_v1,EXT-B,old_python14,historical_finite_categories,,,,,,,,,"{""fail"": 8, ""fail_or_unextractable"": 1, ""pass"": 4, ""tool_unknown"": 1}",EVID-ASSIST-V1
61
+ assistant_review_fable_v1,EXT-B,new_natural64,content,0,64,0,64,,,,,,EVID-ASSIST-V1
62
+ assistant_review_fable_v1,EXT-B,new_natural64,joint,0,64,0,64,,,,,,EVID-ASSIST-V1
63
+ assistant_review_fable_v1,EXT-B,new_natural64,format,3,15,0,18,,,,,,EVID-ASSIST-V1
64
+ assistant_review_fable_v1,EXT-B,new_python32,content,3,29,0,32,,,,,,EVID-ASSIST-V1
65
+ assistant_review_fable_v1,EXT-B,new_python32,joint,3,29,0,32,,,,,,EVID-ASSIST-V1
66
+ assistant_review_fable_v1,EXT-B,new_python32,format,0,0,0,0,,,,,,EVID-ASSIST-V1
67
+ assistant_review_fable_v1,EXT-B,new_python32,interface,31,1,0,32,,,,,,EVID-ASSIST-V1
68
+ assistant_review_fable_v1,EXT-B,new_python32,historical_finite_categories,,,,,,,,,"{""fail"": 21, ""fail_or_unextractable"": 2, ""pass"": 7, ""tool_unknown"": 2}",EVID-ASSIST-V1
69
+ assistant_review_fable_v1,EXT-B,practical27_joint,joint_group_bounds,,,,,8/81,1/9,9.876543,11.111111,,EVID-ASSIST-V1
70
+ assistant_review_fable_v1,EXT-B,new_natural64,category:R1DEV-CONTEXT,0,16,0,16,,,,,,EVID-ASSIST-V1
71
+ assistant_review_fable_v1,EXT-B,new_natural64,category:R1DEV-REWRITE,0,16,0,16,,,,,,EVID-ASSIST-V1
72
+ assistant_review_fable_v1,EXT-B,new_natural64,category:R1DEV-FOLLOWUP,0,16,0,16,,,,,,EVID-ASSIST-V1
73
+ assistant_review_fable_v1,EXT-B,new_natural64,category:R1DEV-CONDITIONAL,0,16,0,16,,,,,,EVID-ASSIST-V1
74
+ assistant_owner_clarification_4_v1,alpha075,old_qa41,content,6,33,2,41,,,,,,EVID-ASSIST-C4
75
+ assistant_owner_clarification_4_v1,alpha075,old_qa41,joint,6,33,2,41,,,,,,EVID-ASSIST-C4
76
+ assistant_owner_clarification_4_v1,alpha075,old_qa41,format,0,1,0,1,,,,,,EVID-ASSIST-C4
77
+ assistant_owner_clarification_4_v1,alpha075,old_practical38,content,10,26,2,38,,,,,,EVID-ASSIST-C4
78
+ assistant_owner_clarification_4_v1,alpha075,old_practical38,joint,10,26,2,38,,,,,,EVID-ASSIST-C4
79
+ assistant_owner_clarification_4_v1,alpha075,old_practical38,format,20,1,0,21,,,,,,EVID-ASSIST-C4
80
+ assistant_owner_clarification_4_v1,alpha075,old_python14,content,0,14,0,14,,,,,,EVID-ASSIST-C4
81
+ assistant_owner_clarification_4_v1,alpha075,old_python14,joint,0,14,0,14,,,,,,EVID-ASSIST-C4
82
+ assistant_owner_clarification_4_v1,alpha075,old_python14,format,3,1,0,4,,,,,,EVID-ASSIST-C4
83
+ assistant_owner_clarification_4_v1,alpha075,old_python14,interface,11,3,0,14,,,,,,EVID-ASSIST-C4
84
+ assistant_owner_clarification_4_v1,alpha075,old_python14,historical_finite_categories,,,,,,,,,"{""not_recorded"": 14}",EVID-ASSIST-C4
85
+ assistant_owner_clarification_4_v1,alpha075,new_natural64,content,3,61,0,64,,,,,,EVID-ASSIST-C4
86
+ assistant_owner_clarification_4_v1,alpha075,new_natural64,joint,3,61,0,64,,,,,,EVID-ASSIST-C4
87
+ assistant_owner_clarification_4_v1,alpha075,new_natural64,format,7,11,0,18,,,,,,EVID-ASSIST-C4
88
+ assistant_owner_clarification_4_v1,alpha075,new_python32,content,0,32,0,32,,,,,,EVID-ASSIST-C4
89
+ assistant_owner_clarification_4_v1,alpha075,new_python32,joint,0,32,0,32,,,,,,EVID-ASSIST-C4
90
+ assistant_owner_clarification_4_v1,alpha075,new_python32,format,0,0,0,0,,,,,,EVID-ASSIST-C4
91
+ assistant_owner_clarification_4_v1,alpha075,new_python32,interface,31,1,0,32,,,,,,EVID-ASSIST-C4
92
+ assistant_owner_clarification_4_v1,alpha075,new_python32,historical_finite_categories,,,,,,,,,"{""fail"": 25, ""fail_or_unextractable"": 1, ""tool_unknown"": 6}",EVID-ASSIST-C4
93
+ assistant_owner_clarification_4_v1,alpha075,practical27_joint,joint_group_bounds,,,,,26/81,10/27,32.098765,37.037037,,EVID-ASSIST-C4
94
+ assistant_owner_clarification_4_v1,alpha075,new_natural64,category:R1DEV-CONTEXT,0,16,0,16,,,,,,EVID-ASSIST-C4
95
+ assistant_owner_clarification_4_v1,alpha075,new_natural64,category:R1DEV-REWRITE,0,16,0,16,,,,,,EVID-ASSIST-C4
96
+ assistant_owner_clarification_4_v1,alpha075,new_natural64,category:R1DEV-FOLLOWUP,2,14,0,16,,,,,,EVID-ASSIST-C4
97
+ assistant_owner_clarification_4_v1,alpha075,new_natural64,category:R1DEV-CONDITIONAL,1,15,0,16,,,,,,EVID-ASSIST-C4
98
+ assistant_owner_clarification_4_v1,EXT-A,old_qa41,content,7,34,0,41,,,,,,EVID-ASSIST-C4
99
+ assistant_owner_clarification_4_v1,EXT-A,old_qa41,joint,7,34,0,41,,,,,,EVID-ASSIST-C4
100
+ assistant_owner_clarification_4_v1,EXT-A,old_qa41,format,0,1,0,1,,,,,,EVID-ASSIST-C4
101
+ assistant_owner_clarification_4_v1,EXT-A,old_practical38,content,18,18,2,38,,,,,,EVID-ASSIST-C4
102
+ assistant_owner_clarification_4_v1,EXT-A,old_practical38,joint,18,18,2,38,,,,,,EVID-ASSIST-C4
103
+ assistant_owner_clarification_4_v1,EXT-A,old_practical38,format,18,3,0,21,,,,,,EVID-ASSIST-C4
104
+ assistant_owner_clarification_4_v1,EXT-A,old_python14,content,1,13,0,14,,,,,,EVID-ASSIST-C4
105
+ assistant_owner_clarification_4_v1,EXT-A,old_python14,joint,1,13,0,14,,,,,,EVID-ASSIST-C4
106
+ assistant_owner_clarification_4_v1,EXT-A,old_python14,format,3,1,0,4,,,,,,EVID-ASSIST-C4
107
+ assistant_owner_clarification_4_v1,EXT-A,old_python14,interface,14,0,0,14,,,,,,EVID-ASSIST-C4
108
+ assistant_owner_clarification_4_v1,EXT-A,old_python14,historical_finite_categories,,,,,,,,,"{""fail"": 7, ""fail_or_unextractable"": 2, ""pass"": 3, ""tool_unknown"": 2}",EVID-ASSIST-C4
109
+ assistant_owner_clarification_4_v1,EXT-A,new_natural64,content,4,60,0,64,,,,,,EVID-ASSIST-C4
110
+ assistant_owner_clarification_4_v1,EXT-A,new_natural64,joint,4,60,0,64,,,,,,EVID-ASSIST-C4
111
+ assistant_owner_clarification_4_v1,EXT-A,new_natural64,format,13,5,0,18,,,,,,EVID-ASSIST-C4
112
+ assistant_owner_clarification_4_v1,EXT-A,new_python32,content,4,28,0,32,,,,,,EVID-ASSIST-C4
113
+ assistant_owner_clarification_4_v1,EXT-A,new_python32,joint,4,28,0,32,,,,,,EVID-ASSIST-C4
114
+ assistant_owner_clarification_4_v1,EXT-A,new_python32,format,0,0,0,0,,,,,,EVID-ASSIST-C4
115
+ assistant_owner_clarification_4_v1,EXT-A,new_python32,interface,32,0,0,32,,,,,,EVID-ASSIST-C4
116
+ assistant_owner_clarification_4_v1,EXT-A,new_python32,historical_finite_categories,,,,,,,,,"{""fail"": 19, ""fail_or_unextractable"": 4, ""pass"": 2, ""tool_unknown"": 7}",EVID-ASSIST-C4
117
+ assistant_owner_clarification_4_v1,EXT-A,practical27_joint,joint_group_bounds,,,,,47/81,17/27,58.024691,62.962963,,EVID-ASSIST-C4
118
+ assistant_owner_clarification_4_v1,EXT-A,new_natural64,category:R1DEV-CONTEXT,0,16,0,16,,,,,,EVID-ASSIST-C4
119
+ assistant_owner_clarification_4_v1,EXT-A,new_natural64,category:R1DEV-REWRITE,2,14,0,16,,,,,,EVID-ASSIST-C4
120
+ assistant_owner_clarification_4_v1,EXT-A,new_natural64,category:R1DEV-FOLLOWUP,2,14,0,16,,,,,,EVID-ASSIST-C4
121
+ assistant_owner_clarification_4_v1,EXT-A,new_natural64,category:R1DEV-CONDITIONAL,0,16,0,16,,,,,,EVID-ASSIST-C4
122
+ assistant_owner_clarification_4_v1,EXT-B,old_qa41,content,5,35,1,41,,,,,,EVID-ASSIST-C4
123
+ assistant_owner_clarification_4_v1,EXT-B,old_qa41,joint,5,35,1,41,,,,,,EVID-ASSIST-C4
124
+ assistant_owner_clarification_4_v1,EXT-B,old_qa41,format,0,1,0,1,,,,,,EVID-ASSIST-C4
125
+ assistant_owner_clarification_4_v1,EXT-B,old_practical38,content,6,31,1,38,,,,,,EVID-ASSIST-C4
126
+ assistant_owner_clarification_4_v1,EXT-B,old_practical38,joint,6,31,1,38,,,,,,EVID-ASSIST-C4
127
+ assistant_owner_clarification_4_v1,EXT-B,old_practical38,format,1,20,0,21,,,,,,EVID-ASSIST-C4
128
+ assistant_owner_clarification_4_v1,EXT-B,old_python14,content,4,10,0,14,,,,,,EVID-ASSIST-C4
129
+ assistant_owner_clarification_4_v1,EXT-B,old_python14,joint,3,11,0,14,,,,,,EVID-ASSIST-C4
130
+ assistant_owner_clarification_4_v1,EXT-B,old_python14,format,2,2,0,4,,,,,,EVID-ASSIST-C4
131
+ assistant_owner_clarification_4_v1,EXT-B,old_python14,interface,13,1,0,14,,,,,,EVID-ASSIST-C4
132
+ assistant_owner_clarification_4_v1,EXT-B,old_python14,historical_finite_categories,,,,,,,,,"{""fail"": 8, ""fail_or_unextractable"": 1, ""pass"": 4, ""tool_unknown"": 1}",EVID-ASSIST-C4
133
+ assistant_owner_clarification_4_v1,EXT-B,new_natural64,content,0,64,0,64,,,,,,EVID-ASSIST-C4
134
+ assistant_owner_clarification_4_v1,EXT-B,new_natural64,joint,0,64,0,64,,,,,,EVID-ASSIST-C4
135
+ assistant_owner_clarification_4_v1,EXT-B,new_natural64,format,3,15,0,18,,,,,,EVID-ASSIST-C4
136
+ assistant_owner_clarification_4_v1,EXT-B,new_python32,content,2,30,0,32,,,,,,EVID-ASSIST-C4
137
+ assistant_owner_clarification_4_v1,EXT-B,new_python32,joint,2,30,0,32,,,,,,EVID-ASSIST-C4
138
+ assistant_owner_clarification_4_v1,EXT-B,new_python32,format,0,0,0,0,,,,,,EVID-ASSIST-C4
139
+ assistant_owner_clarification_4_v1,EXT-B,new_python32,interface,31,1,0,32,,,,,,EVID-ASSIST-C4
140
+ assistant_owner_clarification_4_v1,EXT-B,new_python32,historical_finite_categories,,,,,,,,,"{""fail"": 21, ""fail_or_unextractable"": 2, ""pass"": 7, ""tool_unknown"": 2}",EVID-ASSIST-C4
141
+ assistant_owner_clarification_4_v1,EXT-B,practical27_joint,joint_group_bounds,,,,,8/81,1/9,9.876543,11.111111,,EVID-ASSIST-C4
142
+ assistant_owner_clarification_4_v1,EXT-B,new_natural64,category:R1DEV-CONTEXT,0,16,0,16,,,,,,EVID-ASSIST-C4
143
+ assistant_owner_clarification_4_v1,EXT-B,new_natural64,category:R1DEV-REWRITE,0,16,0,16,,,,,,EVID-ASSIST-C4
144
+ assistant_owner_clarification_4_v1,EXT-B,new_natural64,category:R1DEV-FOLLOWUP,0,16,0,16,,,,,,EVID-ASSIST-C4
145
+ assistant_owner_clarification_4_v1,EXT-B,new_natural64,category:R1DEV-CONDITIONAL,0,16,0,16,,,,,,EVID-ASSIST-C4
tables/PRETRAIN_SOURCE_MIXTURE.csv ADDED
@@ -0,0 +1,15 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ stage,source_display_name,stage_i_node_key,upstream_dataset,upstream_revision,upstream_config,recorded_licence_at_pinned_revision,target_serialized_tokens,selected_serialized_tokens,overshoot_tokens,selected_documents,share_of_stage_selected_pct,transport_repository_recorded,transport_revision_recorded
2
+ stage_a,FineWeb-Edu (dedup),fineweb_edu_dedup,HuggingFaceTB/smollm-corpus,3ba9d605774198c5868892d7a8deda78031a781f,fineweb-edu-dedup,odc-by-1.0,7110526316,7110526955,639,7350945,71.1052,HuggingFaceTB/smollm-corpus,3ba9d605774198c5868892d7a8deda78031a781f
3
+ stage_a,DCLM-Edu,dclm_edu,HuggingFaceTB/dclm-edu,dbad8ad71224482740cd9c9d353591adbf62fe04,default,cc-by-4.0,2031578947,2031579037,90,1593857,20.3158,HuggingFaceTB/dclm-edu,dbad8ad71224482740cd9c9d353591adbf62fe04
4
+ stage_a,Wikipedia (FineWiki EN),finewiki_en,HuggingFaceFW/finewiki,8bd13e72e6a002407649b3e898535f42ceb1aeb9,en,cc-by-sa-4.0,507894737,507896470,1733,557285,5.0790,HuggingFaceFW/finewiki,8bd13e72e6a002407649b3e898535f42ceb1aeb9
5
+ stage_a,Python-Edu,python_gate_c_full,common-pile/stackv2_edu_filtered,c354dbe88469a1153e97c6a63ac50591849654de,default,per-record metadata.license (Software Heritage permissive subset),350000000,350000772,772,487237,3.5000,NOT_RECORDED,NOT_RECORDED
6
+ stage_b,FineWeb-Edu (dedup),fineweb_edu_dedup,HuggingFaceTB/smollm-corpus,3ba9d605774198c5868892d7a8deda78031a781f,fineweb-edu-dedup,odc-by-1.0,1203125000,1203125470,470,1199264,40.1041,HuggingFaceTB/smollm-corpus,3ba9d605774198c5868892d7a8deda78031a781f
7
+ stage_b,DCLM-Edu,dclm_edu,HuggingFaceTB/dclm-edu,dbad8ad71224482740cd9c9d353591adbf62fe04,default,cc-by-4.0,687500000,687500443,443,538590,22.9166,HuggingFaceTB/dclm-edu,dbad8ad71224482740cd9c9d353591adbf62fe04
8
+ stage_b,structured_tutorial (Cosmopedia v2 + FinePhrase tutorial),structured_tutorial,HuggingFaceTB/smollm-corpus + HuggingFaceFW/finephrase,3ba9d605774198c5868892d7a8deda78031a781f + 78cf4a5ed0099214979c094c963e699c19163838,cosmopedia-v2 + tutorial,odc-by-1.0 (both),343750000,343750175,175,530450,11.4583,HuggingFaceTB/smollm-corpus + HuggingFaceFW/finephrase,3ba9d605774198c5868892d7a8deda78031a781f + 78cf4a5ed0099214979c094c963e699c19163838
9
+ stage_b,Python-Edu,python_gate_c_full,common-pile/stackv2_edu_filtered,c354dbe88469a1153e97c6a63ac50591849654de,default,per-record metadata.license (Software Heritage permissive subset),250000000,250000383,383,347145,8.3333,NOT_RECORDED,NOT_RECORDED
10
+ stage_b,Wikipedia (FineWiki EN),finewiki_en,HuggingFaceFW/finewiki,8bd13e72e6a002407649b3e898535f42ceb1aeb9,en,cc-by-sa-4.0,171875000,171877052,2052,189465,5.7292,HuggingFaceFW/finewiki,8bd13e72e6a002407649b3e898535f42ceb1aeb9
11
+ stage_b,PES2O,pes2o,allenai/dolmino-mix-1124,a319f19eef1e257417b11ea8c30da266ae175557,pes2o,odc-by-1.0,171875000,171875364,364,625468,5.7292,allenai/dolmino-mix-1124,c58ab4b6ff990115e1ff3121953754ee2bc29501
12
+ stage_b,StackExchange,stackexchange,allenai/dolmino-mix-1124,a319f19eef1e257417b11ea8c30da266ae175557,stackexchange,cc-by-sa,171875000,171875353,353,336025,5.7292,allenai/dolmino-mix-1124,c58ab4b6ff990115e1ff3121953754ee2bc29501
13
+ stage_a,TOTAL,,,,,,10000000000,10000003234,,9989324,100.0000,,
14
+ stage_b,TOTAL,,,,,,3000000000,3000004240,,3766407,100.0000,,
15
+ both,GRAND TOTAL,,,,,,13000000000,13000007474,,13755731,,,
tables/PUBLIC_BENCHMARK_RESULTS.csv ADDED
@@ -0,0 +1,7 @@
 
 
 
 
 
 
 
 
1
+ result_version,protocol_version,model,task,split,dataset,dataset_revision,documents,expected_documents,candidate_sequences,acc_correct,acc,acc_pct,acc_norm_correct,acc_norm,acc_norm_pct,status,historical_or_author_reported,runner,evidence_id
2
+ FP32_V2,FP32_V2,PetitGPT-alpha075,arc_easy,test,allenai/ai2_arc,210d026faf9955653af8916fad021475a3f00453,2376,2376,9501,1372,0.5774410774410774,57.74,1244,0.5235690235690236,52.36,COMPLETE,null,native_protocol_compatible_evaluator_NOT_installed_lm_eval,EVID-BENCH-01
3
+ FP32_V2,FP32_V2,PetitGPT-alpha075,piqa,validation,baber/piqa,142f6d7367fd9877f0fb3b5734ea6a545f54cdd1,1838,1838,3676,1167,0.6349292709466812,63.49,1145,0.6229597388465724,62.30,COMPLETE,null,native_protocol_compatible_evaluator_NOT_installed_lm_eval,EVID-BENCH-01
4
+ FP32_V2,FP32_V2,SmolLM-135M-Instruct,arc_easy,test,allenai/ai2_arc,210d026faf9955653af8916fad021475a3f00453,2376,2376,9501,1170,0.49242424242424243,49.24,1033,0.43476430976430974,43.48,COMPLETE,null,native_protocol_compatible_evaluator_NOT_installed_lm_eval,EVID-BENCH-01
5
+ FP32_V2,FP32_V2,SmolLM-135M-Instruct,piqa,validation,baber/piqa,142f6d7367fd9877f0fb3b5734ea6a545f54cdd1,1838,1838,3676,1233,0.6708378672470077,67.08,1236,0.6724700761697497,67.25,COMPLETE,null,native_protocol_compatible_evaluator_NOT_installed_lm_eval,EVID-BENCH-01
6
+ FP32_V2,FP32_V2,SmolLM2-135M-Instruct,arc_easy,test,allenai/ai2_arc,210d026faf9955653af8916fad021475a3f00453,2376,2376,9501,1283,0.539983164983165,54.00,1160,0.4882154882154882,48.82,COMPLETE,null,native_protocol_compatible_evaluator_NOT_installed_lm_eval,EVID-BENCH-01
7
+ FP32_V2,FP32_V2,SmolLM2-135M-Instruct,piqa,validation,baber/piqa,142f6d7367fd9877f0fb3b5734ea6a545f54cdd1,1838,1838,3676,1226,0.6670293797606094,66.70,1227,0.6675734494015234,66.76,COMPLETE,null,native_protocol_compatible_evaluator_NOT_installed_lm_eval,EVID-BENCH-01
tokenizer.json ADDED
The diff for this file is too large to render. See raw diff