Publish petitgpt research-v1 native alpha075 model and scoped documentation
Browse files- DOCUMENTATION_LICENSE.md +9 -0
- LICENSE +202 -0
- MODEL_PROVENANCE.json +26 -0
- README.md +123 -0
- RUN_GUIDE.md +125 -0
- SHA256SUMS +21 -0
- SOURCE_NOTICE.md +39 -0
- THIRD_PARTY_NOTICES.md +5 -0
- config.json +13 -0
- inference.py +63 -0
- model.safetensors +3 -0
- requirements-inference-tested.txt +5 -0
- special_tokens_map.json +12 -0
- src/__init__.py +0 -0
- src/accepted_generate.py +18 -0
- src/chat_template.py +507 -0
- src/model.py +470 -0
- src/special_tokens.py +138 -0
- tables/ASSISTANT_RESULTS_VERSIONED.csv +145 -0
- tables/PRETRAIN_SOURCE_MIXTURE.csv +15 -0
- tables/PUBLIC_BENCHMARK_RESULTS.csv +7 -0
- tokenizer.json +0 -0
DOCUMENTATION_LICENSE.md
ADDED
|
@@ -0,0 +1,9 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Licence scope
|
| 2 |
+
|
| 3 |
+
Copyright 2026 Yang Qi. Owner-controlled code and the selected model/tokenizer are licensed under the standard Apache License 2.0 in LICENSE, only for rights Yang Qi is entitled to grant.
|
| 4 |
+
|
| 5 |
+
As an explicit exception to the root code licence, author-written reports and documentation (including README, native run guide and versioned report) are licensed under Creative Commons Attribution 4.0 International (CC BY 4.0): https://creativecommons.org/licenses/by/4.0/ and https://creativecommons.org/licenses/by/4.0/legalcode.en . Attribute Yang Qi and petitgpt, link the licence, and indicate changes. Existing third-party content and notices retain their applicable terms; they are not relicensed.
|
| 6 |
+
|
| 7 |
+
Research describes intended use; it adds no noncommercial or research-only restriction to Apache-licensed artifacts. This grant was approved by the owner through the explicit research-release execution instruction. Earlier PENDING_OWNER_DECISION records remain historical evidence; this does not claim an earlier licence choice.
|
| 8 |
+
|
| 9 |
+
Source metadata is not rights clearance. Weights are not the raw corpus; neither automatic inheritance nor automatic non-application of all dataset terms is asserted. No infringement guarantee or legal certification is given. A disclaimer does not replace applicable permission.
|
LICENSE
ADDED
|
@@ -0,0 +1,202 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
|
| 2 |
+
Apache License
|
| 3 |
+
Version 2.0, January 2004
|
| 4 |
+
http://www.apache.org/licenses/
|
| 5 |
+
|
| 6 |
+
TERMS AND CONDITIONS FOR USE, REPRODUCTION, AND DISTRIBUTION
|
| 7 |
+
|
| 8 |
+
1. Definitions.
|
| 9 |
+
|
| 10 |
+
"License" shall mean the terms and conditions for use, reproduction,
|
| 11 |
+
and distribution as defined by Sections 1 through 9 of this document.
|
| 12 |
+
|
| 13 |
+
"Licensor" shall mean the copyright owner or entity authorized by
|
| 14 |
+
the copyright owner that is granting the License.
|
| 15 |
+
|
| 16 |
+
"Legal Entity" shall mean the union of the acting entity and all
|
| 17 |
+
other entities that control, are controlled by, or are under common
|
| 18 |
+
control with that entity. For the purposes of this definition,
|
| 19 |
+
"control" means (i) the power, direct or indirect, to cause the
|
| 20 |
+
direction or management of such entity, whether by contract or
|
| 21 |
+
otherwise, or (ii) ownership of fifty percent (50%) or more of the
|
| 22 |
+
outstanding shares, or (iii) beneficial ownership of such entity.
|
| 23 |
+
|
| 24 |
+
"You" (or "Your") shall mean an individual or Legal Entity
|
| 25 |
+
exercising permissions granted by this License.
|
| 26 |
+
|
| 27 |
+
"Source" form shall mean the preferred form for making modifications,
|
| 28 |
+
including but not limited to software source code, documentation
|
| 29 |
+
source, and configuration files.
|
| 30 |
+
|
| 31 |
+
"Object" form shall mean any form resulting from mechanical
|
| 32 |
+
transformation or translation of a Source form, including but
|
| 33 |
+
not limited to compiled object code, generated documentation,
|
| 34 |
+
and conversions to other media types.
|
| 35 |
+
|
| 36 |
+
"Work" shall mean the work of authorship, whether in Source or
|
| 37 |
+
Object form, made available under the License, as indicated by a
|
| 38 |
+
copyright notice that is included in or attached to the work
|
| 39 |
+
(an example is provided in the Appendix below).
|
| 40 |
+
|
| 41 |
+
"Derivative Works" shall mean any work, whether in Source or Object
|
| 42 |
+
form, that is based on (or derived from) the Work and for which the
|
| 43 |
+
editorial revisions, annotations, elaborations, or other modifications
|
| 44 |
+
represent, as a whole, an original work of authorship. For the purposes
|
| 45 |
+
of this License, Derivative Works shall not include works that remain
|
| 46 |
+
separable from, or merely link (or bind by name) to the interfaces of,
|
| 47 |
+
the Work and Derivative Works thereof.
|
| 48 |
+
|
| 49 |
+
"Contribution" shall mean any work of authorship, including
|
| 50 |
+
the original version of the Work and any modifications or additions
|
| 51 |
+
to that Work or Derivative Works thereof, that is intentionally
|
| 52 |
+
submitted to Licensor for inclusion in the Work by the copyright owner
|
| 53 |
+
or by an individual or Legal Entity authorized to submit on behalf of
|
| 54 |
+
the copyright owner. For the purposes of this definition, "submitted"
|
| 55 |
+
means any form of electronic, verbal, or written communication sent
|
| 56 |
+
to the Licensor or its representatives, including but not limited to
|
| 57 |
+
communication on electronic mailing lists, source code control systems,
|
| 58 |
+
and issue tracking systems that are managed by, or on behalf of, the
|
| 59 |
+
Licensor for the purpose of discussing and improving the Work, but
|
| 60 |
+
excluding communication that is conspicuously marked or otherwise
|
| 61 |
+
designated in writing by the copyright owner as "Not a Contribution."
|
| 62 |
+
|
| 63 |
+
"Contributor" shall mean Licensor and any individual or Legal Entity
|
| 64 |
+
on behalf of whom a Contribution has been received by Licensor and
|
| 65 |
+
subsequently incorporated within the Work.
|
| 66 |
+
|
| 67 |
+
2. Grant of Copyright License. Subject to the terms and conditions of
|
| 68 |
+
this License, each Contributor hereby grants to You a perpetual,
|
| 69 |
+
worldwide, non-exclusive, no-charge, royalty-free, irrevocable
|
| 70 |
+
copyright license to reproduce, prepare Derivative Works of,
|
| 71 |
+
publicly display, publicly perform, sublicense, and distribute the
|
| 72 |
+
Work and such Derivative Works in Source or Object form.
|
| 73 |
+
|
| 74 |
+
3. Grant of Patent License. Subject to the terms and conditions of
|
| 75 |
+
this License, each Contributor hereby grants to You a perpetual,
|
| 76 |
+
worldwide, non-exclusive, no-charge, royalty-free, irrevocable
|
| 77 |
+
(except as stated in this section) patent license to make, have made,
|
| 78 |
+
use, offer to sell, sell, import, and otherwise transfer the Work,
|
| 79 |
+
where such license applies only to those patent claims licensable
|
| 80 |
+
by such Contributor that are necessarily infringed by their
|
| 81 |
+
Contribution(s) alone or by combination of their Contribution(s)
|
| 82 |
+
with the Work to which such Contribution(s) was submitted. If You
|
| 83 |
+
institute patent litigation against any entity (including a
|
| 84 |
+
cross-claim or counterclaim in a lawsuit) alleging that the Work
|
| 85 |
+
or a Contribution incorporated within the Work constitutes direct
|
| 86 |
+
or contributory patent infringement, then any patent licenses
|
| 87 |
+
granted to You under this License for that Work shall terminate
|
| 88 |
+
as of the date such litigation is filed.
|
| 89 |
+
|
| 90 |
+
4. Redistribution. You may reproduce and distribute copies of the
|
| 91 |
+
Work or Derivative Works thereof in any medium, with or without
|
| 92 |
+
modifications, and in Source or Object form, provided that You
|
| 93 |
+
meet the following conditions:
|
| 94 |
+
|
| 95 |
+
(a) You must give any other recipients of the Work or
|
| 96 |
+
Derivative Works a copy of this License; and
|
| 97 |
+
|
| 98 |
+
(b) You must cause any modified files to carry prominent notices
|
| 99 |
+
stating that You changed the files; and
|
| 100 |
+
|
| 101 |
+
(c) You must retain, in the Source form of any Derivative Works
|
| 102 |
+
that You distribute, all copyright, patent, trademark, and
|
| 103 |
+
attribution notices from the Source form of the Work,
|
| 104 |
+
excluding those notices that do not pertain to any part of
|
| 105 |
+
the Derivative Works; and
|
| 106 |
+
|
| 107 |
+
(d) If the Work includes a "NOTICE" text file as part of its
|
| 108 |
+
distribution, then any Derivative Works that You distribute must
|
| 109 |
+
include a readable copy of the attribution notices contained
|
| 110 |
+
within such NOTICE file, excluding those notices that do not
|
| 111 |
+
pertain to any part of the Derivative Works, in at least one
|
| 112 |
+
of the following places: within a NOTICE text file distributed
|
| 113 |
+
as part of the Derivative Works; within the Source form or
|
| 114 |
+
documentation, if provided along with the Derivative Works; or,
|
| 115 |
+
within a display generated by the Derivative Works, if and
|
| 116 |
+
wherever such third-party notices normally appear. The contents
|
| 117 |
+
of the NOTICE file are for informational purposes only and
|
| 118 |
+
do not modify the License. You may add Your own attribution
|
| 119 |
+
notices within Derivative Works that You distribute, alongside
|
| 120 |
+
or as an addendum to the NOTICE text from the Work, provided
|
| 121 |
+
that such additional attribution notices cannot be construed
|
| 122 |
+
as modifying the License.
|
| 123 |
+
|
| 124 |
+
You may add Your own copyright statement to Your modifications and
|
| 125 |
+
may provide additional or different license terms and conditions
|
| 126 |
+
for use, reproduction, or distribution of Your modifications, or
|
| 127 |
+
for any such Derivative Works as a whole, provided Your use,
|
| 128 |
+
reproduction, and distribution of the Work otherwise complies with
|
| 129 |
+
the conditions stated in this License.
|
| 130 |
+
|
| 131 |
+
5. Submission of Contributions. Unless You explicitly state otherwise,
|
| 132 |
+
any Contribution intentionally submitted for inclusion in the Work
|
| 133 |
+
by You to the Licensor shall be under the terms and conditions of
|
| 134 |
+
this License, without any additional terms or conditions.
|
| 135 |
+
Notwithstanding the above, nothing herein shall supersede or modify
|
| 136 |
+
the terms of any separate license agreement you may have executed
|
| 137 |
+
with Licensor regarding such Contributions.
|
| 138 |
+
|
| 139 |
+
6. Trademarks. This License does not grant permission to use the trade
|
| 140 |
+
names, trademarks, service marks, or product names of the Licensor,
|
| 141 |
+
except as required for reasonable and customary use in describing the
|
| 142 |
+
origin of the Work and reproducing the content of the NOTICE file.
|
| 143 |
+
|
| 144 |
+
7. Disclaimer of Warranty. Unless required by applicable law or
|
| 145 |
+
agreed to in writing, Licensor provides the Work (and each
|
| 146 |
+
Contributor provides its Contributions) on an "AS IS" BASIS,
|
| 147 |
+
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or
|
| 148 |
+
implied, including, without limitation, any warranties or conditions
|
| 149 |
+
of TITLE, NON-INFRINGEMENT, MERCHANTABILITY, or FITNESS FOR A
|
| 150 |
+
PARTICULAR PURPOSE. You are solely responsible for determining the
|
| 151 |
+
appropriateness of using or redistributing the Work and assume any
|
| 152 |
+
risks associated with Your exercise of permissions under this License.
|
| 153 |
+
|
| 154 |
+
8. Limitation of Liability. In no event and under no legal theory,
|
| 155 |
+
whether in tort (including negligence), contract, or otherwise,
|
| 156 |
+
unless required by applicable law (such as deliberate and grossly
|
| 157 |
+
negligent acts) or agreed to in writing, shall any Contributor be
|
| 158 |
+
liable to You for damages, including any direct, indirect, special,
|
| 159 |
+
incidental, or consequential damages of any character arising as a
|
| 160 |
+
result of this License or out of the use or inability to use the
|
| 161 |
+
Work (including but not limited to damages for loss of goodwill,
|
| 162 |
+
work stoppage, computer failure or malfunction, or any and all
|
| 163 |
+
other commercial damages or losses), even if such Contributor
|
| 164 |
+
has been advised of the possibility of such damages.
|
| 165 |
+
|
| 166 |
+
9. Accepting Warranty or Additional Liability. While redistributing
|
| 167 |
+
the Work or Derivative Works thereof, You may choose to offer,
|
| 168 |
+
and charge a fee for, acceptance of support, warranty, indemnity,
|
| 169 |
+
or other liability obligations and/or rights consistent with this
|
| 170 |
+
License. However, in accepting such obligations, You may act only
|
| 171 |
+
on Your own behalf and on Your sole responsibility, not on behalf
|
| 172 |
+
of any other Contributor, and only if You agree to indemnify,
|
| 173 |
+
defend, and hold each Contributor harmless for any liability
|
| 174 |
+
incurred by, or claims asserted against, such Contributor by reason
|
| 175 |
+
of your accepting any such warranty or additional liability.
|
| 176 |
+
|
| 177 |
+
END OF TERMS AND CONDITIONS
|
| 178 |
+
|
| 179 |
+
APPENDIX: How to apply the Apache License to your work.
|
| 180 |
+
|
| 181 |
+
To apply the Apache License to your work, attach the following
|
| 182 |
+
boilerplate notice, with the fields enclosed by brackets "[]"
|
| 183 |
+
replaced with your own identifying information. (Don't include
|
| 184 |
+
the brackets!) The text should be enclosed in the appropriate
|
| 185 |
+
comment syntax for the file format. We also recommend that a
|
| 186 |
+
file or class name and description of purpose be included on the
|
| 187 |
+
same "printed page" as the copyright notice for easier
|
| 188 |
+
identification within third-party archives.
|
| 189 |
+
|
| 190 |
+
Copyright [yyyy] [name of copyright owner]
|
| 191 |
+
|
| 192 |
+
Licensed under the Apache License, Version 2.0 (the "License");
|
| 193 |
+
you may not use this file except in compliance with the License.
|
| 194 |
+
You may obtain a copy of the License at
|
| 195 |
+
|
| 196 |
+
http://www.apache.org/licenses/LICENSE-2.0
|
| 197 |
+
|
| 198 |
+
Unless required by applicable law or agreed to in writing, software
|
| 199 |
+
distributed under the License is distributed on an "AS IS" BASIS,
|
| 200 |
+
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
| 201 |
+
See the License for the specific language governing permissions and
|
| 202 |
+
limitations under the License.
|
MODEL_PROVENANCE.json
ADDED
|
@@ -0,0 +1,26 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"project": "petitgpt",
|
| 3 |
+
"author": "Yang Qi",
|
| 4 |
+
"checkpoint": "alpha075",
|
| 5 |
+
"lineage": [
|
| 6 |
+
"Base",
|
| 7 |
+
"P2 step750",
|
| 8 |
+
"P3 step320",
|
| 9 |
+
"interpolation alpha=0.75"
|
| 10 |
+
],
|
| 11 |
+
"excluded_updates": [
|
| 12 |
+
"DeepSeek-response-KD",
|
| 13 |
+
"unified Base-SFT",
|
| 14 |
+
"DPO",
|
| 15 |
+
"soft-KD",
|
| 16 |
+
"LoRA"
|
| 17 |
+
],
|
| 18 |
+
"source_checkpoint_sha256": "1da85cc329d55e92dacf51c36623779558c4c6c9a39d78a34a064f61fcddbe97",
|
| 19 |
+
"model_safetensors_sha256": "4396efb7a52b047e7fdf513e46d1b401dfc70582d3aca1f9cb5a07e97d426ef1",
|
| 20 |
+
"tokenizer_sha256": "d8f84df58928023edebd809e152b3b38a0dac53b9f887bd2455f427661e9b9ce",
|
| 21 |
+
"accepted_archive_sha256": "0c844962c6ebfcc1cd6b17cdb19c05d88936fa280f0173c080ff864bb54dca1e",
|
| 22 |
+
"unique_parameters": 124635456,
|
| 23 |
+
"release_model_operations": 0,
|
| 24 |
+
"license": "Apache-2.0 only for author-controlled rights; see notices",
|
| 25 |
+
"tokenizer_source_linkage": "Digest-linked to six pinned corpus releases; PES2O and StackExchange not in tokenizer corpus"
|
| 26 |
+
}
|
README.md
ADDED
|
@@ -0,0 +1,123 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: apache-2.0
|
| 3 |
+
tags:
|
| 4 |
+
- petitgpt
|
| 5 |
+
- native-pytorch
|
| 6 |
+
- research
|
| 7 |
+
---
|
| 8 |
+
|
| 9 |
+
# petitgpt
|
| 10 |
+
|
| 11 |
+
Author: Yang Qi. Selected checkpoint: alpha075.
|
| 12 |
+
|
| 13 |
+
## Identity
|
| 14 |
+
|
| 15 |
+
| Field | Value |
|
| 16 |
+
|---|---|
|
| 17 |
+
| Status | accepted native research inference artifact |
|
| 18 |
+
| Unique parameters | 124,635,456 (124.6M) |
|
| 19 |
+
| Layers / width / FFN | 30 / 576 / 1536 |
|
| 20 |
+
| Attention | 9 query heads, 3 key/value heads (GQA), head dim 64 |
|
| 21 |
+
| Vocabulary / context | 32,000 / 2,048 |
|
| 22 |
+
| Embeddings | tied input/output |
|
| 23 |
+
| Normalization | RMSNorm, epsilon 1e-6 |
|
| 24 |
+
| Positions | RoPE, theta 10000, full head rotation |
|
| 25 |
+
| Dropout | 0.0 |
|
| 26 |
+
| Stored weights | FP32 |
|
| 27 |
+
|
| 28 |
+
Complete checkpoint-derived settings ship in the bundle's `config.json`. The three core modules (model, chat template, token contract) are byte-identical to the project originals. Checkpoint and archive hashes are in MODEL_PROVENANCE.json.
|
| 29 |
+
|
| 30 |
+
## Provenance
|
| 31 |
+
|
| 32 |
+
The selected weights are a **parameter interpolation**, not a training step:
|
| 33 |
+
|
| 34 |
+
> `theta = theta_P2_step750 + 0.75 · (theta_P3_step320 − theta_P2_step750)`
|
| 35 |
+
|
| 36 |
+
Ancestry: tokenizer release → Stage A pretraining (steps 0–38,146) → Stage B continued pretraining (steps 38,146–49,590, exact full-state resume, accepted as Base) → P2 concise-instruction SFT (750 updates, weights-only initialization from Base) → P3 basic-instruction adaptation (parent B is **step 320**, not the step-640 endpoint) → this interpolation, executed with **zero optimizer updates and zero backward passes** → a numerically unchanged FP32 export.
|
| 37 |
+
|
| 38 |
+
This checkpoint **does not contain** later DeepSeek-response-KD, unified Base-SFT, DPO, soft-KD or LoRA branch updates. Several later branches were initialized *from* it (a one-pass behaviour mix, a DPO pilot, a chosen-answer CE control, and two response-distillation runs); others were not — a unified SFT curve started from the accepted pretrained Base, a loss-allocation A/B split from that curve's step 403, and a shared-tokenizer soft-KD lab ran entirely on external models. A preference-data build used this model's generations but produced no checkpoint. Branch exposures must not be summed into this model's training history.
|
| 39 |
+
|
| 40 |
+
## Training data
|
| 41 |
+
|
| 42 |
+
Pretraining consumed 13,000,005,634 retained packed tokens over 13,755,731 documents, of which the optimizer stepped over 12,999,720,960 model-input positions, one exposure per block, with no replay of the executed token-position traversal.
|
| 43 |
+
|
| 44 |
+
**Stage A (10,000,003,234 selected serialized tokens, 4 sources):** FineWeb-Edu dedup 71.11%, DCLM-Edu 20.32%, Wikipedia (FineWiki EN) 5.08%, Python-Edu 3.50%.
|
| 45 |
+
|
| 46 |
+
**Stage B (3,000,004,240 selected serialized tokens, 7 sources):** FineWeb-Edu dedup 40.10%, DCLM-Edu 22.92%, structured tutorial content 11.46%, Python-Edu 8.33%, Wikipedia 5.73%, PES2O 5.73%, StackExchange 5.73%.
|
| 47 |
+
|
| 48 |
+
Upstream datasets, pinned revisions and the licence string recorded at each pinned revision are in `tables/PRETRAIN_SOURCE_MIXTURE.csv`. Those recorded strings are evidence of what the builder captured at that revision; they are not a legal determination, are not asserted to be today's terms, and do not by themselves determine the licence of trained weights.
|
| 49 |
+
|
| 50 |
+
Post-training used seven instruction subsets. These are values of a row-level `source` column inside one pinned collection, `HuggingFaceTB/smol-smoltalk` at revision `f73fe857d519ff6ac5af2ea67c4d3834da7b8bcc`, config `default`, train split, established by digest joins through the project's own census and cleanup records. The publisher's card at that pinned revision carries a flat Apache-2.0 badge; its parent collection limits that grant to four newly generated subsets and refers readers to the original dataset for each incorporated public dataset. Four of the seven labels correspond to the newly generated subsets. Of the three incorporated components, one declares `apache-2.0`, one declares `odc-by`, and one declares no licence in its card metadata. Component notices were read from current publisher pages, not from revisions contemporaneous with this training run, and **no component revision is established**. No licence determination is made here.
|
| 51 |
+
|
| 52 |
+
## Intended use, and use it is not intended for
|
| 53 |
+
|
| 54 |
+
**Intended:** research and engineering study of a small from-scratch language model — reproducing the recorded measurements, inspecting the pipeline, and analysing failure modes.
|
| 55 |
+
|
| 56 |
+
**Not intended:** a general assistant, a production system, anything safety- or correctness-certified, or a source of factual answers. Do not execute code it generates without independent review.
|
| 57 |
+
|
| 58 |
+
**Not evaluated at all:** long-context work, multilingual behaviour, tool use, extended multi-turn dialogue, safety and refusal behaviour, factual currency, retrieval, and any public generative benchmark.
|
| 59 |
+
|
| 60 |
+
## Evaluation — public multiple-choice likelihood
|
| 61 |
+
|
| 62 |
+
Frozen FP32 results, copied byte-identically from the accepted measurement and **not recomputed**:
|
| 63 |
+
|
| 64 |
+
| Dataset | Split / documents | acc | acc_norm |
|
| 65 |
+
|---|---|---|---|
|
| 66 |
+
| ARC-Easy | test / 2,376 | 1372/2376 = 0.5774410774410774 | 1244/2376 = 0.5235690235690236 |
|
| 67 |
+
| PIQA | validation / 1,838 | 1167/1838 = 0.6349292709466812 | 1145/1838 = 0.6229597388465724 |
|
| 68 |
+
|
| 69 |
+
Protocol: zero-shot raw `Question: …\nAnswer:` completion scored by candidate-answer likelihood. No chat template, no role tokens, no few-shot examples, no BOS insertion, no scored EOS, no generation, no cleanup. FP32 parameters and forward with autocast disabled, TF32 off for matmul and cuDNN, MATH SDPA, batch size 1 unpadded, no KV cache, no compile. `acc` is the first argmax of summed continuation log-likelihood; `acc_norm` divides by `len()` of the **original** answer text in Unicode characters, not tokenizer length; ties take the first index.
|
| 70 |
+
|
| 71 |
+
Comparators measured under the identical protocol on the same rows: SmolLM-135M-Instruct 0.4924 / 0.6708 and SmolLM2-135M-Instruct 0.5400 / 0.6670 (acc, ARC-Easy / PIQA).
|
| 72 |
+
|
| 73 |
+
**Qualifications.** The evaluator is a native protocol-compatible implementation pinned to a specific lm-evaluation-harness commit; the harness package was not installed and a full installed-harness run is not claimed. Both datasets are prior project diagnostics with no contamination audit — they are **not untouched final tests**. Training data, compute, tokenizers and architectures are unmatched across the three models. This is a protocol-bounded descriptive comparison; no significance test was run. Multiple-choice accuracy does not establish free-generation reliability.
|
| 74 |
+
|
| 75 |
+
## Evaluation — historical full-answer assistant review
|
| 76 |
+
|
| 77 |
+
A separate evaluation family scored generated text across 189 prompts per model (567 answers, 565 distinct prompt/output/contract units) over three models. **Two named versions exist and must not be combined in one table:** `assistant_review_fable_v1` (the original 567 final records) and `assistant_owner_clarification_4_v1` (four explicit final-content decisions, every other axis preserved). The version shown below is `assistant_owner_clarification_4_v1`.
|
| 78 |
+
|
| 79 |
+
| Slice | Content / joint (true / false / unknown) | Other axes |
|
| 80 |
+
|---|---|---|
|
| 81 |
+
| old_qa41 | 6 / 33 / 2 | format 0/1 |
|
| 82 |
+
| old_practical38 | 10 / 26 / 2 | explicit format 20/21 |
|
| 83 |
+
| old_python14 | 0 / 14 / 0 | interface 11/14; finite execution `not_recorded` |
|
| 84 |
+
| new_natural64 | 3 / 61 / 0 | format 7/18 |
|
| 85 |
+
| new_python32 | 0 / 32 / 0 | interface 31/32 |
|
| 86 |
+
|
| 87 |
+
Practical joint bounds: **26/81 .. 10/27** (= 52/162 .. 60/162). This denominator is **27 equally weighted dialogue groups over 38 scored turns** — rows are averaged inside a group first — not an ordinary row average. These are exact unknown-retention bounds, **not confidence intervals**.
|
| 88 |
+
|
| 89 |
+
**Qualifications.** Judgments are model-assisted, not human adjudication. The chronology was: an initial pass over the 565 units with model metadata masked and its own recorded limitations; then a pass with the mapping visible that produced 17 consistency edits; then four owner clarifications forming a separate version. It was therefore neither strictly blinded throughout nor fully label-visible throughout. Development sets were reused across many runs. `not_recorded` means the field is unavailable in this imported view — it does not establish that no historical function test was ever run. Unsupported-builtin results stay unknown and are never converted into demonstrated failures.
|
| 90 |
+
|
| 91 |
+
In these specific Python diagnostics the model produced a correct function **interface** in 42 of 46 prompts and a correct **whole answer** in 0 of 46. This does not establish that it can never write correct code.
|
| 92 |
+
|
| 93 |
+
Ordinary QA and complete natural-task generation remained limited across the evaluated configurations. Individual results differ by suite and label version and are reported with their source rather than reduced to a cross-version maximum; see the technical report §10.1.
|
| 94 |
+
|
| 95 |
+
## Export parity
|
| 96 |
+
|
| 97 |
+
Eight frozen fixture pairs (four `bf16_native`, four `fp32_math`) matched exactly on prompt IDs, boundaries, full-shape logits, greedy output IDs and stop reason, with a maximum absolute logit difference of **0 within each profile**. All 213 named state entries and 60 non-persistent rotary buffers matched after strict load and safetensors reload.
|
| 98 |
+
|
| 99 |
+
This is a **numerical parity check between source and export under the same profile**. It is not a quality test, not a semantic evaluation, and it does not assert that the two profiles agree with each other. Local import closure was demonstrated once in a fresh isolated process on the measured environment; that is **not** a clean-machine, CPU, cross-hardware or fresh-installation test.
|
| 100 |
+
|
| 101 |
+
## Format support
|
| 102 |
+
|
| 103 |
+
Native PyTorch CUDA inference only. **No** Transformers `AutoModel`, GGUF, ONNX, vLLM or llama.cpp compatibility is implemented or tested. `special_tokens_map.json` is descriptive native metadata. A CUDA GPU is required.
|
| 104 |
+
|
| 105 |
+
## Licence and distribution
|
| 106 |
+
|
| 107 |
+
Copyright 2026 Yang Qi. Owner-controlled code and the selected model/tokenizer are licensed under the standard Apache License 2.0 in LICENSE, only for rights Yang Qi is entitled to grant.
|
| 108 |
+
|
| 109 |
+
As an explicit exception to the root code licence, author-written reports and documentation (including README, native run guide and versioned report) are licensed under Creative Commons Attribution 4.0 International (CC BY 4.0): https://creativecommons.org/licenses/by/4.0/ and https://creativecommons.org/licenses/by/4.0/legalcode.en . Attribute Yang Qi and petitgpt, link the licence, and indicate changes. Existing third-party content and notices retain their applicable terms; they are not relicensed.
|
| 110 |
+
|
| 111 |
+
Research describes intended use; it adds no noncommercial or research-only restriction to Apache-licensed artifacts. This grant was approved by the owner through the explicit research-release execution instruction. Earlier PENDING_OWNER_DECISION records remain historical evidence; this does not claim an earlier licence choice.
|
| 112 |
+
|
| 113 |
+
Source metadata is not rights clearance. Weights are not the raw corpus; neither automatic inheritance nor automatic non-application of all dataset terms is asserted. No infringement guarantee or legal certification is given. A disclaimer does not replace applicable permission.
|
| 114 |
+
|
| 115 |
+
## Recorded runtime
|
| 116 |
+
|
| 117 |
+
Python 3.10.12, torch 2.11.0+cu126, numpy 2.2.6, tokenizers 0.22.2, safetensors 0.8.0, NVIDIA GeForce RTX 4090. Matching versions do not guarantee bit-identical results on other hardware or untested software; recorded driver versions differ across project phases, which is a recorded difference rather than a resolved equivalence.
|
| 118 |
+
|
| 119 |
+
## Files and report
|
| 120 |
+
|
| 121 |
+
See [native run guide](RUN_GUIDE.md), [source notice](SOURCE_NOTICE.md), [third-party notices](THIRD_PARTY_NOTICES.md) and [file manifest](SHA256SUMS).
|
| 122 |
+
|
| 123 |
+
GitHub project: https://github.com/yangqi0/petitgpt . The complete research-v1 report is staged for `docs/petitgpt-v1/TECHNICAL_REPORT.md` there; publication is blocked by missing GitHub authentication. That report is not yet publicly available, and no live report link is claimed.
|
RUN_GUIDE.md
ADDED
|
@@ -0,0 +1,125 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# petitgpt native run guide
|
| 2 |
+
|
| 3 |
+
Author: Yang Qi. Documentation: CC BY 4.0. Download the loose model repository files together, preserving the src/ directory. The exact release file list and hashes are in SHA256SUMS. Model weights, tokenizer, config and executable code retain the accepted export bytes.
|
| 4 |
+
|
| 5 |
+
## 2. Requirements
|
| 6 |
+
|
| 7 |
+
A **CUDA GPU is required** by this CLI. The tested configuration is Python 3.10.12 with:
|
| 8 |
+
|
| 9 |
+
```
|
| 10 |
+
torch==2.11.0+cu126
|
| 11 |
+
numpy==2.2.6
|
| 12 |
+
tokenizers==0.22.2
|
| 13 |
+
safetensors==0.8.0
|
| 14 |
+
```
|
| 15 |
+
|
| 16 |
+
on an NVIDIA GeForce RTX 4090 (driver 580.178.04). Prepare the environment separately, from locally supplied wheels:
|
| 17 |
+
|
| 18 |
+
```sh
|
| 19 |
+
python -m pip install --no-index --find-links /path/to/local/wheelhouse \
|
| 20 |
+
-r /path/to/bundle/requirements-inference-tested.txt
|
| 21 |
+
```
|
| 22 |
+
|
| 23 |
+
No packages were installed during the export itself. Matching versions do not guarantee bit-identical results on arbitrary hardware or untested software.
|
| 24 |
+
|
| 25 |
+
## 3. Command line
|
| 26 |
+
|
| 27 |
+
Replace `/path/to/bundle` with the real extracted location.
|
| 28 |
+
|
| 29 |
+
```sh
|
| 30 |
+
python /path/to/bundle/inference.py \
|
| 31 |
+
--model-directory /path/to/bundle \
|
| 32 |
+
--prompt "Say hello in one sentence." \
|
| 33 |
+
--profile bf16_native \
|
| 34 |
+
--max-new-tokens 32
|
| 35 |
+
|
| 36 |
+
python /path/to/bundle/inference.py \
|
| 37 |
+
--model-directory /path/to/bundle \
|
| 38 |
+
--messages-json /path/to/messages.json \
|
| 39 |
+
--profile fp32_math \
|
| 40 |
+
--max-new-tokens 32
|
| 41 |
+
```
|
| 42 |
+
|
| 43 |
+
A synthetic `messages.json`:
|
| 44 |
+
|
| 45 |
+
```json
|
| 46 |
+
[{"role":"system","content":"Use short sentences."},
|
| 47 |
+
{"role":"user","content":"My name is Lin."},
|
| 48 |
+
{"role":"assistant","content":"Hello, Lin."},
|
| 49 |
+
{"role":"user","content":"What name did I give you?"}]
|
| 50 |
+
```
|
| 51 |
+
|
| 52 |
+
> **Illustrative and unexecuted.** The greeting above is an authored example showing command syntax only. It was written for this guide, was not run, and is not a demonstration of model quality. Do not put frozen evaluation prompts or their outputs into a public demo.
|
| 53 |
+
|
| 54 |
+
## 4. Python API
|
| 55 |
+
|
| 56 |
+
With the extracted bundle directory on `sys.path`:
|
| 57 |
+
|
| 58 |
+
```python
|
| 59 |
+
from inference import load_bundle, generate
|
| 60 |
+
|
| 61 |
+
model, tokenizer = load_bundle("/path/to/bundle")
|
| 62 |
+
model = model.to("cuda").eval()
|
| 63 |
+
result = generate(
|
| 64 |
+
model, tokenizer,
|
| 65 |
+
[{"role": "user", "content": "Say hello."}],
|
| 66 |
+
cap=32,
|
| 67 |
+
profile="fp32_math",
|
| 68 |
+
)
|
| 69 |
+
print(result["output_text_raw_including_terminal_eos"])
|
| 70 |
+
```
|
| 71 |
+
|
| 72 |
+
## 5. Input contract
|
| 73 |
+
|
| 74 |
+
`messages` must be a JSON array of objects with **exactly** the fields `role` and `content`. A conversation is an optional initial system turn followed by alternating user/assistant turns, **ending in user**. A plain `--prompt` becomes one user message.
|
| 75 |
+
|
| 76 |
+
Rejected, by design, with an explicit error rather than a repair:
|
| 77 |
+
|
| 78 |
+
| Case | Recorded rejection |
|
| 79 |
+
|---|---|
|
| 80 |
+
| Missing `content` | `Each message must contain only role and content` |
|
| 81 |
+
| Unsupported role (e.g. `tool`) | `message 0 has invalid role 'tool'` |
|
| 82 |
+
| Any extra message field | `Each message must contain only role and content` |
|
| 83 |
+
| Conversation ending in `assistant` | `chat must end with a non-empty user turn; got 'assistant'` |
|
| 84 |
+
| Prompt + budget over 2,048 | `Context overflow: prompt plus token budget exceeds 2048` |
|
| 85 |
+
| `max_new_tokens` outside 1..384 | `max_new_tokens must be 1..384` |
|
| 86 |
+
|
| 87 |
+
`default_system=None`: a supplied system turn and full history are retained, with **no injected default and no text normalization**. Literal special-token spellings inside content — for example the literal spelling `[EOS]` — are encoded as ordinary text and cannot inject control IDs.
|
| 88 |
+
|
| 89 |
+
Native token structure:
|
| 90 |
+
|
| 91 |
+
```
|
| 92 |
+
[BOS] <|system|> system <|user|> user <|assistant|> assistant [EOS] … <|user|> user <|assistant|>
|
| 93 |
+
```
|
| 94 |
+
|
| 95 |
+
The system segment is omitted when absent. `[BOS]` occurs once; `[EOS]` closes completed assistant turns; no duplicate assistant prefix is added. IDs are fixed: `[PAD]=0`, `[UNK]=1`, `[BOS]=2`, `[EOS]=3`, `<|system|>=4`, `<|user|>=5`, `<|assistant|>=6`.
|
| 96 |
+
|
| 97 |
+
## 6. Precision profiles
|
| 98 |
+
|
| 99 |
+
Both profiles store FP32 parameters. They differ only in the forward numerical path, and **they are not asserted to agree with each other**.
|
| 100 |
+
|
| 101 |
+
| | `bf16_native` | `fp32_math` |
|
| 102 |
+
|---|---|---|
|
| 103 |
+
| Forward | CUDA BF16 autocast | FP32, autocast off |
|
| 104 |
+
| matmul TF32 | off | off |
|
| 105 |
+
| cuDNN TF32 | **on** | off |
|
| 106 |
+
| `float32_matmul_precision` | highest | highest |
|
| 107 |
+
| SDPA backend | native backends enabled | MATH |
|
| 108 |
+
|
| 109 |
+
## 7. Decoding and stopping
|
| 110 |
+
|
| 111 |
+
Greedy only: `temperature=0`, `top_k=0`, `top_p=1`, `EOS=3`. Default `max_new_tokens` is **384**; an explicit lower integer in `1..384` is supported. Prompt length plus budget must fit 2,048 — there is no cropping.
|
| 112 |
+
|
| 113 |
+
Output JSON preserves the generated token IDs, the raw text **including the terminal EOS**, and a stop reason of `eos` or `max_new_tokens`. There is no answer cleanup, no fact fixing, no best-of-N, and no retry. A token cap can truncate an answer mid-sentence; that is the recorded behaviour, not a defect.
|
| 114 |
+
|
| 115 |
+
## 8. Format support and limits
|
| 116 |
+
|
| 117 |
+
Native PyTorch CUDA inference only. There is **no implemented or tested** Transformers `AutoModel`, GGUF, ONNX, vLLM or llama.cpp path. `special_tokens_map.json` is descriptive native metadata, not a loader contract.
|
| 118 |
+
|
| 119 |
+
Local import closure was demonstrated once, in a fresh process with a temporary working directory and an empty `PYTHONPATH`, under an audit hook that denied network access and out-of-bundle repository access; no denied access occurred. That establishes closure **on the measured environment only**. It is not a clean-machine test, not a CPU test, not a cross-hardware test, and not a fresh dependency-installation test.
|
| 120 |
+
|
| 121 |
+
The recorded export parity is likewise **profile-specific**: within each of the two profiles the export reproduced its source exactly, and no claim is made that the two profiles agree with each other, nor that any Transformers, GGUF, ONNX, vLLM or llama.cpp path exists.
|
| 122 |
+
|
| 123 |
+
## 9. Before you use output
|
| 124 |
+
|
| 125 |
+
This is a research artifact, not a safety- or correctness-certified assistant. In the recorded Python diagnostics the model produced a valid function interface far more often than a correct whole answer. **Do not execute generated code without independent review.** Backend details and the bounded parity evidence live in a separate private evidence archive and are not part of this bundle.
|
SHA256SUMS
ADDED
|
@@ -0,0 +1,21 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
74322f799773266d15acf47c40165d2c7e5bf9df8e5ca7a076dd98e51ee41656 DOCUMENTATION_LICENSE.md
|
| 2 |
+
cfc7749b96f63bd31c3c42b5c471bf756814053e847c10f3eb003417bc523d30 LICENSE
|
| 3 |
+
936b54b619edeb38bc5ff13e01e4bb1682f7d5c6f8fdf78756cc4c4c9f953aaf MODEL_PROVENANCE.json
|
| 4 |
+
d6ce84427abcf0f0a3d1d45475c5282c6e5bd9b5b88d9055ac502698cee66605 README.md
|
| 5 |
+
b6e7c48e4c287517c1c4e0013ed0f8b33f49cae479d185fa7af734e057ab28a4 RUN_GUIDE.md
|
| 6 |
+
46dc9f7c632aa5b7ffeb5678d76dbe03a9f78b70a8d721f8a3d96dda4a0f5876 SOURCE_NOTICE.md
|
| 7 |
+
aeb343a8a11ef0bffa736b360e73273d5a05f48429ae8346a97517155b75c3ef THIRD_PARTY_NOTICES.md
|
| 8 |
+
fe6cd464a09a0605f6ead38867c0497edc71de31fa45446f790e3f9a4f937369 config.json
|
| 9 |
+
17e307631b8e5fd8be8c9d7e992632a6dfbc26708287318fca1259f6d1dd548d inference.py
|
| 10 |
+
4396efb7a52b047e7fdf513e46d1b401dfc70582d3aca1f9cb5a07e97d426ef1 model.safetensors
|
| 11 |
+
dd6b7d954f0153cc6898a9aa0f45901e3727906172e0045d9bc776daf2f7163e requirements-inference-tested.txt
|
| 12 |
+
79bc49c15e40cfc59779b5559d452c058a06f684a2867d97c3c0190933d8801f special_tokens_map.json
|
| 13 |
+
e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855 src/__init__.py
|
| 14 |
+
5c4214f0d2a985ee7b24da431ce1bf521c96ae6f7dfa816ef800b505eb68013c src/accepted_generate.py
|
| 15 |
+
c25068dcb7d3a907dba7b7405d53a2962d91886abcaa25b0d5832a82e8bc821b src/chat_template.py
|
| 16 |
+
2bc9fa8ae16636837c4a2937301a2419d0ac92faa2cc27560dacbd29a5144dc2 src/model.py
|
| 17 |
+
f767b864d7c8e0cb5e2c166c4f019f3f14666dcbf2b944c7909db050e4cf1e96 src/special_tokens.py
|
| 18 |
+
4dd0077babf854389d09623c154844b46f977ca64753f06ecf50662297a2caf5 tables/ASSISTANT_RESULTS_VERSIONED.csv
|
| 19 |
+
b452952e2762f6fe07484e995a7d590f8659c5d5cf60d13e1785cd975a48ace9 tables/PRETRAIN_SOURCE_MIXTURE.csv
|
| 20 |
+
5350ff2d0b30e67645d3b088974f893a9943d9c5f7407f45548917f1c54da297 tables/PUBLIC_BENCHMARK_RESULTS.csv
|
| 21 |
+
d8f84df58928023edebd809e152b3b38a0dac53b9f887bd2455f427661e9b9ce tokenizer.json
|
SOURCE_NOTICE.md
ADDED
|
@@ -0,0 +1,39 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Source notice
|
| 2 |
+
|
| 3 |
+
petitgpt by Yang Qi; selected checkpoint alpha075. Parameter lineage: Base -> P2 step750 -> P3 step320 -> interpolation (0.75 toward P3). Later DeepSeek-response-KD, unified Base-SFT, DPO, soft-KD and LoRA updates are absent.
|
| 4 |
+
|
| 5 |
+
Pretraining source and tokenizer linkage have been established by the existing digest records. The tokenizer corpus uses six of the eight frozen releases; PES2O and StackExchange occur in pretraining, not tokenizer training. Selected source counts are not per-source consumed-token measurements. No new corpus audit is claimed.
|
| 6 |
+
|
| 7 |
+
## Pretraining and tokenizer sources
|
| 8 |
+
|
| 9 |
+
The following terms are builder-recorded metadata at pinned revisions, not independent verification of every licensor's rights. Dataset authors and contributors retain their rights. Exact mixture and transport revisions are preserved in the companion PRETRAIN_SOURCE_MIXTURE.csv.
|
| 10 |
+
|
| 11 |
+
| Dataset / config | Pinned revision | Recorded terms |
|
| 12 |
+
|---|---|---|
|
| 13 |
+
| HuggingFaceTB/smollm-corpus / fineweb-edu-dedup | `3ba9d605774198c5868892d7a8deda78031a781f` | odc-by-1.0 |
|
| 14 |
+
| HuggingFaceTB/dclm-edu / default | `dbad8ad71224482740cd9c9d353591adbf62fe04` | cc-by-4.0 |
|
| 15 |
+
| HuggingFaceFW/finewiki / en | `8bd13e72e6a002407649b3e898535f42ceb1aeb9` | cc-by-sa-4.0 |
|
| 16 |
+
| common-pile/stackv2_edu_filtered / default | `c354dbe88469a1153e97c6a63ac50591849654de` | per-record metadata.license (Software Heritage permissive subset) |
|
| 17 |
+
| HuggingFaceTB/smollm-corpus + HuggingFaceFW/finephrase / cosmopedia-v2 + tutorial | `3ba9d605774198c5868892d7a8deda78031a781f + 78cf4a5ed0099214979c094c963e699c19163838` | odc-by-1.0 (both) |
|
| 18 |
+
| allenai/dolmino-mix-1124 / pes2o | `a319f19eef1e257417b11ea8c30da266ae175557` | odc-by-1.0 |
|
| 19 |
+
| allenai/dolmino-mix-1124 / stackexchange | `a319f19eef1e257417b11ea8c30da266ae175557` | cc-by-sa |
|
| 20 |
+
|
| 21 |
+
Publisher datasets: https://huggingface.co/datasets/HuggingFaceTB/smollm-corpus ; https://huggingface.co/datasets/HuggingFaceTB/dclm-edu ; https://huggingface.co/datasets/HuggingFaceFW/finewiki ; https://huggingface.co/datasets/common-pile/stackv2_edu_filtered ; https://huggingface.co/datasets/HuggingFaceFW/finephrase ; https://huggingface.co/datasets/allenai/dolmino-mix-1124 .
|
| 22 |
+
|
| 23 |
+
## Instruction components
|
| 24 |
+
|
| 25 |
+
The collection is HuggingFaceTB/smol-smoltalk, revision `f73fe857d519ff6ac5af2ea67c4d3834da7b8bcc`, default/train. Its pinned card has an Apache-2.0 badge. P2's seven labels are source-column values, not separate repositories. Parent/component correspondence is documentary, not a per-row join; exact component revisions remain unestablished. Parent and component notices below are CURRENT_ONLY observations retrieved on 2026-09-10, not terms proven contemporaneous with training.
|
| 26 |
+
|
| 27 |
+
- **openhermes-50k**: teknium/OpenHermes-2.5. Recorded metadata/grant: no licence in component-card metadata. component dataset card metadata at head b82037821055c377bed0d495e72e46de3bc72e84 (retrieved 2026-09-10T17:53:40Z)
|
| 28 |
+
- **smol-contraints**: Smol-contraints. Recorded metadata/grant: Apache-2.0. parent collection card at head 5feaf2fd3ffca7c237fc38d1861bc30365d48ffa (retrieved 2026-09-10T17:52:48Z): "All the new datasets (Smol-Magpie-Ultra, Smol-contraints, Smol-rewrite, Smol-summarize) are licensed under Apache 2.0."
|
| 29 |
+
- **smollm-rewrite-30k**: Smol-rewrite. Recorded metadata/grant: Apache-2.0. parent collection card at head 5feaf2fd3ffca7c237fc38d1861bc30365d48ffa (retrieved 2026-09-10T17:52:48Z): "All the new datasets (Smol-Magpie-Ultra, Smol-contraints, Smol-rewrite, Smol-summarize) are licensed under Apache 2.0."
|
| 30 |
+
- **smol-magpie-ultra-short**: Smol-Magpie-Ultra. Recorded metadata/grant: Apache-2.0. parent collection card at head 5feaf2fd3ffca7c237fc38d1861bc30365d48ffa (retrieved 2026-09-10T17:52:48Z): "All the new datasets (Smol-Magpie-Ultra, Smol-contraints, Smol-rewrite, Smol-summarize) are licensed under Apache 2.0."
|
| 31 |
+
- **self-oss-instruct**: bigcode/self-oss-instruct-sc2-exec-filter-50k. Recorded metadata/grant: odc-by. component dataset card metadata at head 356bb069eee815daa6e23e9a282eeefe1490ad44 (retrieved 2026-09-10T17:53:40Z)
|
| 32 |
+
- **smol-summarize-20k**: Smol-summarize. Recorded metadata/grant: Apache-2.0. parent collection card at head 5feaf2fd3ffca7c237fc38d1861bc30365d48ffa (retrieved 2026-09-10T17:52:48Z): "All the new datasets (Smol-Magpie-Ultra, Smol-contraints, Smol-rewrite, Smol-summarize) are licensed under Apache 2.0."
|
| 33 |
+
- **everyday-conversations**: HuggingFaceTB/everyday-conversations-llama3.1-2k. Recorded metadata/grant: apache-2.0. component dataset card metadata at head 14f543216b9ba42b6b951dc5bd199460d193b162 (retrieved 2026-09-10T17:53:41Z)
|
| 34 |
+
|
| 35 |
+
The parent publisher limits its Apache grant to its newly generated subsets and points to component-specific terms for incorporated datasets. The self-oss-instruct ODC-By metadata differs from the collection badge; a declaration difference is not proof of legal incompatibility. OpenHermes-2.5 refers to component-specific licences; its empty card licence metadata does not resolve those terms, and exact upstream-component rows for this subset remain unresolved. Ordinary historical unknowns have not been newly resolved.
|
| 36 |
+
|
| 37 |
+
Sources: https://huggingface.co/datasets/HuggingFaceTB/smol-smoltalk/tree/f73fe857d519ff6ac5af2ea67c4d3834da7b8bcc ; https://huggingface.co/datasets/HuggingFaceTB/smoltalk ; https://huggingface.co/datasets/teknium/OpenHermes-2.5 ; https://huggingface.co/datasets/bigcode/self-oss-instruct-sc2-exec-filter-50k ; https://huggingface.co/datasets/HuggingFaceTB/everyday-conversations-llama3.1-2k .
|
| 38 |
+
|
| 39 |
+
No raw corpus, frozen evaluation prompts or model answers are distributed. Author licensing does not relicense upstream datasets or clear third-party rights. See DOCUMENTATION_LICENSE.md for the scoped grant.
|
THIRD_PARTY_NOTICES.md
ADDED
|
@@ -0,0 +1,5 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Third-party notices
|
| 2 |
+
|
| 3 |
+
Existing third-party code, attribution headers and notices retain their own applicable terms; the root Apache-2.0 licence grants only rights controlled by Yang Qi. The native code is copied byte-for-byte from the accepted export, with its existing comments preserved. Dependencies are declared, not vendored: PyTorch, NumPy, Hugging Face Tokenizers and Safetensors retain their upstream licences and notices. This release makes no blanket claim about all existing repository code.
|
| 4 |
+
|
| 5 |
+
Dataset authors and contributors are acknowledged in SOURCE_NOTICE.md, including ODC-By components, Wikipedia/FineWiki and StackExchange share-alike metadata, per-record Python licence metadata, and unresolved OpenHermes component terms. Dataset terms are not replaced by the model's Apache-2.0 declaration.
|
config.json
ADDED
|
@@ -0,0 +1,13 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"vocab_size": 32000,
|
| 3 |
+
"n_layers": 30,
|
| 4 |
+
"d_model": 576,
|
| 5 |
+
"n_heads": 9,
|
| 6 |
+
"n_kv_heads": 3,
|
| 7 |
+
"d_ff": 1536,
|
| 8 |
+
"max_seq_len": 2048,
|
| 9 |
+
"dropout": 0.0,
|
| 10 |
+
"tie_embeddings": true,
|
| 11 |
+
"rope_theta": 10000.0,
|
| 12 |
+
"rope_pct": 1.0
|
| 13 |
+
}
|
inference.py
ADDED
|
@@ -0,0 +1,63 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Local native alpha075 loading and precision profiles."""
|
| 2 |
+
import argparse,json
|
| 3 |
+
from pathlib import Path
|
| 4 |
+
from contextlib import contextmanager
|
| 5 |
+
import torch
|
| 6 |
+
from torch.nn.attention import sdpa_kernel,SDPBackend
|
| 7 |
+
from safetensors.torch import load_model
|
| 8 |
+
from src.model import GPT,gpt_config_from_checkpoint_dict,audit_gpt_parameter_count
|
| 9 |
+
from src.chat_template import load_chat_tokenizer,encode_prompt
|
| 10 |
+
from src.accepted_generate import generate as accepted_generate
|
| 11 |
+
|
| 12 |
+
@contextmanager
|
| 13 |
+
def precision(profile):
|
| 14 |
+
if profile not in ('bf16_native','fp32_math'): raise ValueError('Unsupported precision profile')
|
| 15 |
+
torch.set_float32_matmul_precision('highest')
|
| 16 |
+
torch.backends.cuda.matmul.allow_tf32=False
|
| 17 |
+
torch.backends.cudnn.allow_tf32=profile=='bf16_native'
|
| 18 |
+
with sdpa_kernel([SDPBackend.MATH] if profile=='fp32_math' else [SDPBackend.FLASH_ATTENTION,SDPBackend.EFFICIENT_ATTENTION,SDPBackend.MATH,SDPBackend.CUDNN_ATTENTION]):
|
| 19 |
+
with torch.autocast('cuda',dtype=torch.bfloat16,enabled=profile=='bf16_native'):
|
| 20 |
+
yield
|
| 21 |
+
|
| 22 |
+
def load_bundle(model_dir):
|
| 23 |
+
root=Path(model_dir).resolve()
|
| 24 |
+
cfg=gpt_config_from_checkpoint_dict(json.loads((root/'config.json').read_text()))
|
| 25 |
+
model=GPT(cfg).eval()
|
| 26 |
+
load_model(model,str(root/'model.safetensors'),strict=True,device='cpu')
|
| 27 |
+
audit=audit_gpt_parameter_count(model,cfg)
|
| 28 |
+
if audit['actual_total']!=124635456: raise ValueError('Parameter identity guard failed')
|
| 29 |
+
if model.tok_emb.weight is not model.lm_head.weight: raise ValueError('Embedding tie lost')
|
| 30 |
+
if model.tok_emb.weight.data_ptr()!=model.lm_head.weight.data_ptr(): raise ValueError('Storage tie lost')
|
| 31 |
+
if any(p.dtype!=torch.float32 for p in model.parameters()): raise ValueError('Expected FP32 weights')
|
| 32 |
+
return model,load_chat_tokenizer(str(root/'tokenizer.json'))
|
| 33 |
+
|
| 34 |
+
def validate_input(tok,messages,cap):
|
| 35 |
+
if not isinstance(messages,list) or not messages: raise ValueError('Messages must be a non-empty JSON array')
|
| 36 |
+
if any(not isinstance(m,dict) or set(m)!={'role','content'} for m in messages): raise ValueError('Each message must contain only role and content')
|
| 37 |
+
if isinstance(cap,bool) or not isinstance(cap,int) or not 1<=cap<=384: raise ValueError('max_new_tokens must be 1..384')
|
| 38 |
+
ids=encode_prompt(tok,messages,default_system=None,mode='full_context')
|
| 39 |
+
if len(ids)+cap>2048: raise ValueError('Context overflow: prompt plus token budget exceeds 2048')
|
| 40 |
+
return ids
|
| 41 |
+
|
| 42 |
+
@torch.inference_mode()
|
| 43 |
+
def generate(model,tok,messages,cap=384,profile='bf16_native'):
|
| 44 |
+
validate_input(tok,messages,cap)
|
| 45 |
+
model.eval()
|
| 46 |
+
with precision(profile):
|
| 47 |
+
return accepted_generate(model,tok,messages,cap,autocast_enabled=profile=='bf16_native')
|
| 48 |
+
|
| 49 |
+
def main():
|
| 50 |
+
p=argparse.ArgumentParser(description=__doc__)
|
| 51 |
+
p.add_argument('--model-directory',required=True,type=Path)
|
| 52 |
+
g=p.add_mutually_exclusive_group(required=True)
|
| 53 |
+
g.add_argument('--messages-json',type=Path);g.add_argument('--prompt')
|
| 54 |
+
p.add_argument('--profile',choices=['bf16_native','fp32_math'],default='bf16_native')
|
| 55 |
+
p.add_argument('--max-new-tokens',type=int,default=384)
|
| 56 |
+
a=p.parse_args()
|
| 57 |
+
messages=json.loads(a.messages_json.read_text()) if a.messages_json else [{'role':'user','content':a.prompt}]
|
| 58 |
+
tok=load_chat_tokenizer(str(a.model_directory/'tokenizer.json'))
|
| 59 |
+
validate_input(tok,messages,a.max_new_tokens)
|
| 60 |
+
model,tok=load_bundle(a.model_directory)
|
| 61 |
+
model=model.to('cuda').eval()
|
| 62 |
+
print(json.dumps(generate(model,tok,messages,a.max_new_tokens,a.profile),ensure_ascii=False))
|
| 63 |
+
if __name__=='__main__': main()
|
model.safetensors
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:4396efb7a52b047e7fdf513e46d1b401dfc70582d3aca1f9cb5a07e97d426ef1
|
| 3 |
+
size 498562192
|
requirements-inference-tested.txt
ADDED
|
@@ -0,0 +1,5 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Python 3.10.12; CUDA 12.6 PyTorch build
|
| 2 |
+
torch==2.11.0+cu126
|
| 3 |
+
numpy==2.2.6
|
| 4 |
+
tokenizers==0.22.2
|
| 5 |
+
safetensors==0.8.0
|
special_tokens_map.json
ADDED
|
@@ -0,0 +1,12 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"format": "descriptive native metadata; no HF AutoModel claim",
|
| 3 |
+
"tokens": {
|
| 4 |
+
"[PAD]": 0,
|
| 5 |
+
"[UNK]": 1,
|
| 6 |
+
"[BOS]": 2,
|
| 7 |
+
"[EOS]": 3,
|
| 8 |
+
"<|system|>": 4,
|
| 9 |
+
"<|user|>": 5,
|
| 10 |
+
"<|assistant|>": 6
|
| 11 |
+
}
|
| 12 |
+
}
|
src/__init__.py
ADDED
|
File without changes
|
src/accepted_generate.py
ADDED
|
@@ -0,0 +1,18 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
import torch
|
| 2 |
+
from src.chat_template import encode_prompt
|
| 3 |
+
|
| 4 |
+
@torch.inference_mode()
|
| 5 |
+
def generate(model,tok,messages,cap,autocast_enabled=True):
|
| 6 |
+
ids=encode_prompt(tok,messages,default_system=None,mode='full_context')
|
| 7 |
+
assert len(ids)+cap<=2048
|
| 8 |
+
with torch.autocast('cuda',dtype=torch.bfloat16,enabled=autocast_enabled):
|
| 9 |
+
gen=model.generate(torch.tensor([ids],device='cuda'),max_new_tokens=cap,temperature=0,top_k=0,top_p=1,eos_id=3)
|
| 10 |
+
new=gen[0,len(ids):].tolist();eos=bool(new and new[-1]==3);before=new[:-1] if eos else new
|
| 11 |
+
text=tok.decode(before,skip_special_tokens=False)
|
| 12 |
+
grams=[tuple(before[i:i+4]) for i in range(max(0,len(before)-3))]
|
| 13 |
+
return {'prompt_token_ids':ids,'generated_token_ids':new,'output_text_for_scoring':text,
|
| 14 |
+
'output_text_raw_including_terminal_eos':tok.decode(new,skip_special_tokens=False),
|
| 15 |
+
'stop_reason':'eos' if eos else 'max_new_tokens','generated_tokens_including_eos':len(new),
|
| 16 |
+
'empty_output':not text,'repeated_4gram_fraction':1-len(set(grams))/len(grams) if grams else 0.,
|
| 17 |
+
'unexpected_control_token_count':sum(v in {0,2,4,5,6} for v in before),
|
| 18 |
+
'answer_cleanup':False,'constrained_decoding':False,'max_new_tokens':cap}
|
src/chat_template.py
ADDED
|
@@ -0,0 +1,507 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Single source of truth for the chat format, shared by SFT / distill / DPO /
|
| 2 |
+
GRPO training AND their sampling code (previously seven duplicated copies).
|
| 3 |
+
|
| 4 |
+
Token-level template — role boundaries are special tokens, not plain text:
|
| 5 |
+
|
| 6 |
+
[BOS] <|system|> {system} <|user|> {q1} <|assistant|> {a1} [EOS] <|user|> {q2} ...
|
| 7 |
+
|
| 8 |
+
Design rules:
|
| 9 |
+
- Role tokens delimit turns. BPE can never merge across a special token, so a
|
| 10 |
+
conversation encodes to the same ids whether it is built turn-by-turn during
|
| 11 |
+
training or as a generation prompt at inference (the old plain-text
|
| 12 |
+
"User: ...\\n\\n" template drifted at every segment boundary).
|
| 13 |
+
- [EOS] appears ONLY after assistant turns: it means "assistant finished,
|
| 14 |
+
stop generating" — the same stop semantics as document ends in pretraining.
|
| 15 |
+
System/user turns need no terminator; the next role token is the boundary.
|
| 16 |
+
- The supervised span is exactly each assistant turn's content tokens plus its
|
| 17 |
+
trailing [EOS] (so the model is explicitly taught to stop).
|
| 18 |
+
- Content is encoded with `tokenizer.encode_special_tokens = True`
|
| 19 |
+
(see `load_chat_tokenizer`), so literal "[EOS]" / "<|user|>" strings inside
|
| 20 |
+
user or corpus text are tokenized as plain text and can never inject real
|
| 21 |
+
control tokens.
|
| 22 |
+
"""
|
| 23 |
+
|
| 24 |
+
from __future__ import annotations
|
| 25 |
+
|
| 26 |
+
from typing import Any
|
| 27 |
+
|
| 28 |
+
from tokenizers import Tokenizer
|
| 29 |
+
|
| 30 |
+
from src.special_tokens import (
|
| 31 |
+
ASSISTANT_ID,
|
| 32 |
+
BOS_ID,
|
| 33 |
+
EOS_ID,
|
| 34 |
+
PAD_ID,
|
| 35 |
+
SYSTEM_ID,
|
| 36 |
+
USER_ID,
|
| 37 |
+
assert_tokenizer_contract,
|
| 38 |
+
)
|
| 39 |
+
|
| 40 |
+
DEFAULT_SYSTEM = "You are a helpful assistant."
|
| 41 |
+
|
| 42 |
+
IGNORE_INDEX = -100
|
| 43 |
+
|
| 44 |
+
|
| 45 |
+
# -------------------------
|
| 46 |
+
# Tokenizer loading
|
| 47 |
+
# -------------------------
|
| 48 |
+
def configure_chat_tokenizer(tok: Tokenizer) -> Tokenizer:
|
| 49 |
+
"""Make `tok.encode` treat special-token strings in raw text as plain text.
|
| 50 |
+
|
| 51 |
+
All special tokens in this pipeline are inserted by ID by the code below,
|
| 52 |
+
never parsed out of content — this closes the prompt-injection hole where a
|
| 53 |
+
document containing the literal string "[EOS]" would encode to the real
|
| 54 |
+
EOS id.
|
| 55 |
+
"""
|
| 56 |
+
tok.encode_special_tokens = True
|
| 57 |
+
return tok
|
| 58 |
+
|
| 59 |
+
|
| 60 |
+
def load_chat_tokenizer(tokenizer_path: str) -> Tokenizer:
|
| 61 |
+
"""Load tokenizer.json, assert the hardcoded special-token IDs, and disable
|
| 62 |
+
special-token matching in raw text. Every chat-stage script should load its
|
| 63 |
+
tokenizer through this."""
|
| 64 |
+
assert_tokenizer_contract(tokenizer_path)
|
| 65 |
+
return configure_chat_tokenizer(Tokenizer.from_file(tokenizer_path))
|
| 66 |
+
|
| 67 |
+
|
| 68 |
+
# -------------------------
|
| 69 |
+
# Text cleaning
|
| 70 |
+
# -------------------------
|
| 71 |
+
def norm_newlines(s: str) -> str:
|
| 72 |
+
return (s or "").replace("\r\n", "\n").replace("\r", "\n")
|
| 73 |
+
|
| 74 |
+
|
| 75 |
+
def clean_text(s: str) -> str:
|
| 76 |
+
# for system/user text: strip leading/trailing whitespace
|
| 77 |
+
return norm_newlines(s).strip()
|
| 78 |
+
|
| 79 |
+
|
| 80 |
+
def clean_text_assistant(s: str) -> str:
|
| 81 |
+
# IMPORTANT: do not strip assistant text (keeps code indentation / markdown
|
| 82 |
+
# formatting); only trailing whitespace goes (EOS follows immediately).
|
| 83 |
+
return norm_newlines(s).rstrip()
|
| 84 |
+
|
| 85 |
+
|
| 86 |
+
def _normalized_messages(
|
| 87 |
+
messages: list[dict[str, str]],
|
| 88 |
+
default_system: str | None,
|
| 89 |
+
*,
|
| 90 |
+
expected_end: str | None,
|
| 91 |
+
) -> list[dict[str, str]]:
|
| 92 |
+
"""Clean and validate the canonical chat state machine.
|
| 93 |
+
|
| 94 |
+
Normally a chat has an initial system turn, followed by alternating user
|
| 95 |
+
and assistant turns. Explicit default_system=None preserves source content
|
| 96 |
+
verbatim and permits a missing system turn without injecting one. ``expected_end`` is ``"user"`` for
|
| 97 |
+
a generation prompt and ``"assistant"`` for a complete training sample.
|
| 98 |
+
No malformed turn is silently skipped.
|
| 99 |
+
"""
|
| 100 |
+
out: list[dict[str, str]] = []
|
| 101 |
+
for index, message in enumerate(messages or []):
|
| 102 |
+
if not isinstance(message, dict):
|
| 103 |
+
raise ValueError(f"message {index} must be an object")
|
| 104 |
+
role_value = message.get("role")
|
| 105 |
+
role = role_value.strip().lower() if isinstance(role_value, str) else ""
|
| 106 |
+
if role not in ("system", "user", "assistant"):
|
| 107 |
+
raise ValueError(f"message {index} has invalid role {role_value!r}")
|
| 108 |
+
raw = message.get("content")
|
| 109 |
+
if not isinstance(raw, str):
|
| 110 |
+
raise ValueError(f"message {index} content must be a string")
|
| 111 |
+
text = raw if default_system is None else (
|
| 112 |
+
clean_text_assistant(raw) if role == "assistant" else clean_text(raw)
|
| 113 |
+
)
|
| 114 |
+
if not text.strip():
|
| 115 |
+
if index == 0 and role == "system" and default_system is not None:
|
| 116 |
+
fallback = clean_text(default_system)
|
| 117 |
+
if not fallback:
|
| 118 |
+
raise ValueError(
|
| 119 |
+
"chat requires a non-empty initial system turn or non-empty default_system"
|
| 120 |
+
)
|
| 121 |
+
text = fallback
|
| 122 |
+
else:
|
| 123 |
+
raise ValueError(f"message {index} ({role}) content must be non-empty")
|
| 124 |
+
out.append({"role": role, "content": text})
|
| 125 |
+
|
| 126 |
+
if (not out or out[0]["role"] != "system") and default_system is not None:
|
| 127 |
+
fallback = clean_text(default_system)
|
| 128 |
+
if not fallback:
|
| 129 |
+
raise ValueError(
|
| 130 |
+
"chat requires a non-empty initial system turn or non-empty default_system"
|
| 131 |
+
)
|
| 132 |
+
out.insert(0, {"role": "system", "content": fallback})
|
| 133 |
+
|
| 134 |
+
has_system = bool(out) and out[0]["role"] == "system"
|
| 135 |
+
if not out or (has_system and len(out) == 1):
|
| 136 |
+
raise ValueError("chat requires at least one non-empty user turn")
|
| 137 |
+
for index, message in enumerate(out):
|
| 138 |
+
turn_index = index - int(has_system)
|
| 139 |
+
expected = "system" if has_system and index == 0 else (
|
| 140 |
+
"user" if turn_index % 2 == 0 else "assistant"
|
| 141 |
+
)
|
| 142 |
+
if message["role"] != expected:
|
| 143 |
+
raise ValueError(
|
| 144 |
+
"invalid chat role order: initial system must be followed by "
|
| 145 |
+
f"alternating user/assistant turns (index {index}: expected "
|
| 146 |
+
f"{expected!r}, got {message['role']!r})"
|
| 147 |
+
)
|
| 148 |
+
if expected_end is not None and out[-1]["role"] != expected_end:
|
| 149 |
+
raise ValueError(
|
| 150 |
+
f"chat must end with a non-empty {expected_end} turn; got {out[-1]['role']!r}"
|
| 151 |
+
)
|
| 152 |
+
return out
|
| 153 |
+
|
| 154 |
+
|
| 155 |
+
def prepare_prompt_messages(
|
| 156 |
+
messages: list[dict[str, str]],
|
| 157 |
+
default_system: str | None = DEFAULT_SYSTEM,
|
| 158 |
+
) -> list[dict[str, str]]:
|
| 159 |
+
"""Explicitly turn a valid conversation/example into a USER-ending prompt.
|
| 160 |
+
|
| 161 |
+
Callers sampling from a complete SFT example must opt in to removing its
|
| 162 |
+
final assistant answer. ``encode_prompt`` itself never drops turns.
|
| 163 |
+
"""
|
| 164 |
+
normalized = _normalized_messages(messages, default_system, expected_end=None)
|
| 165 |
+
if normalized[-1]["role"] == "assistant":
|
| 166 |
+
normalized = normalized[:-1]
|
| 167 |
+
if normalized[-1]["role"] != "user":
|
| 168 |
+
raise ValueError("generation prompt context must end with a user turn")
|
| 169 |
+
return normalized
|
| 170 |
+
|
| 171 |
+
|
| 172 |
+
# -------------------------
|
| 173 |
+
# Core encoding
|
| 174 |
+
# -------------------------
|
| 175 |
+
_ROLE_TOKEN_ID = {"system": SYSTEM_ID, "user": USER_ID, "assistant": ASSISTANT_ID}
|
| 176 |
+
|
| 177 |
+
|
| 178 |
+
def _encode_normalized_chat(
|
| 179 |
+
tok: Tokenizer, messages: list[dict[str, str]]
|
| 180 |
+
) -> tuple[list[int], list[int]]:
|
| 181 |
+
ids: list[int] = [BOS_ID]
|
| 182 |
+
labels: list[int] = [IGNORE_INDEX]
|
| 183 |
+
for index, message in enumerate(messages):
|
| 184 |
+
role = message["role"]
|
| 185 |
+
content_ids = tok.encode(message["content"]).ids
|
| 186 |
+
if not content_ids:
|
| 187 |
+
raise ValueError(f"message {index} ({role}) must encode to at least one token")
|
| 188 |
+
ids.append(_ROLE_TOKEN_ID[role])
|
| 189 |
+
labels.append(IGNORE_INDEX)
|
| 190 |
+
if role == "assistant":
|
| 191 |
+
ids.extend(content_ids)
|
| 192 |
+
labels.extend(content_ids)
|
| 193 |
+
ids.append(EOS_ID)
|
| 194 |
+
labels.append(EOS_ID)
|
| 195 |
+
else:
|
| 196 |
+
ids.extend(content_ids)
|
| 197 |
+
labels.extend([IGNORE_INDEX] * len(content_ids))
|
| 198 |
+
return ids, labels
|
| 199 |
+
|
| 200 |
+
|
| 201 |
+
def encode_chat(
|
| 202 |
+
tok: Tokenizer,
|
| 203 |
+
messages: list[dict[str, str]],
|
| 204 |
+
default_system: str | None = DEFAULT_SYSTEM,
|
| 205 |
+
) -> tuple[list[int], list[int]]:
|
| 206 |
+
"""Encode a full conversation for training.
|
| 207 |
+
|
| 208 |
+
Returns (ids, labels), same length. labels[i] == ids[i] on supervised
|
| 209 |
+
positions (every assistant turn's content + its trailing EOS) and
|
| 210 |
+
IGNORE_INDEX everywhere else (BOS, role tokens, system/user content).
|
| 211 |
+
"""
|
| 212 |
+
msgs = _normalized_messages(messages, default_system, expected_end="assistant")
|
| 213 |
+
return _encode_normalized_chat(tok, msgs)
|
| 214 |
+
|
| 215 |
+
|
| 216 |
+
def encode_prompt(
|
| 217 |
+
tok: Tokenizer,
|
| 218 |
+
messages: list[dict[str, str]],
|
| 219 |
+
default_system: str | None = DEFAULT_SYSTEM,
|
| 220 |
+
mode: str = "full_context",
|
| 221 |
+
) -> list[int]:
|
| 222 |
+
"""Encode a generation prompt: context ending in the ``<|assistant|>`` cue.
|
| 223 |
+
|
| 224 |
+
mode:
|
| 225 |
+
- "full_context": the entire validated USER-ending context (earlier
|
| 226 |
+
assistant turns keep their [EOS]).
|
| 227 |
+
- "last_user": system turn + last user turn only.
|
| 228 |
+
|
| 229 |
+
The returned ids start with BOS and end with ASSISTANT_ID, exactly matching
|
| 230 |
+
the training-time prefix for an assistant turn.
|
| 231 |
+
"""
|
| 232 |
+
msgs = _normalized_messages(messages, default_system, expected_end="user")
|
| 233 |
+
|
| 234 |
+
if mode == "last_user":
|
| 235 |
+
msgs = [msgs[0], msgs[-1]] if msgs[0]["role"] == "system" else [msgs[-1]]
|
| 236 |
+
elif mode != "full_context":
|
| 237 |
+
raise ValueError(f"unknown prompt mode: {mode}")
|
| 238 |
+
|
| 239 |
+
ids, _ = _encode_normalized_chat(tok, msgs)
|
| 240 |
+
ids.append(ASSISTANT_ID)
|
| 241 |
+
return ids
|
| 242 |
+
|
| 243 |
+
|
| 244 |
+
def encode_completion(
|
| 245 |
+
tok: Tokenizer,
|
| 246 |
+
messages: list[dict[str, str]],
|
| 247 |
+
completion: str,
|
| 248 |
+
default_system: str | None = DEFAULT_SYSTEM,
|
| 249 |
+
) -> tuple[list[int], list[int]]:
|
| 250 |
+
"""Encode (prompt context + one assistant completion) for DPO-style scoring.
|
| 251 |
+
|
| 252 |
+
`messages` is the shared context (must end with a user turn); `completion`
|
| 253 |
+
is a plain assistant string. Returns (ids, labels) where the supervised
|
| 254 |
+
span is the completion tokens + trailing EOS — logps therefore include the
|
| 255 |
+
stop decision.
|
| 256 |
+
"""
|
| 257 |
+
ids = encode_prompt(tok, messages, default_system, mode="full_context")
|
| 258 |
+
labels = [IGNORE_INDEX] * len(ids)
|
| 259 |
+
if not isinstance(completion, str) or not completion.strip():
|
| 260 |
+
raise ValueError("DPO completion must be a non-empty, non-whitespace string")
|
| 261 |
+
comp_ids = tok.encode(completion if default_system is None else clean_text_assistant(completion)).ids
|
| 262 |
+
if not comp_ids:
|
| 263 |
+
raise ValueError("DPO completion must encode to at least one token")
|
| 264 |
+
ids.extend(comp_ids)
|
| 265 |
+
labels.extend(comp_ids)
|
| 266 |
+
ids.append(EOS_ID)
|
| 267 |
+
labels.append(EOS_ID)
|
| 268 |
+
return ids, labels
|
| 269 |
+
|
| 270 |
+
|
| 271 |
+
def truncate_chat_sequence(
|
| 272 |
+
ids: list[int],
|
| 273 |
+
labels: list[int] | None,
|
| 274 |
+
max_len: int,
|
| 275 |
+
) -> tuple[list[int], list[int] | None]:
|
| 276 |
+
"""Validate and truncate only at complete user-turn boundaries.
|
| 277 |
+
|
| 278 |
+
The prefix is BOS, plus the initial SYSTEM turn when present. It is always
|
| 279 |
+
retained; source-preserving conversations may begin BOS + USER. The remainder is the largest recent suffix that begins at a real
|
| 280 |
+
``USER_ID`` marker and runs through the original sequence end. A training
|
| 281 |
+
sequence must end with an assistant EOS; a prompt must end with the final
|
| 282 |
+
assistant cue. Literal special-token text cannot create these boundaries
|
| 283 |
+
because chat tokenizers disable special-token matching for raw content.
|
| 284 |
+
|
| 285 |
+
If the system prefix plus the latest complete user-led suffix cannot fit,
|
| 286 |
+
this function raises instead of cutting content, role markers, assistant
|
| 287 |
+
completions, or EOS targets.
|
| 288 |
+
"""
|
| 289 |
+
if max_len <= 0:
|
| 290 |
+
raise ValueError("max_len must be positive")
|
| 291 |
+
if labels is not None and len(labels) != len(ids):
|
| 292 |
+
raise ValueError("ids and labels must have identical lengths")
|
| 293 |
+
if len(ids) < 3 or ids[0] != BOS_ID or ids[1] not in {SYSTEM_ID, USER_ID}:
|
| 294 |
+
raise ValueError("chat sequence must start with BOS_ID and a SYSTEM or USER turn")
|
| 295 |
+
|
| 296 |
+
forbidden_content_ids = {PAD_ID, BOS_ID, EOS_ID, SYSTEM_ID, USER_ID, ASSISTANT_ID}
|
| 297 |
+
role_ids = {SYSTEM_ID, USER_ID, ASSISTANT_ID}
|
| 298 |
+
first_role = 1
|
| 299 |
+
if ids[1] == SYSTEM_ID:
|
| 300 |
+
first_role = next((index for index in range(2, len(ids)) if ids[index] in role_ids), len(ids))
|
| 301 |
+
if first_role == 2:
|
| 302 |
+
raise ValueError("initial system content must contain at least one token")
|
| 303 |
+
if any(token_id in forbidden_content_ids for token_id in ids[2:first_role]):
|
| 304 |
+
raise ValueError("system content contains a structural special-token ID")
|
| 305 |
+
if first_role == len(ids) or ids[first_role] != USER_ID:
|
| 306 |
+
raise ValueError("chat sequence must contain a user turn after the initial system turn")
|
| 307 |
+
|
| 308 |
+
user_starts: list[int] = []
|
| 309 |
+
cursor = first_role
|
| 310 |
+
while cursor < len(ids):
|
| 311 |
+
if ids[cursor] != USER_ID:
|
| 312 |
+
raise ValueError("chat role sequence must alternate USER and ASSISTANT")
|
| 313 |
+
user_starts.append(cursor)
|
| 314 |
+
user_content_start = cursor + 1
|
| 315 |
+
assistant_pos = next(
|
| 316 |
+
(
|
| 317 |
+
index
|
| 318 |
+
for index in range(user_content_start, len(ids))
|
| 319 |
+
if ids[index] in forbidden_content_ids
|
| 320 |
+
),
|
| 321 |
+
len(ids),
|
| 322 |
+
)
|
| 323 |
+
if assistant_pos == user_content_start:
|
| 324 |
+
raise ValueError("user content must contain at least one token")
|
| 325 |
+
if assistant_pos == len(ids) or ids[assistant_pos] != ASSISTANT_ID:
|
| 326 |
+
raise ValueError("each user turn must be followed by an assistant turn")
|
| 327 |
+
|
| 328 |
+
assistant_content_start = assistant_pos + 1
|
| 329 |
+
if assistant_content_start == len(ids):
|
| 330 |
+
if labels is not None:
|
| 331 |
+
raise ValueError("training chat must end with assistant content and EOS")
|
| 332 |
+
cursor = len(ids)
|
| 333 |
+
break
|
| 334 |
+
|
| 335 |
+
eos_pos = next(
|
| 336 |
+
(
|
| 337 |
+
index
|
| 338 |
+
for index in range(assistant_content_start, len(ids))
|
| 339 |
+
if ids[index] in forbidden_content_ids
|
| 340 |
+
),
|
| 341 |
+
len(ids),
|
| 342 |
+
)
|
| 343 |
+
if eos_pos == assistant_content_start:
|
| 344 |
+
raise ValueError("assistant content must contain at least one token")
|
| 345 |
+
if eos_pos == len(ids) or ids[eos_pos] != EOS_ID:
|
| 346 |
+
raise ValueError("each assistant completion must end with EOS")
|
| 347 |
+
cursor = eos_pos + 1
|
| 348 |
+
if cursor == len(ids):
|
| 349 |
+
if labels is None:
|
| 350 |
+
raise ValueError("generation prompt must end with an assistant cue")
|
| 351 |
+
break
|
| 352 |
+
|
| 353 |
+
if labels is not None:
|
| 354 |
+
if ids[-1] != EOS_ID:
|
| 355 |
+
raise ValueError("training chat must end with a supervised assistant EOS")
|
| 356 |
+
if labels[-1] != EOS_ID:
|
| 357 |
+
raise ValueError("final assistant EOS must be supervised")
|
| 358 |
+
|
| 359 |
+
prefix_end = first_role
|
| 360 |
+
chosen_start = next(
|
| 361 |
+
(start for start in user_starts if prefix_end + (len(ids) - start) <= max_len),
|
| 362 |
+
None,
|
| 363 |
+
)
|
| 364 |
+
if chosen_start is None:
|
| 365 |
+
minimum = prefix_end + (len(ids) - user_starts[-1])
|
| 366 |
+
raise ValueError(
|
| 367 |
+
"chat sequence does not fit without cutting the system prefix or latest "
|
| 368 |
+
f"user-led suffix (requires at least {minimum} tokens, max_len={max_len})"
|
| 369 |
+
)
|
| 370 |
+
|
| 371 |
+
if chosen_start == prefix_end:
|
| 372 |
+
return list(ids), list(labels) if labels is not None else None
|
| 373 |
+
|
| 374 |
+
kept_ids = ids[:prefix_end] + ids[chosen_start:]
|
| 375 |
+
if labels is None:
|
| 376 |
+
return kept_ids, None
|
| 377 |
+
return kept_ids, labels[:prefix_end] + labels[chosen_start:]
|
| 378 |
+
|
| 379 |
+
|
| 380 |
+
def pad_or_truncate(
|
| 381 |
+
ids: list[int],
|
| 382 |
+
labels: list[int],
|
| 383 |
+
seq_len: int,
|
| 384 |
+
pad_id: int = PAD_ID,
|
| 385 |
+
) -> tuple[list[int], list[int]]:
|
| 386 |
+
"""Structure-aware fixed-length shaping plus right padding."""
|
| 387 |
+
ids, maybe_labels = truncate_chat_sequence(ids, labels, seq_len)
|
| 388 |
+
assert maybe_labels is not None
|
| 389 |
+
labels = maybe_labels
|
| 390 |
+
pad_n = seq_len - len(ids)
|
| 391 |
+
return ids + [pad_id] * pad_n, labels + [IGNORE_INDEX] * pad_n
|
| 392 |
+
|
| 393 |
+
|
| 394 |
+
def count_chat_tokens(
|
| 395 |
+
tok: Tokenizer,
|
| 396 |
+
messages: list[dict[str, str]],
|
| 397 |
+
default_system: str | None = DEFAULT_SYSTEM,
|
| 398 |
+
) -> int:
|
| 399 |
+
"""Exact token count of the training encoding (for mix/token budgeting)."""
|
| 400 |
+
ids, _ = encode_chat(tok, messages, default_system)
|
| 401 |
+
return len(ids)
|
| 402 |
+
|
| 403 |
+
|
| 404 |
+
def extract_last_user_and_ref(messages: list[dict[str, str]]) -> tuple[str, str]:
|
| 405 |
+
"""Return (last_user_text, last_assistant_text_if_any) — for sample logs."""
|
| 406 |
+
last_user = ""
|
| 407 |
+
ref = ""
|
| 408 |
+
for m in reversed(messages or []):
|
| 409 |
+
if (m.get("role") or "").strip().lower() == "user":
|
| 410 |
+
last_user = clean_text(m.get("content", ""))
|
| 411 |
+
break
|
| 412 |
+
for m in reversed(messages or []):
|
| 413 |
+
if (m.get("role") or "").strip().lower() == "assistant":
|
| 414 |
+
ref = clean_text_assistant(m.get("content", ""))
|
| 415 |
+
break
|
| 416 |
+
return last_user, ref
|
| 417 |
+
|
| 418 |
+
|
| 419 |
+
def decode_completion(tok: Tokenizer, ids: list[int]) -> str:
|
| 420 |
+
"""Decode generated completion ids (the tokens AFTER the prompt), dropping
|
| 421 |
+
a trailing EOS if present. Replaces the old fragile rfind('Assistant: ')."""
|
| 422 |
+
if ids and ids[-1] == EOS_ID:
|
| 423 |
+
ids = ids[:-1]
|
| 424 |
+
return tok.decode(ids).strip() if ids else ""
|
| 425 |
+
|
| 426 |
+
|
| 427 |
+
# -------------------------
|
| 428 |
+
# Refusal detection (shared by SFT/distill example weighting)
|
| 429 |
+
# -------------------------
|
| 430 |
+
def is_refusal_text(text: str, patterns: list[str]) -> bool:
|
| 431 |
+
"""If assistant content contains any refusal-ish substring, treat as refusal."""
|
| 432 |
+
t = (text or "").strip().lower()
|
| 433 |
+
if not t:
|
| 434 |
+
return False
|
| 435 |
+
for p in patterns:
|
| 436 |
+
p2 = p.strip().lower()
|
| 437 |
+
if p2 and p2 in t:
|
| 438 |
+
return True
|
| 439 |
+
return False
|
| 440 |
+
|
| 441 |
+
|
| 442 |
+
def compute_example_weight_from_messages(
|
| 443 |
+
messages: list[dict[str, str]],
|
| 444 |
+
refusal_downweight: float,
|
| 445 |
+
refusal_patterns: list[str],
|
| 446 |
+
refusal_mode: str,
|
| 447 |
+
) -> float:
|
| 448 |
+
"""Scalar loss weight for a training example; downweights refusal-looking
|
| 449 |
+
assistant turns (refusal_mode="contains_any")."""
|
| 450 |
+
if refusal_downweight >= 1.0:
|
| 451 |
+
return 1.0
|
| 452 |
+
if refusal_downweight <= 0.0:
|
| 453 |
+
return 0.0
|
| 454 |
+
if refusal_mode != "contains_any":
|
| 455 |
+
raise ValueError(f"unknown refusal_mode: {refusal_mode}")
|
| 456 |
+
for m in messages or []:
|
| 457 |
+
if (m.get("role") or "").strip().lower() == "assistant":
|
| 458 |
+
if is_refusal_text(m.get("content", ""), refusal_patterns):
|
| 459 |
+
return refusal_downweight
|
| 460 |
+
return 1.0
|
| 461 |
+
|
| 462 |
+
|
| 463 |
+
def build_example(
|
| 464 |
+
ex: dict[str, Any],
|
| 465 |
+
tok: Tokenizer,
|
| 466 |
+
seq_len: int,
|
| 467 |
+
default_system: str | None,
|
| 468 |
+
refusal_downweight: float,
|
| 469 |
+
refusal_patterns: list[str],
|
| 470 |
+
refusal_mode: str,
|
| 471 |
+
pad_id: int = PAD_ID,
|
| 472 |
+
):
|
| 473 |
+
"""One SFT/distill training example -> (input_ids, labels, example_weight).
|
| 474 |
+
|
| 475 |
+
Tensors are torch.long of length seq_len; labels use IGNORE_INDEX outside
|
| 476 |
+
the supervised assistant spans. Honors meta.bucket safety exemption and
|
| 477 |
+
meta.weight multipliers exactly as before.
|
| 478 |
+
"""
|
| 479 |
+
import torch
|
| 480 |
+
|
| 481 |
+
messages = ex.get("messages") or []
|
| 482 |
+
if not messages:
|
| 483 |
+
raise ValueError("missing messages")
|
| 484 |
+
|
| 485 |
+
meta = ex.get("meta") or {}
|
| 486 |
+
bucket = str(meta.get("bucket", "")).strip()
|
| 487 |
+
# Do NOT downweight refusals inside the safety bucket (otherwise safety
|
| 488 |
+
# examples get muted).
|
| 489 |
+
refusal_dw_eff = 1.0 if bucket in ("D_safety", "D") else refusal_downweight
|
| 490 |
+
ex_weight = compute_example_weight_from_messages(
|
| 491 |
+
messages, refusal_dw_eff, refusal_patterns, refusal_mode
|
| 492 |
+
)
|
| 493 |
+
w0 = meta.get("weight", None) if isinstance(meta, dict) else None
|
| 494 |
+
if isinstance(w0, (int, float)):
|
| 495 |
+
ex_weight *= float(w0)
|
| 496 |
+
|
| 497 |
+
ids, labels = encode_chat(tok, messages, default_system)
|
| 498 |
+
if default_system is None and len(ids) > seq_len:
|
| 499 |
+
raise ValueError("source-preserving SFT rejects overlength conversations")
|
| 500 |
+
ids, labels = pad_or_truncate(ids, labels, seq_len, pad_id)
|
| 501 |
+
if all(label == IGNORE_INDEX for label in labels):
|
| 502 |
+
raise ValueError("SFT example requires a retained non-empty assistant turn")
|
| 503 |
+
return (
|
| 504 |
+
torch.tensor(ids, dtype=torch.long),
|
| 505 |
+
torch.tensor(labels, dtype=torch.long),
|
| 506 |
+
float(ex_weight),
|
| 507 |
+
)
|
src/model.py
ADDED
|
@@ -0,0 +1,470 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
from __future__ import annotations
|
| 2 |
+
|
| 3 |
+
from dataclasses import dataclass
|
| 4 |
+
import math
|
| 5 |
+
|
| 6 |
+
import torch
|
| 7 |
+
import torch.nn as nn
|
| 8 |
+
import torch.nn.functional as F
|
| 9 |
+
|
| 10 |
+
|
| 11 |
+
@dataclass
|
| 12 |
+
class GPTConfig:
|
| 13 |
+
vocab_size: int = 32000
|
| 14 |
+
n_layers: int = 30
|
| 15 |
+
d_model: int = 576
|
| 16 |
+
n_heads: int = 9
|
| 17 |
+
n_kv_heads: int = 3 # GQA KV heads; == n_heads is plain MHA (pre-GQA checkpoints)
|
| 18 |
+
d_ff: int = 1536 # SwiGLU: ~2.67x d_model (MobileLLM/SmolLM2-135M deep-thin shape)
|
| 19 |
+
max_seq_len: int = 2048
|
| 20 |
+
dropout: float = 0.0
|
| 21 |
+
tie_embeddings: bool = True
|
| 22 |
+
|
| 23 |
+
# RoPE (rotary positional embedding)
|
| 24 |
+
rope_theta: float = 10000.0
|
| 25 |
+
rope_pct: float = 1.0 # fraction of head_dim to rotate (1.0 = full head_dim)
|
| 26 |
+
|
| 27 |
+
|
| 28 |
+
CANONICAL_DENSE_PARAMETER_COUNT = 124_635_456
|
| 29 |
+
_CANONICAL_PARAMETERIZATION = {
|
| 30 |
+
"vocab_size": 32_000,
|
| 31 |
+
"n_layers": 30,
|
| 32 |
+
"d_model": 576,
|
| 33 |
+
"n_heads": 9,
|
| 34 |
+
"n_kv_heads": 3,
|
| 35 |
+
"d_ff": 1_536,
|
| 36 |
+
"tie_embeddings": True,
|
| 37 |
+
}
|
| 38 |
+
|
| 39 |
+
|
| 40 |
+
def expected_gpt_parameter_count(cfg: GPTConfig) -> int:
|
| 41 |
+
"""Derive the unique parameter count for the dense bias-free GPT."""
|
| 42 |
+
integer_fields = {
|
| 43 |
+
"vocab_size": cfg.vocab_size,
|
| 44 |
+
"n_layers": cfg.n_layers,
|
| 45 |
+
"d_model": cfg.d_model,
|
| 46 |
+
"n_heads": cfg.n_heads,
|
| 47 |
+
"n_kv_heads": cfg.n_kv_heads,
|
| 48 |
+
"d_ff": cfg.d_ff,
|
| 49 |
+
}
|
| 50 |
+
for name, value in integer_fields.items():
|
| 51 |
+
if isinstance(value, bool) or not isinstance(value, int) or value <= 0:
|
| 52 |
+
raise ValueError(f"GPTConfig.{name} must be a positive integer")
|
| 53 |
+
if cfg.d_model % cfg.n_heads:
|
| 54 |
+
raise ValueError("GPTConfig.d_model must be divisible by n_heads")
|
| 55 |
+
if cfg.n_heads % cfg.n_kv_heads:
|
| 56 |
+
raise ValueError("GPTConfig.n_heads must be divisible by n_kv_heads")
|
| 57 |
+
|
| 58 |
+
head_dim = cfg.d_model // cfg.n_heads
|
| 59 |
+
kv_dim = cfg.n_kv_heads * head_dim
|
| 60 |
+
token_matrices = 1 if cfg.tie_embeddings else 2
|
| 61 |
+
embeddings = token_matrices * cfg.vocab_size * cfg.d_model
|
| 62 |
+
# q + output projections are d_model x d_model; k and v are d_model x kv_dim (GQA)
|
| 63 |
+
attention = 2 * cfg.d_model * cfg.d_model + 2 * cfg.d_model * kv_dim
|
| 64 |
+
swiglu = 3 * cfg.d_model * cfg.d_ff
|
| 65 |
+
block_norms = 2 * cfg.d_model
|
| 66 |
+
final_norm = cfg.d_model
|
| 67 |
+
return int(embeddings + cfg.n_layers * (attention + swiglu + block_norms) + final_norm)
|
| 68 |
+
|
| 69 |
+
|
| 70 |
+
def audit_gpt_parameter_count(model: nn.Module, cfg: GPTConfig) -> dict[str, int | bool | str]:
|
| 71 |
+
"""Fail fast on implementation/config drift and return manifest metadata."""
|
| 72 |
+
expected = expected_gpt_parameter_count(cfg)
|
| 73 |
+
actual = int(sum(parameter.numel() for parameter in model.parameters()))
|
| 74 |
+
trainable = int(
|
| 75 |
+
sum(parameter.numel() for parameter in model.parameters() if parameter.requires_grad)
|
| 76 |
+
)
|
| 77 |
+
if actual != expected:
|
| 78 |
+
raise RuntimeError(
|
| 79 |
+
"GPT parameter count disagrees with the architecture-derived count: "
|
| 80 |
+
f"actual={actual:,}, expected={expected:,}"
|
| 81 |
+
)
|
| 82 |
+
|
| 83 |
+
canonical = all(
|
| 84 |
+
getattr(cfg, field) == expected_value
|
| 85 |
+
for field, expected_value in _CANONICAL_PARAMETERIZATION.items()
|
| 86 |
+
)
|
| 87 |
+
if canonical and actual != CANONICAL_DENSE_PARAMETER_COUNT:
|
| 88 |
+
raise RuntimeError(
|
| 89 |
+
"canonical PetitGPT parameter count mismatch: "
|
| 90 |
+
f"actual={actual:,}, expected={CANONICAL_DENSE_PARAMETER_COUNT:,}"
|
| 91 |
+
)
|
| 92 |
+
|
| 93 |
+
return {
|
| 94 |
+
"status": "passed",
|
| 95 |
+
"counting_method": "unique_parameter_objects_excluding_buffers",
|
| 96 |
+
"actual_total": actual,
|
| 97 |
+
"actual_trainable": trainable,
|
| 98 |
+
"derived_expected_total": expected,
|
| 99 |
+
"canonical_parameterization": canonical,
|
| 100 |
+
"canonical_expected_total": CANONICAL_DENSE_PARAMETER_COUNT,
|
| 101 |
+
"canonical_match": canonical and actual == CANONICAL_DENSE_PARAMETER_COUNT,
|
| 102 |
+
}
|
| 103 |
+
|
| 104 |
+
|
| 105 |
+
def gpt_config_from_checkpoint_dict(cfg_dict: dict) -> GPTConfig:
|
| 106 |
+
"""Rebuild a GPTConfig from a checkpoint's serialized config dict.
|
| 107 |
+
|
| 108 |
+
Pre-GQA checkpoints carry no n_kv_heads; absence means plain MHA
|
| 109 |
+
(n_kv_heads == n_heads), whose fused-QKV weight layout is unchanged.
|
| 110 |
+
"""
|
| 111 |
+
cfg_dict = dict(cfg_dict)
|
| 112 |
+
cfg_dict.setdefault("n_kv_heads", cfg_dict["n_heads"])
|
| 113 |
+
return GPTConfig(**cfg_dict)
|
| 114 |
+
|
| 115 |
+
|
| 116 |
+
class RMSNorm(nn.Module):
|
| 117 |
+
def __init__(self, dim: int, eps: float = 1e-6):
|
| 118 |
+
super().__init__()
|
| 119 |
+
self.eps = eps
|
| 120 |
+
self.weight = nn.Parameter(torch.ones(dim))
|
| 121 |
+
|
| 122 |
+
def forward(self, x: torch.Tensor) -> torch.Tensor:
|
| 123 |
+
# x: [B, T, C]
|
| 124 |
+
rms = x.pow(2).mean(dim=-1, keepdim=True).add(self.eps).rsqrt()
|
| 125 |
+
return x * rms * self.weight
|
| 126 |
+
|
| 127 |
+
|
| 128 |
+
def _rotate_half(x: torch.Tensor) -> torch.Tensor:
|
| 129 |
+
# x: [..., D]. Half-split layout (Llama/GPT-NeoX): pairs are (i, i+D/2),
|
| 130 |
+
# matching the `cat([freqs, freqs])` cos/sin cache below.
|
| 131 |
+
half = x.shape[-1] // 2
|
| 132 |
+
x1 = x[..., :half]
|
| 133 |
+
x2 = x[..., half:]
|
| 134 |
+
return torch.cat((-x2, x1), dim=-1)
|
| 135 |
+
|
| 136 |
+
|
| 137 |
+
class RotaryEmbedding(nn.Module):
|
| 138 |
+
"""Precomputes RoPE cos/sin caches up to max_seq_len."""
|
| 139 |
+
|
| 140 |
+
def __init__(self, head_dim: int, max_seq_len: int, theta: float = 10000.0, pct: float = 1.0):
|
| 141 |
+
super().__init__()
|
| 142 |
+
if head_dim % 2 != 0:
|
| 143 |
+
raise ValueError(f"RoPE requires even head_dim, got {head_dim}")
|
| 144 |
+
self.head_dim = int(head_dim)
|
| 145 |
+
self.max_seq_len = int(max_seq_len)
|
| 146 |
+
self.theta = float(theta)
|
| 147 |
+
self.pct = float(pct)
|
| 148 |
+
|
| 149 |
+
rope_dim = int(self.head_dim * self.pct)
|
| 150 |
+
rope_dim = rope_dim - (rope_dim % 2)
|
| 151 |
+
rope_dim = max(0, min(rope_dim, self.head_dim))
|
| 152 |
+
self.rope_dim = rope_dim
|
| 153 |
+
|
| 154 |
+
if self.rope_dim > 0:
|
| 155 |
+
inv_freq = 1.0 / (
|
| 156 |
+
self.theta ** (torch.arange(0, self.rope_dim, 2).float() / self.rope_dim)
|
| 157 |
+
)
|
| 158 |
+
t = torch.arange(self.max_seq_len, dtype=torch.float32)
|
| 159 |
+
freqs = torch.outer(t, inv_freq) # [T, rope_dim/2]
|
| 160 |
+
emb = torch.cat([freqs, freqs], dim=-1) # [T, rope_dim]
|
| 161 |
+
cos = emb.cos()
|
| 162 |
+
sin = emb.sin()
|
| 163 |
+
else:
|
| 164 |
+
cos = torch.empty(self.max_seq_len, 0, dtype=torch.float32)
|
| 165 |
+
sin = torch.empty(self.max_seq_len, 0, dtype=torch.float32)
|
| 166 |
+
|
| 167 |
+
self.register_buffer("cos_cached", cos, persistent=False)
|
| 168 |
+
self.register_buffer("sin_cached", sin, persistent=False)
|
| 169 |
+
|
| 170 |
+
def forward(
|
| 171 |
+
self, q: torch.Tensor, k: torch.Tensor, seq_len: int, offset: int = 0
|
| 172 |
+
) -> tuple[torch.Tensor, torch.Tensor]:
|
| 173 |
+
"""Apply RoPE to q,k. q,k: [B, nH, T, Hd].
|
| 174 |
+
|
| 175 |
+
`offset` is the absolute position of the first token in q,k — nonzero
|
| 176 |
+
during KV-cached incremental decoding, where the new tokens sit at
|
| 177 |
+
positions [offset, offset+seq_len).
|
| 178 |
+
"""
|
| 179 |
+
end = offset + seq_len
|
| 180 |
+
if end > self.max_seq_len:
|
| 181 |
+
raise ValueError(
|
| 182 |
+
f"position {end} exceeds max_seq_len={self.max_seq_len} for RoPE cache"
|
| 183 |
+
)
|
| 184 |
+
if self.rope_dim == 0:
|
| 185 |
+
return q, k
|
| 186 |
+
|
| 187 |
+
cos = self.cos_cached[offset:end].to(dtype=q.dtype, device=q.device) # [T, rope_dim]
|
| 188 |
+
sin = self.sin_cached[offset:end].to(dtype=q.dtype, device=q.device) # [T, rope_dim]
|
| 189 |
+
cos = cos.unsqueeze(0).unsqueeze(0) # [1,1,T,rope_dim]
|
| 190 |
+
sin = sin.unsqueeze(0).unsqueeze(0)
|
| 191 |
+
|
| 192 |
+
q1, q2 = q[..., : self.rope_dim], q[..., self.rope_dim :]
|
| 193 |
+
k1, k2 = k[..., : self.rope_dim], k[..., self.rope_dim :]
|
| 194 |
+
|
| 195 |
+
q1 = q1 * cos + _rotate_half(q1) * sin
|
| 196 |
+
k1 = k1 * cos + _rotate_half(k1) * sin
|
| 197 |
+
|
| 198 |
+
q = torch.cat([q1, q2], dim=-1)
|
| 199 |
+
k = torch.cat([k1, k2], dim=-1)
|
| 200 |
+
return q, k
|
| 201 |
+
|
| 202 |
+
|
| 203 |
+
class SwiGLU(nn.Module):
|
| 204 |
+
def __init__(self, cfg: GPTConfig):
|
| 205 |
+
super().__init__()
|
| 206 |
+
self.w1 = nn.Linear(cfg.d_model, cfg.d_ff, bias=False)
|
| 207 |
+
self.w3 = nn.Linear(cfg.d_model, cfg.d_ff, bias=False)
|
| 208 |
+
self.w2 = nn.Linear(cfg.d_ff, cfg.d_model, bias=False)
|
| 209 |
+
self.drop = nn.Dropout(cfg.dropout)
|
| 210 |
+
|
| 211 |
+
def forward(self, x: torch.Tensor) -> torch.Tensor:
|
| 212 |
+
x = F.silu(self.w1(x)) * self.w3(x)
|
| 213 |
+
x = self.w2(x)
|
| 214 |
+
return self.drop(x)
|
| 215 |
+
|
| 216 |
+
|
| 217 |
+
class CausalSelfAttention(nn.Module):
|
| 218 |
+
def __init__(self, cfg: GPTConfig):
|
| 219 |
+
super().__init__()
|
| 220 |
+
assert cfg.d_model % cfg.n_heads == 0
|
| 221 |
+
assert cfg.n_heads % cfg.n_kv_heads == 0
|
| 222 |
+
self.cfg = cfg
|
| 223 |
+
self.head_dim = cfg.d_model // cfg.n_heads
|
| 224 |
+
self.kv_dim = cfg.n_kv_heads * self.head_dim
|
| 225 |
+
|
| 226 |
+
# QKV fused: one matmul instead of three. K/V carry n_kv_heads (GQA);
|
| 227 |
+
# n_kv_heads == n_heads is plain MHA with the historical 3*d_model layout.
|
| 228 |
+
self.qkv = nn.Linear(cfg.d_model, cfg.d_model + 2 * self.kv_dim, bias=False)
|
| 229 |
+
# residual branch output projection
|
| 230 |
+
self.proj = nn.Linear(cfg.d_model, cfg.d_model, bias=False)
|
| 231 |
+
self.drop = nn.Dropout(cfg.dropout)
|
| 232 |
+
|
| 233 |
+
self.rope = RotaryEmbedding(
|
| 234 |
+
head_dim=self.head_dim,
|
| 235 |
+
max_seq_len=cfg.max_seq_len,
|
| 236 |
+
theta=cfg.rope_theta,
|
| 237 |
+
pct=cfg.rope_pct,
|
| 238 |
+
)
|
| 239 |
+
|
| 240 |
+
@staticmethod
|
| 241 |
+
def _incremental_mask(T: int, past_len: int, device: torch.device) -> torch.Tensor:
|
| 242 |
+
"""Bottom-right causal mask [T, past_len+T] (True = attend) for decoding
|
| 243 |
+
T new queries against past_len cached keys plus the new keys."""
|
| 244 |
+
q_pos = past_len + torch.arange(T, device=device)
|
| 245 |
+
k_pos = torch.arange(past_len + T, device=device)
|
| 246 |
+
return k_pos[None, :] <= q_pos[:, None]
|
| 247 |
+
|
| 248 |
+
def forward(
|
| 249 |
+
self,
|
| 250 |
+
x: torch.Tensor,
|
| 251 |
+
past_kv: tuple[torch.Tensor, torch.Tensor] | None = None,
|
| 252 |
+
use_cache: bool = False,
|
| 253 |
+
):
|
| 254 |
+
"""x: [B, T, C]. With no cache this is byte-identical to a plain causal
|
| 255 |
+
forward and returns the output tensor. With `use_cache` (or a supplied
|
| 256 |
+
`past_kv`) it also returns the updated (k, v) for this layer."""
|
| 257 |
+
B, T, C = x.shape
|
| 258 |
+
past_len = 0 if past_kv is None else past_kv[0].size(2)
|
| 259 |
+
if past_len + T > self.cfg.max_seq_len:
|
| 260 |
+
raise ValueError(
|
| 261 |
+
f"cache length {past_len + T} exceeds max_seq_len={self.cfg.max_seq_len}"
|
| 262 |
+
)
|
| 263 |
+
|
| 264 |
+
qkv = self.qkv(x) # [B, T, C + 2*kv_dim]
|
| 265 |
+
q, k, v = qkv.split([C, self.kv_dim, self.kv_dim], dim=-1)
|
| 266 |
+
|
| 267 |
+
q = q.view(B, T, self.cfg.n_heads, self.head_dim).transpose(1, 2) # [B,nH,T,Hd]
|
| 268 |
+
k = k.view(B, T, self.cfg.n_kv_heads, self.head_dim).transpose(1, 2) # [B,nKV,T,Hd]
|
| 269 |
+
v = v.view(B, T, self.cfg.n_kv_heads, self.head_dim).transpose(1, 2)
|
| 270 |
+
|
| 271 |
+
# RoPE rotates only the new tokens, at their absolute positions.
|
| 272 |
+
q, k = self.rope(q, k, seq_len=T, offset=past_len)
|
| 273 |
+
|
| 274 |
+
# Prepend cached keys/values (already rotated when they were new). The
|
| 275 |
+
# cache stays un-expanded at n_kv_heads so its memory reflects GQA.
|
| 276 |
+
if past_kv is not None:
|
| 277 |
+
k = torch.cat([past_kv[0], k], dim=2)
|
| 278 |
+
v = torch.cat([past_kv[1], v], dim=2)
|
| 279 |
+
present = (k, v) if use_cache else None
|
| 280 |
+
|
| 281 |
+
# Expand grouped KV heads to the full head count for attention. KV head g
|
| 282 |
+
# serves query heads [g*rep, (g+1)*rep) — repeat_interleave matches SDPA's
|
| 283 |
+
# enable_gqa grouping (torch >= 2.5), which can replace this someday.
|
| 284 |
+
if self.cfg.n_kv_heads != self.cfg.n_heads:
|
| 285 |
+
rep = self.cfg.n_heads // self.cfg.n_kv_heads
|
| 286 |
+
k = k.repeat_interleave(rep, dim=1)
|
| 287 |
+
v = v.repeat_interleave(rep, dim=1)
|
| 288 |
+
|
| 289 |
+
dropout_p = float(self.cfg.dropout) if (self.training and self.cfg.dropout > 0) else 0.0
|
| 290 |
+
|
| 291 |
+
if q.device.type == "cuda":
|
| 292 |
+
if past_len == 0:
|
| 293 |
+
y = F.scaled_dot_product_attention(
|
| 294 |
+
q, k, v, attn_mask=None, dropout_p=dropout_p, is_causal=True
|
| 295 |
+
)
|
| 296 |
+
else:
|
| 297 |
+
y = F.scaled_dot_product_attention(
|
| 298 |
+
q,
|
| 299 |
+
k,
|
| 300 |
+
v,
|
| 301 |
+
attn_mask=self._incremental_mask(T, past_len, q.device),
|
| 302 |
+
dropout_p=dropout_p,
|
| 303 |
+
)
|
| 304 |
+
else:
|
| 305 |
+
scale = 1.0 / math.sqrt(self.head_dim)
|
| 306 |
+
att = torch.matmul(q * scale, k.transpose(-2, -1)) # [B,nH,T,past_len+T]
|
| 307 |
+
if past_len == 0:
|
| 308 |
+
mask = torch.triu(torch.ones((T, T), device=q.device, dtype=torch.bool), diagonal=1)
|
| 309 |
+
att = att.masked_fill(mask, float("-inf"))
|
| 310 |
+
else:
|
| 311 |
+
allow = self._incremental_mask(T, past_len, q.device) # [T, past_len+T]
|
| 312 |
+
att = att.masked_fill(~allow, float("-inf"))
|
| 313 |
+
att = F.softmax(att, dim=-1)
|
| 314 |
+
if dropout_p > 0.0:
|
| 315 |
+
att = F.dropout(att, p=dropout_p)
|
| 316 |
+
y = torch.matmul(att, v)
|
| 317 |
+
|
| 318 |
+
y = y.transpose(1, 2).contiguous().view(B, T, C)
|
| 319 |
+
y = self.drop(self.proj(y))
|
| 320 |
+
if use_cache:
|
| 321 |
+
return y, present
|
| 322 |
+
return y
|
| 323 |
+
|
| 324 |
+
|
| 325 |
+
class Block(nn.Module):
|
| 326 |
+
def __init__(self, cfg: GPTConfig):
|
| 327 |
+
super().__init__()
|
| 328 |
+
self.norm1 = RMSNorm(cfg.d_model)
|
| 329 |
+
self.attn = CausalSelfAttention(cfg)
|
| 330 |
+
self.norm2 = RMSNorm(cfg.d_model)
|
| 331 |
+
self.mlp = SwiGLU(cfg)
|
| 332 |
+
|
| 333 |
+
def forward(
|
| 334 |
+
self,
|
| 335 |
+
x: torch.Tensor,
|
| 336 |
+
past_kv: tuple[torch.Tensor, torch.Tensor] | None = None,
|
| 337 |
+
use_cache: bool = False,
|
| 338 |
+
):
|
| 339 |
+
if use_cache or past_kv is not None:
|
| 340 |
+
attn_out, present = self.attn(self.norm1(x), past_kv=past_kv, use_cache=True)
|
| 341 |
+
x = x + attn_out
|
| 342 |
+
x = x + self.mlp(self.norm2(x))
|
| 343 |
+
return x, present
|
| 344 |
+
x = x + self.attn(self.norm1(x))
|
| 345 |
+
x = x + self.mlp(self.norm2(x))
|
| 346 |
+
return x
|
| 347 |
+
|
| 348 |
+
|
| 349 |
+
class GPT(nn.Module):
|
| 350 |
+
def __init__(self, cfg: GPTConfig):
|
| 351 |
+
super().__init__()
|
| 352 |
+
self.cfg = cfg
|
| 353 |
+
|
| 354 |
+
self.tok_emb = nn.Embedding(cfg.vocab_size, cfg.d_model)
|
| 355 |
+
self.drop = nn.Dropout(cfg.dropout)
|
| 356 |
+
|
| 357 |
+
self.blocks = nn.ModuleList([Block(cfg) for _ in range(cfg.n_layers)])
|
| 358 |
+
self.norm_f = RMSNorm(cfg.d_model)
|
| 359 |
+
|
| 360 |
+
self.lm_head = nn.Linear(cfg.d_model, cfg.vocab_size, bias=False)
|
| 361 |
+
if cfg.tie_embeddings:
|
| 362 |
+
self.lm_head.weight = self.tok_emb.weight
|
| 363 |
+
|
| 364 |
+
# Base init everywhere...
|
| 365 |
+
self.apply(self._init_weights)
|
| 366 |
+
# ...then scale init ONLY on residual-branch output projections: attn.proj and mlp.w2
|
| 367 |
+
self._init_residual_projections()
|
| 368 |
+
|
| 369 |
+
def _init_weights(self, m: nn.Module):
|
| 370 |
+
if isinstance(m, (nn.Linear, nn.Embedding)):
|
| 371 |
+
torch.nn.init.normal_(m.weight, mean=0.0, std=0.02)
|
| 372 |
+
|
| 373 |
+
def _init_residual_projections(self):
|
| 374 |
+
std = 0.02 / math.sqrt(2.0 * float(self.cfg.n_layers))
|
| 375 |
+
for blk in self.blocks:
|
| 376 |
+
torch.nn.init.normal_(blk.attn.proj.weight, mean=0.0, std=std)
|
| 377 |
+
torch.nn.init.normal_(blk.mlp.w2.weight, mean=0.0, std=std)
|
| 378 |
+
|
| 379 |
+
def forward(
|
| 380 |
+
self,
|
| 381 |
+
input_ids: torch.Tensor,
|
| 382 |
+
past_kv: list[tuple[torch.Tensor, torch.Tensor]] | None = None,
|
| 383 |
+
use_cache: bool = False,
|
| 384 |
+
):
|
| 385 |
+
"""Default call `model(input_ids)` returns logits [B, T, V] — unchanged.
|
| 386 |
+
|
| 387 |
+
For incremental decoding, pass `use_cache=True` to also get a per-layer
|
| 388 |
+
list of (k, v) tensors, and feed it back as `past_kv` with only the new
|
| 389 |
+
token(s) on the next call. See `generate`.
|
| 390 |
+
"""
|
| 391 |
+
B, T = input_ids.shape
|
| 392 |
+
past_len = 0 if past_kv is None else past_kv[0][0].size(2)
|
| 393 |
+
if past_len + T > self.cfg.max_seq_len:
|
| 394 |
+
raise ValueError(
|
| 395 |
+
f"T={T} with cache={past_len} exceeds max_seq_len={self.cfg.max_seq_len}"
|
| 396 |
+
)
|
| 397 |
+
if T < 1:
|
| 398 |
+
raise ValueError("Empty sequence")
|
| 399 |
+
caching = use_cache or (past_kv is not None)
|
| 400 |
+
|
| 401 |
+
x = self.tok_emb(input_ids)
|
| 402 |
+
x = self.drop(x)
|
| 403 |
+
presents: list[tuple[torch.Tensor, torch.Tensor]] = []
|
| 404 |
+
for i, blk in enumerate(self.blocks):
|
| 405 |
+
layer_past = past_kv[i] if past_kv is not None else None
|
| 406 |
+
if caching:
|
| 407 |
+
x, present = blk(x, past_kv=layer_past, use_cache=True)
|
| 408 |
+
presents.append(present)
|
| 409 |
+
else:
|
| 410 |
+
x = blk(x)
|
| 411 |
+
x = self.norm_f(x)
|
| 412 |
+
logits = self.lm_head(x)
|
| 413 |
+
if caching:
|
| 414 |
+
return logits, presents
|
| 415 |
+
return logits
|
| 416 |
+
|
| 417 |
+
@staticmethod
|
| 418 |
+
def _sample_token(
|
| 419 |
+
logits: torch.Tensor, temperature: float, top_k: int, top_p: float
|
| 420 |
+
) -> torch.Tensor:
|
| 421 |
+
"""logits: [B, V] -> next token [B, 1]. temperature<=0 is greedy."""
|
| 422 |
+
if temperature <= 0:
|
| 423 |
+
return logits.argmax(dim=-1, keepdim=True)
|
| 424 |
+
logits = logits / temperature
|
| 425 |
+
if top_k and top_k > 0:
|
| 426 |
+
k = min(int(top_k), logits.size(-1))
|
| 427 |
+
thresh = torch.topk(logits, k, dim=-1).values[:, -1, None]
|
| 428 |
+
logits = logits.masked_fill(logits < thresh, float("-inf"))
|
| 429 |
+
if top_p and top_p < 1.0:
|
| 430 |
+
sorted_logits, sorted_idx = torch.sort(logits, descending=True, dim=-1)
|
| 431 |
+
cum = torch.softmax(sorted_logits, dim=-1).cumsum(dim=-1)
|
| 432 |
+
drop_sorted = cum > top_p
|
| 433 |
+
drop_sorted[..., 0] = False
|
| 434 |
+
drop = torch.zeros_like(drop_sorted).scatter(-1, sorted_idx, drop_sorted)
|
| 435 |
+
logits = logits.masked_fill(drop, float("-inf"))
|
| 436 |
+
probs = torch.softmax(logits, dim=-1)
|
| 437 |
+
return torch.multinomial(probs, num_samples=1)
|
| 438 |
+
|
| 439 |
+
@torch.no_grad()
|
| 440 |
+
def generate(
|
| 441 |
+
self,
|
| 442 |
+
input_ids: torch.Tensor,
|
| 443 |
+
max_new_tokens: int,
|
| 444 |
+
*,
|
| 445 |
+
temperature: float = 1.0,
|
| 446 |
+
top_k: int = 0,
|
| 447 |
+
top_p: float = 1.0,
|
| 448 |
+
eos_id: int | None = None,
|
| 449 |
+
) -> torch.Tensor:
|
| 450 |
+
"""KV-cached incremental decoding. input_ids: [B, T] -> [B, T + n].
|
| 451 |
+
|
| 452 |
+
Prefills the prompt once, then feeds one new token per step against the
|
| 453 |
+
cache (O(T) forwards of length 1) instead of re-running the full growing
|
| 454 |
+
sequence each step. Stops early if all rows emit `eos_id`.
|
| 455 |
+
"""
|
| 456 |
+
was_training = self.training
|
| 457 |
+
self.eval()
|
| 458 |
+
logits, past = self.forward(input_ids, use_cache=True)
|
| 459 |
+
out = input_ids
|
| 460 |
+
for _ in range(int(max_new_tokens)):
|
| 461 |
+
next_tok = self._sample_token(logits[:, -1, :], temperature, top_k, top_p)
|
| 462 |
+
out = torch.cat([out, next_tok], dim=1)
|
| 463 |
+
if eos_id is not None and bool((next_tok.squeeze(1) == eos_id).all()):
|
| 464 |
+
break
|
| 465 |
+
if out.size(1) >= self.cfg.max_seq_len:
|
| 466 |
+
break
|
| 467 |
+
logits, past = self.forward(next_tok, past_kv=past, use_cache=True)
|
| 468 |
+
if was_training:
|
| 469 |
+
self.train()
|
| 470 |
+
return out
|
src/special_tokens.py
ADDED
|
@@ -0,0 +1,138 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Canonical special-token IDs for the whole pipeline.
|
| 2 |
+
|
| 3 |
+
Every training/data script hardcodes these IDs (see CLAUDE.md):
|
| 4 |
+
|
| 5 |
+
[PAD]=0 [UNK]=1 [BOS]=2 [EOS]=3
|
| 6 |
+
<|system|>=4 <|user|>=5 <|assistant|>=6
|
| 7 |
+
|
| 8 |
+
The three role tokens delimit chat turns at the *token* level (see
|
| 9 |
+
src/chat_template.py) — BPE can never merge across a special token, so
|
| 10 |
+
training-time and inference-time encodings agree by construction.
|
| 11 |
+
|
| 12 |
+
`assert_special_token_ids` turns the silent assumption into a loud startup
|
| 13 |
+
check: if the tokenizer at --tokenizer_path is ever retrained and the IDs move,
|
| 14 |
+
scripts fail immediately instead of training with a misaligned loss mask or a
|
| 15 |
+
broken EOS stop condition.
|
| 16 |
+
"""
|
| 17 |
+
|
| 18 |
+
from __future__ import annotations
|
| 19 |
+
|
| 20 |
+
import json
|
| 21 |
+
from os import PathLike
|
| 22 |
+
|
| 23 |
+
PAD_ID = 0
|
| 24 |
+
UNK_ID = 1
|
| 25 |
+
BOS_ID = 2
|
| 26 |
+
EOS_ID = 3
|
| 27 |
+
SYSTEM_ID = 4
|
| 28 |
+
USER_ID = 5
|
| 29 |
+
ASSISTANT_ID = 6
|
| 30 |
+
CANONICAL_VOCAB_SIZE = 32_000
|
| 31 |
+
|
| 32 |
+
PAD_TOKEN = "[PAD]"
|
| 33 |
+
UNK_TOKEN = "[UNK]"
|
| 34 |
+
BOS_TOKEN = "[BOS]"
|
| 35 |
+
EOS_TOKEN = "[EOS]"
|
| 36 |
+
SYSTEM_TOKEN = "<|system|>"
|
| 37 |
+
USER_TOKEN = "<|user|>"
|
| 38 |
+
ASSISTANT_TOKEN = "<|assistant|>"
|
| 39 |
+
|
| 40 |
+
SPECIAL_TOKEN_IDS: dict[str, int] = {
|
| 41 |
+
PAD_TOKEN: PAD_ID,
|
| 42 |
+
UNK_TOKEN: UNK_ID,
|
| 43 |
+
BOS_TOKEN: BOS_ID,
|
| 44 |
+
EOS_TOKEN: EOS_ID,
|
| 45 |
+
SYSTEM_TOKEN: SYSTEM_ID,
|
| 46 |
+
USER_TOKEN: USER_ID,
|
| 47 |
+
ASSISTANT_TOKEN: ASSISTANT_ID,
|
| 48 |
+
}
|
| 49 |
+
|
| 50 |
+
# Ordered by ID — the exact list BpeTrainer must receive so IDs come out right.
|
| 51 |
+
SPECIAL_TOKENS: list[str] = sorted(SPECIAL_TOKEN_IDS, key=SPECIAL_TOKEN_IDS.get)
|
| 52 |
+
|
| 53 |
+
|
| 54 |
+
def assert_special_token_ids(tokenizer_path: str) -> None:
|
| 55 |
+
"""Validate the exact registered-special-token contract.
|
| 56 |
+
|
| 57 |
+
Reads the file as plain JSON (no `tokenizers` import needed) so it is cheap
|
| 58 |
+
to call from any script entry point. Merely finding the strings in the BPE
|
| 59 |
+
vocabulary is insufficient: all seven must be registered as special tokens,
|
| 60 |
+
and no eighth registered special token is allowed.
|
| 61 |
+
"""
|
| 62 |
+
with open(tokenizer_path, encoding="utf-8") as f:
|
| 63 |
+
obj = json.load(f)
|
| 64 |
+
|
| 65 |
+
special_entries = [
|
| 66 |
+
entry
|
| 67 |
+
for entry in (obj.get("added_tokens") or [])
|
| 68 |
+
if isinstance(entry, dict) and entry.get("special") is True
|
| 69 |
+
]
|
| 70 |
+
registered_tokens = [entry.get("content") for entry in special_entries]
|
| 71 |
+
if len(special_entries) != len(SPECIAL_TOKEN_IDS) or set(registered_tokens) != set(
|
| 72 |
+
SPECIAL_TOKEN_IDS
|
| 73 |
+
):
|
| 74 |
+
raise ValueError(
|
| 75 |
+
f"{tokenizer_path} must register exactly these seven special tokens: "
|
| 76 |
+
f"{list(SPECIAL_TOKEN_IDS)}; got {registered_tokens!r}"
|
| 77 |
+
)
|
| 78 |
+
|
| 79 |
+
added = {entry["content"]: entry.get("id") for entry in special_entries}
|
| 80 |
+
for token, expected in SPECIAL_TOKEN_IDS.items():
|
| 81 |
+
got = added.get(token)
|
| 82 |
+
if got != expected:
|
| 83 |
+
raise ValueError(
|
| 84 |
+
f"special token {token!r} has id {got!r} in {tokenizer_path}, but the "
|
| 85 |
+
f"pipeline hardcodes {expected}. Retrain the tokenizer with "
|
| 86 |
+
f"tokenizer/tokenizer_training/train_tokenizer.py (its --strict_special_ids "
|
| 87 |
+
f"default enforces this layout) or reconcile src/special_tokens.py."
|
| 88 |
+
)
|
| 89 |
+
|
| 90 |
+
|
| 91 |
+
def assert_tokenizer_contract(tokenizer_path: str | PathLike[str]) -> None:
|
| 92 |
+
"""Fail fast unless ``tokenizer.json`` is the canonical production artifact.
|
| 93 |
+
|
| 94 |
+
This is deliberately stricter than :func:`assert_special_token_ids`, which
|
| 95 |
+
remains useful for tiny unit-test tokenizers. Production entry points must
|
| 96 |
+
call this function so a valid-looking seven-token map cannot hide an old
|
| 97 |
+
vocabulary, normalizer, prefix-space rule, or automatic BOS/EOS processor.
|
| 98 |
+
"""
|
| 99 |
+
path = str(tokenizer_path)
|
| 100 |
+
assert_special_token_ids(path)
|
| 101 |
+
with open(path, encoding="utf-8") as f:
|
| 102 |
+
obj = json.load(f)
|
| 103 |
+
|
| 104 |
+
model = obj.get("model") or {}
|
| 105 |
+
if model.get("type") != "BPE":
|
| 106 |
+
raise ValueError(f"{path} must use a BPE model; got {model.get('type')!r}")
|
| 107 |
+
if model.get("unk_token") != UNK_TOKEN:
|
| 108 |
+
raise ValueError(
|
| 109 |
+
f"{path} BPE unk_token must be {UNK_TOKEN!r}; got {model.get('unk_token')!r}"
|
| 110 |
+
)
|
| 111 |
+
|
| 112 |
+
model_vocab = model.get("vocab") or {}
|
| 113 |
+
token_ids = set(model_vocab.values())
|
| 114 |
+
token_ids.update(
|
| 115 |
+
entry.get("id")
|
| 116 |
+
for entry in (obj.get("added_tokens") or [])
|
| 117 |
+
if isinstance(entry, dict) and isinstance(entry.get("id"), int)
|
| 118 |
+
)
|
| 119 |
+
expected_ids = set(range(CANONICAL_VOCAB_SIZE))
|
| 120 |
+
if token_ids != expected_ids:
|
| 121 |
+
raise ValueError(
|
| 122 |
+
f"{path} runtime vocab IDs must be exactly 0..{CANONICAL_VOCAB_SIZE - 1} "
|
| 123 |
+
f"(vocab_size={CANONICAL_VOCAB_SIZE}); got {len(token_ids)} unique IDs"
|
| 124 |
+
)
|
| 125 |
+
|
| 126 |
+
if obj.get("normalizer") is not None:
|
| 127 |
+
raise ValueError(f"{path} must not configure a tokenizer normalizer")
|
| 128 |
+
if obj.get("post_processor") is not None:
|
| 129 |
+
raise ValueError(f"{path} must not configure an automatic BOS/EOS post-processor")
|
| 130 |
+
|
| 131 |
+
pre = obj.get("pre_tokenizer") or {}
|
| 132 |
+
if pre.get("type") != "ByteLevel" or pre.get("add_prefix_space") is not False:
|
| 133 |
+
raise ValueError(
|
| 134 |
+
f"{path} pre_tokenizer must be ByteLevel(add_prefix_space=False); got {pre!r}"
|
| 135 |
+
)
|
| 136 |
+
decoder = obj.get("decoder") or {}
|
| 137 |
+
if decoder.get("type") != "ByteLevel":
|
| 138 |
+
raise ValueError(f"{path} decoder must be ByteLevel; got {decoder!r}")
|
tables/ASSISTANT_RESULTS_VERSIONED.csv
ADDED
|
@@ -0,0 +1,145 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
score_version,model,slice,axis,true,false,unknown,n,bound_lower,bound_upper,bound_lower_pct,bound_upper_pct,finite_execution_categories,evidence_id
|
| 2 |
+
assistant_review_fable_v1,alpha075,old_qa41,content,7,33,1,41,,,,,,EVID-ASSIST-V1
|
| 3 |
+
assistant_review_fable_v1,alpha075,old_qa41,joint,7,33,1,41,,,,,,EVID-ASSIST-V1
|
| 4 |
+
assistant_review_fable_v1,alpha075,old_qa41,format,0,1,0,1,,,,,,EVID-ASSIST-V1
|
| 5 |
+
assistant_review_fable_v1,alpha075,old_practical38,content,10,26,2,38,,,,,,EVID-ASSIST-V1
|
| 6 |
+
assistant_review_fable_v1,alpha075,old_practical38,joint,10,26,2,38,,,,,,EVID-ASSIST-V1
|
| 7 |
+
assistant_review_fable_v1,alpha075,old_practical38,format,20,1,0,21,,,,,,EVID-ASSIST-V1
|
| 8 |
+
assistant_review_fable_v1,alpha075,old_python14,content,0,14,0,14,,,,,,EVID-ASSIST-V1
|
| 9 |
+
assistant_review_fable_v1,alpha075,old_python14,joint,0,14,0,14,,,,,,EVID-ASSIST-V1
|
| 10 |
+
assistant_review_fable_v1,alpha075,old_python14,format,3,1,0,4,,,,,,EVID-ASSIST-V1
|
| 11 |
+
assistant_review_fable_v1,alpha075,old_python14,interface,11,3,0,14,,,,,,EVID-ASSIST-V1
|
| 12 |
+
assistant_review_fable_v1,alpha075,old_python14,historical_finite_categories,,,,,,,,,"{""not_recorded"": 14}",EVID-ASSIST-V1
|
| 13 |
+
assistant_review_fable_v1,alpha075,new_natural64,content,2,62,0,64,,,,,,EVID-ASSIST-V1
|
| 14 |
+
assistant_review_fable_v1,alpha075,new_natural64,joint,2,62,0,64,,,,,,EVID-ASSIST-V1
|
| 15 |
+
assistant_review_fable_v1,alpha075,new_natural64,format,7,11,0,18,,,,,,EVID-ASSIST-V1
|
| 16 |
+
assistant_review_fable_v1,alpha075,new_python32,content,0,32,0,32,,,,,,EVID-ASSIST-V1
|
| 17 |
+
assistant_review_fable_v1,alpha075,new_python32,joint,0,32,0,32,,,,,,EVID-ASSIST-V1
|
| 18 |
+
assistant_review_fable_v1,alpha075,new_python32,format,0,0,0,0,,,,,,EVID-ASSIST-V1
|
| 19 |
+
assistant_review_fable_v1,alpha075,new_python32,interface,31,1,0,32,,,,,,EVID-ASSIST-V1
|
| 20 |
+
assistant_review_fable_v1,alpha075,new_python32,historical_finite_categories,,,,,,,,,"{""fail"": 25, ""fail_or_unextractable"": 1, ""tool_unknown"": 6}",EVID-ASSIST-V1
|
| 21 |
+
assistant_review_fable_v1,alpha075,practical27_joint,joint_group_bounds,,,,,26/81,10/27,32.098765,37.037037,,EVID-ASSIST-V1
|
| 22 |
+
assistant_review_fable_v1,alpha075,new_natural64,category:R1DEV-CONTEXT,0,16,0,16,,,,,,EVID-ASSIST-V1
|
| 23 |
+
assistant_review_fable_v1,alpha075,new_natural64,category:R1DEV-REWRITE,0,16,0,16,,,,,,EVID-ASSIST-V1
|
| 24 |
+
assistant_review_fable_v1,alpha075,new_natural64,category:R1DEV-FOLLOWUP,1,15,0,16,,,,,,EVID-ASSIST-V1
|
| 25 |
+
assistant_review_fable_v1,alpha075,new_natural64,category:R1DEV-CONDITIONAL,1,15,0,16,,,,,,EVID-ASSIST-V1
|
| 26 |
+
assistant_review_fable_v1,EXT-A,old_qa41,content,7,34,0,41,,,,,,EVID-ASSIST-V1
|
| 27 |
+
assistant_review_fable_v1,EXT-A,old_qa41,joint,7,34,0,41,,,,,,EVID-ASSIST-V1
|
| 28 |
+
assistant_review_fable_v1,EXT-A,old_qa41,format,0,1,0,1,,,,,,EVID-ASSIST-V1
|
| 29 |
+
assistant_review_fable_v1,EXT-A,old_practical38,content,18,18,2,38,,,,,,EVID-ASSIST-V1
|
| 30 |
+
assistant_review_fable_v1,EXT-A,old_practical38,joint,18,18,2,38,,,,,,EVID-ASSIST-V1
|
| 31 |
+
assistant_review_fable_v1,EXT-A,old_practical38,format,18,3,0,21,,,,,,EVID-ASSIST-V1
|
| 32 |
+
assistant_review_fable_v1,EXT-A,old_python14,content,2,12,0,14,,,,,,EVID-ASSIST-V1
|
| 33 |
+
assistant_review_fable_v1,EXT-A,old_python14,joint,2,12,0,14,,,,,,EVID-ASSIST-V1
|
| 34 |
+
assistant_review_fable_v1,EXT-A,old_python14,format,3,1,0,4,,,,,,EVID-ASSIST-V1
|
| 35 |
+
assistant_review_fable_v1,EXT-A,old_python14,interface,14,0,0,14,,,,,,EVID-ASSIST-V1
|
| 36 |
+
assistant_review_fable_v1,EXT-A,old_python14,historical_finite_categories,,,,,,,,,"{""fail"": 7, ""fail_or_unextractable"": 2, ""pass"": 3, ""tool_unknown"": 2}",EVID-ASSIST-V1
|
| 37 |
+
assistant_review_fable_v1,EXT-A,new_natural64,content,4,60,0,64,,,,,,EVID-ASSIST-V1
|
| 38 |
+
assistant_review_fable_v1,EXT-A,new_natural64,joint,4,60,0,64,,,,,,EVID-ASSIST-V1
|
| 39 |
+
assistant_review_fable_v1,EXT-A,new_natural64,format,13,5,0,18,,,,,,EVID-ASSIST-V1
|
| 40 |
+
assistant_review_fable_v1,EXT-A,new_python32,content,4,28,0,32,,,,,,EVID-ASSIST-V1
|
| 41 |
+
assistant_review_fable_v1,EXT-A,new_python32,joint,4,28,0,32,,,,,,EVID-ASSIST-V1
|
| 42 |
+
assistant_review_fable_v1,EXT-A,new_python32,format,0,0,0,0,,,,,,EVID-ASSIST-V1
|
| 43 |
+
assistant_review_fable_v1,EXT-A,new_python32,interface,32,0,0,32,,,,,,EVID-ASSIST-V1
|
| 44 |
+
assistant_review_fable_v1,EXT-A,new_python32,historical_finite_categories,,,,,,,,,"{""fail"": 19, ""fail_or_unextractable"": 4, ""pass"": 2, ""tool_unknown"": 7}",EVID-ASSIST-V1
|
| 45 |
+
assistant_review_fable_v1,EXT-A,practical27_joint,joint_group_bounds,,,,,47/81,17/27,58.024691,62.962963,,EVID-ASSIST-V1
|
| 46 |
+
assistant_review_fable_v1,EXT-A,new_natural64,category:R1DEV-CONTEXT,0,16,0,16,,,,,,EVID-ASSIST-V1
|
| 47 |
+
assistant_review_fable_v1,EXT-A,new_natural64,category:R1DEV-REWRITE,2,14,0,16,,,,,,EVID-ASSIST-V1
|
| 48 |
+
assistant_review_fable_v1,EXT-A,new_natural64,category:R1DEV-FOLLOWUP,2,14,0,16,,,,,,EVID-ASSIST-V1
|
| 49 |
+
assistant_review_fable_v1,EXT-A,new_natural64,category:R1DEV-CONDITIONAL,0,16,0,16,,,,,,EVID-ASSIST-V1
|
| 50 |
+
assistant_review_fable_v1,EXT-B,old_qa41,content,5,35,1,41,,,,,,EVID-ASSIST-V1
|
| 51 |
+
assistant_review_fable_v1,EXT-B,old_qa41,joint,5,35,1,41,,,,,,EVID-ASSIST-V1
|
| 52 |
+
assistant_review_fable_v1,EXT-B,old_qa41,format,0,1,0,1,,,,,,EVID-ASSIST-V1
|
| 53 |
+
assistant_review_fable_v1,EXT-B,old_practical38,content,6,31,1,38,,,,,,EVID-ASSIST-V1
|
| 54 |
+
assistant_review_fable_v1,EXT-B,old_practical38,joint,6,31,1,38,,,,,,EVID-ASSIST-V1
|
| 55 |
+
assistant_review_fable_v1,EXT-B,old_practical38,format,1,20,0,21,,,,,,EVID-ASSIST-V1
|
| 56 |
+
assistant_review_fable_v1,EXT-B,old_python14,content,4,10,0,14,,,,,,EVID-ASSIST-V1
|
| 57 |
+
assistant_review_fable_v1,EXT-B,old_python14,joint,3,11,0,14,,,,,,EVID-ASSIST-V1
|
| 58 |
+
assistant_review_fable_v1,EXT-B,old_python14,format,2,2,0,4,,,,,,EVID-ASSIST-V1
|
| 59 |
+
assistant_review_fable_v1,EXT-B,old_python14,interface,13,1,0,14,,,,,,EVID-ASSIST-V1
|
| 60 |
+
assistant_review_fable_v1,EXT-B,old_python14,historical_finite_categories,,,,,,,,,"{""fail"": 8, ""fail_or_unextractable"": 1, ""pass"": 4, ""tool_unknown"": 1}",EVID-ASSIST-V1
|
| 61 |
+
assistant_review_fable_v1,EXT-B,new_natural64,content,0,64,0,64,,,,,,EVID-ASSIST-V1
|
| 62 |
+
assistant_review_fable_v1,EXT-B,new_natural64,joint,0,64,0,64,,,,,,EVID-ASSIST-V1
|
| 63 |
+
assistant_review_fable_v1,EXT-B,new_natural64,format,3,15,0,18,,,,,,EVID-ASSIST-V1
|
| 64 |
+
assistant_review_fable_v1,EXT-B,new_python32,content,3,29,0,32,,,,,,EVID-ASSIST-V1
|
| 65 |
+
assistant_review_fable_v1,EXT-B,new_python32,joint,3,29,0,32,,,,,,EVID-ASSIST-V1
|
| 66 |
+
assistant_review_fable_v1,EXT-B,new_python32,format,0,0,0,0,,,,,,EVID-ASSIST-V1
|
| 67 |
+
assistant_review_fable_v1,EXT-B,new_python32,interface,31,1,0,32,,,,,,EVID-ASSIST-V1
|
| 68 |
+
assistant_review_fable_v1,EXT-B,new_python32,historical_finite_categories,,,,,,,,,"{""fail"": 21, ""fail_or_unextractable"": 2, ""pass"": 7, ""tool_unknown"": 2}",EVID-ASSIST-V1
|
| 69 |
+
assistant_review_fable_v1,EXT-B,practical27_joint,joint_group_bounds,,,,,8/81,1/9,9.876543,11.111111,,EVID-ASSIST-V1
|
| 70 |
+
assistant_review_fable_v1,EXT-B,new_natural64,category:R1DEV-CONTEXT,0,16,0,16,,,,,,EVID-ASSIST-V1
|
| 71 |
+
assistant_review_fable_v1,EXT-B,new_natural64,category:R1DEV-REWRITE,0,16,0,16,,,,,,EVID-ASSIST-V1
|
| 72 |
+
assistant_review_fable_v1,EXT-B,new_natural64,category:R1DEV-FOLLOWUP,0,16,0,16,,,,,,EVID-ASSIST-V1
|
| 73 |
+
assistant_review_fable_v1,EXT-B,new_natural64,category:R1DEV-CONDITIONAL,0,16,0,16,,,,,,EVID-ASSIST-V1
|
| 74 |
+
assistant_owner_clarification_4_v1,alpha075,old_qa41,content,6,33,2,41,,,,,,EVID-ASSIST-C4
|
| 75 |
+
assistant_owner_clarification_4_v1,alpha075,old_qa41,joint,6,33,2,41,,,,,,EVID-ASSIST-C4
|
| 76 |
+
assistant_owner_clarification_4_v1,alpha075,old_qa41,format,0,1,0,1,,,,,,EVID-ASSIST-C4
|
| 77 |
+
assistant_owner_clarification_4_v1,alpha075,old_practical38,content,10,26,2,38,,,,,,EVID-ASSIST-C4
|
| 78 |
+
assistant_owner_clarification_4_v1,alpha075,old_practical38,joint,10,26,2,38,,,,,,EVID-ASSIST-C4
|
| 79 |
+
assistant_owner_clarification_4_v1,alpha075,old_practical38,format,20,1,0,21,,,,,,EVID-ASSIST-C4
|
| 80 |
+
assistant_owner_clarification_4_v1,alpha075,old_python14,content,0,14,0,14,,,,,,EVID-ASSIST-C4
|
| 81 |
+
assistant_owner_clarification_4_v1,alpha075,old_python14,joint,0,14,0,14,,,,,,EVID-ASSIST-C4
|
| 82 |
+
assistant_owner_clarification_4_v1,alpha075,old_python14,format,3,1,0,4,,,,,,EVID-ASSIST-C4
|
| 83 |
+
assistant_owner_clarification_4_v1,alpha075,old_python14,interface,11,3,0,14,,,,,,EVID-ASSIST-C4
|
| 84 |
+
assistant_owner_clarification_4_v1,alpha075,old_python14,historical_finite_categories,,,,,,,,,"{""not_recorded"": 14}",EVID-ASSIST-C4
|
| 85 |
+
assistant_owner_clarification_4_v1,alpha075,new_natural64,content,3,61,0,64,,,,,,EVID-ASSIST-C4
|
| 86 |
+
assistant_owner_clarification_4_v1,alpha075,new_natural64,joint,3,61,0,64,,,,,,EVID-ASSIST-C4
|
| 87 |
+
assistant_owner_clarification_4_v1,alpha075,new_natural64,format,7,11,0,18,,,,,,EVID-ASSIST-C4
|
| 88 |
+
assistant_owner_clarification_4_v1,alpha075,new_python32,content,0,32,0,32,,,,,,EVID-ASSIST-C4
|
| 89 |
+
assistant_owner_clarification_4_v1,alpha075,new_python32,joint,0,32,0,32,,,,,,EVID-ASSIST-C4
|
| 90 |
+
assistant_owner_clarification_4_v1,alpha075,new_python32,format,0,0,0,0,,,,,,EVID-ASSIST-C4
|
| 91 |
+
assistant_owner_clarification_4_v1,alpha075,new_python32,interface,31,1,0,32,,,,,,EVID-ASSIST-C4
|
| 92 |
+
assistant_owner_clarification_4_v1,alpha075,new_python32,historical_finite_categories,,,,,,,,,"{""fail"": 25, ""fail_or_unextractable"": 1, ""tool_unknown"": 6}",EVID-ASSIST-C4
|
| 93 |
+
assistant_owner_clarification_4_v1,alpha075,practical27_joint,joint_group_bounds,,,,,26/81,10/27,32.098765,37.037037,,EVID-ASSIST-C4
|
| 94 |
+
assistant_owner_clarification_4_v1,alpha075,new_natural64,category:R1DEV-CONTEXT,0,16,0,16,,,,,,EVID-ASSIST-C4
|
| 95 |
+
assistant_owner_clarification_4_v1,alpha075,new_natural64,category:R1DEV-REWRITE,0,16,0,16,,,,,,EVID-ASSIST-C4
|
| 96 |
+
assistant_owner_clarification_4_v1,alpha075,new_natural64,category:R1DEV-FOLLOWUP,2,14,0,16,,,,,,EVID-ASSIST-C4
|
| 97 |
+
assistant_owner_clarification_4_v1,alpha075,new_natural64,category:R1DEV-CONDITIONAL,1,15,0,16,,,,,,EVID-ASSIST-C4
|
| 98 |
+
assistant_owner_clarification_4_v1,EXT-A,old_qa41,content,7,34,0,41,,,,,,EVID-ASSIST-C4
|
| 99 |
+
assistant_owner_clarification_4_v1,EXT-A,old_qa41,joint,7,34,0,41,,,,,,EVID-ASSIST-C4
|
| 100 |
+
assistant_owner_clarification_4_v1,EXT-A,old_qa41,format,0,1,0,1,,,,,,EVID-ASSIST-C4
|
| 101 |
+
assistant_owner_clarification_4_v1,EXT-A,old_practical38,content,18,18,2,38,,,,,,EVID-ASSIST-C4
|
| 102 |
+
assistant_owner_clarification_4_v1,EXT-A,old_practical38,joint,18,18,2,38,,,,,,EVID-ASSIST-C4
|
| 103 |
+
assistant_owner_clarification_4_v1,EXT-A,old_practical38,format,18,3,0,21,,,,,,EVID-ASSIST-C4
|
| 104 |
+
assistant_owner_clarification_4_v1,EXT-A,old_python14,content,1,13,0,14,,,,,,EVID-ASSIST-C4
|
| 105 |
+
assistant_owner_clarification_4_v1,EXT-A,old_python14,joint,1,13,0,14,,,,,,EVID-ASSIST-C4
|
| 106 |
+
assistant_owner_clarification_4_v1,EXT-A,old_python14,format,3,1,0,4,,,,,,EVID-ASSIST-C4
|
| 107 |
+
assistant_owner_clarification_4_v1,EXT-A,old_python14,interface,14,0,0,14,,,,,,EVID-ASSIST-C4
|
| 108 |
+
assistant_owner_clarification_4_v1,EXT-A,old_python14,historical_finite_categories,,,,,,,,,"{""fail"": 7, ""fail_or_unextractable"": 2, ""pass"": 3, ""tool_unknown"": 2}",EVID-ASSIST-C4
|
| 109 |
+
assistant_owner_clarification_4_v1,EXT-A,new_natural64,content,4,60,0,64,,,,,,EVID-ASSIST-C4
|
| 110 |
+
assistant_owner_clarification_4_v1,EXT-A,new_natural64,joint,4,60,0,64,,,,,,EVID-ASSIST-C4
|
| 111 |
+
assistant_owner_clarification_4_v1,EXT-A,new_natural64,format,13,5,0,18,,,,,,EVID-ASSIST-C4
|
| 112 |
+
assistant_owner_clarification_4_v1,EXT-A,new_python32,content,4,28,0,32,,,,,,EVID-ASSIST-C4
|
| 113 |
+
assistant_owner_clarification_4_v1,EXT-A,new_python32,joint,4,28,0,32,,,,,,EVID-ASSIST-C4
|
| 114 |
+
assistant_owner_clarification_4_v1,EXT-A,new_python32,format,0,0,0,0,,,,,,EVID-ASSIST-C4
|
| 115 |
+
assistant_owner_clarification_4_v1,EXT-A,new_python32,interface,32,0,0,32,,,,,,EVID-ASSIST-C4
|
| 116 |
+
assistant_owner_clarification_4_v1,EXT-A,new_python32,historical_finite_categories,,,,,,,,,"{""fail"": 19, ""fail_or_unextractable"": 4, ""pass"": 2, ""tool_unknown"": 7}",EVID-ASSIST-C4
|
| 117 |
+
assistant_owner_clarification_4_v1,EXT-A,practical27_joint,joint_group_bounds,,,,,47/81,17/27,58.024691,62.962963,,EVID-ASSIST-C4
|
| 118 |
+
assistant_owner_clarification_4_v1,EXT-A,new_natural64,category:R1DEV-CONTEXT,0,16,0,16,,,,,,EVID-ASSIST-C4
|
| 119 |
+
assistant_owner_clarification_4_v1,EXT-A,new_natural64,category:R1DEV-REWRITE,2,14,0,16,,,,,,EVID-ASSIST-C4
|
| 120 |
+
assistant_owner_clarification_4_v1,EXT-A,new_natural64,category:R1DEV-FOLLOWUP,2,14,0,16,,,,,,EVID-ASSIST-C4
|
| 121 |
+
assistant_owner_clarification_4_v1,EXT-A,new_natural64,category:R1DEV-CONDITIONAL,0,16,0,16,,,,,,EVID-ASSIST-C4
|
| 122 |
+
assistant_owner_clarification_4_v1,EXT-B,old_qa41,content,5,35,1,41,,,,,,EVID-ASSIST-C4
|
| 123 |
+
assistant_owner_clarification_4_v1,EXT-B,old_qa41,joint,5,35,1,41,,,,,,EVID-ASSIST-C4
|
| 124 |
+
assistant_owner_clarification_4_v1,EXT-B,old_qa41,format,0,1,0,1,,,,,,EVID-ASSIST-C4
|
| 125 |
+
assistant_owner_clarification_4_v1,EXT-B,old_practical38,content,6,31,1,38,,,,,,EVID-ASSIST-C4
|
| 126 |
+
assistant_owner_clarification_4_v1,EXT-B,old_practical38,joint,6,31,1,38,,,,,,EVID-ASSIST-C4
|
| 127 |
+
assistant_owner_clarification_4_v1,EXT-B,old_practical38,format,1,20,0,21,,,,,,EVID-ASSIST-C4
|
| 128 |
+
assistant_owner_clarification_4_v1,EXT-B,old_python14,content,4,10,0,14,,,,,,EVID-ASSIST-C4
|
| 129 |
+
assistant_owner_clarification_4_v1,EXT-B,old_python14,joint,3,11,0,14,,,,,,EVID-ASSIST-C4
|
| 130 |
+
assistant_owner_clarification_4_v1,EXT-B,old_python14,format,2,2,0,4,,,,,,EVID-ASSIST-C4
|
| 131 |
+
assistant_owner_clarification_4_v1,EXT-B,old_python14,interface,13,1,0,14,,,,,,EVID-ASSIST-C4
|
| 132 |
+
assistant_owner_clarification_4_v1,EXT-B,old_python14,historical_finite_categories,,,,,,,,,"{""fail"": 8, ""fail_or_unextractable"": 1, ""pass"": 4, ""tool_unknown"": 1}",EVID-ASSIST-C4
|
| 133 |
+
assistant_owner_clarification_4_v1,EXT-B,new_natural64,content,0,64,0,64,,,,,,EVID-ASSIST-C4
|
| 134 |
+
assistant_owner_clarification_4_v1,EXT-B,new_natural64,joint,0,64,0,64,,,,,,EVID-ASSIST-C4
|
| 135 |
+
assistant_owner_clarification_4_v1,EXT-B,new_natural64,format,3,15,0,18,,,,,,EVID-ASSIST-C4
|
| 136 |
+
assistant_owner_clarification_4_v1,EXT-B,new_python32,content,2,30,0,32,,,,,,EVID-ASSIST-C4
|
| 137 |
+
assistant_owner_clarification_4_v1,EXT-B,new_python32,joint,2,30,0,32,,,,,,EVID-ASSIST-C4
|
| 138 |
+
assistant_owner_clarification_4_v1,EXT-B,new_python32,format,0,0,0,0,,,,,,EVID-ASSIST-C4
|
| 139 |
+
assistant_owner_clarification_4_v1,EXT-B,new_python32,interface,31,1,0,32,,,,,,EVID-ASSIST-C4
|
| 140 |
+
assistant_owner_clarification_4_v1,EXT-B,new_python32,historical_finite_categories,,,,,,,,,"{""fail"": 21, ""fail_or_unextractable"": 2, ""pass"": 7, ""tool_unknown"": 2}",EVID-ASSIST-C4
|
| 141 |
+
assistant_owner_clarification_4_v1,EXT-B,practical27_joint,joint_group_bounds,,,,,8/81,1/9,9.876543,11.111111,,EVID-ASSIST-C4
|
| 142 |
+
assistant_owner_clarification_4_v1,EXT-B,new_natural64,category:R1DEV-CONTEXT,0,16,0,16,,,,,,EVID-ASSIST-C4
|
| 143 |
+
assistant_owner_clarification_4_v1,EXT-B,new_natural64,category:R1DEV-REWRITE,0,16,0,16,,,,,,EVID-ASSIST-C4
|
| 144 |
+
assistant_owner_clarification_4_v1,EXT-B,new_natural64,category:R1DEV-FOLLOWUP,0,16,0,16,,,,,,EVID-ASSIST-C4
|
| 145 |
+
assistant_owner_clarification_4_v1,EXT-B,new_natural64,category:R1DEV-CONDITIONAL,0,16,0,16,,,,,,EVID-ASSIST-C4
|
tables/PRETRAIN_SOURCE_MIXTURE.csv
ADDED
|
@@ -0,0 +1,15 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
stage,source_display_name,stage_i_node_key,upstream_dataset,upstream_revision,upstream_config,recorded_licence_at_pinned_revision,target_serialized_tokens,selected_serialized_tokens,overshoot_tokens,selected_documents,share_of_stage_selected_pct,transport_repository_recorded,transport_revision_recorded
|
| 2 |
+
stage_a,FineWeb-Edu (dedup),fineweb_edu_dedup,HuggingFaceTB/smollm-corpus,3ba9d605774198c5868892d7a8deda78031a781f,fineweb-edu-dedup,odc-by-1.0,7110526316,7110526955,639,7350945,71.1052,HuggingFaceTB/smollm-corpus,3ba9d605774198c5868892d7a8deda78031a781f
|
| 3 |
+
stage_a,DCLM-Edu,dclm_edu,HuggingFaceTB/dclm-edu,dbad8ad71224482740cd9c9d353591adbf62fe04,default,cc-by-4.0,2031578947,2031579037,90,1593857,20.3158,HuggingFaceTB/dclm-edu,dbad8ad71224482740cd9c9d353591adbf62fe04
|
| 4 |
+
stage_a,Wikipedia (FineWiki EN),finewiki_en,HuggingFaceFW/finewiki,8bd13e72e6a002407649b3e898535f42ceb1aeb9,en,cc-by-sa-4.0,507894737,507896470,1733,557285,5.0790,HuggingFaceFW/finewiki,8bd13e72e6a002407649b3e898535f42ceb1aeb9
|
| 5 |
+
stage_a,Python-Edu,python_gate_c_full,common-pile/stackv2_edu_filtered,c354dbe88469a1153e97c6a63ac50591849654de,default,per-record metadata.license (Software Heritage permissive subset),350000000,350000772,772,487237,3.5000,NOT_RECORDED,NOT_RECORDED
|
| 6 |
+
stage_b,FineWeb-Edu (dedup),fineweb_edu_dedup,HuggingFaceTB/smollm-corpus,3ba9d605774198c5868892d7a8deda78031a781f,fineweb-edu-dedup,odc-by-1.0,1203125000,1203125470,470,1199264,40.1041,HuggingFaceTB/smollm-corpus,3ba9d605774198c5868892d7a8deda78031a781f
|
| 7 |
+
stage_b,DCLM-Edu,dclm_edu,HuggingFaceTB/dclm-edu,dbad8ad71224482740cd9c9d353591adbf62fe04,default,cc-by-4.0,687500000,687500443,443,538590,22.9166,HuggingFaceTB/dclm-edu,dbad8ad71224482740cd9c9d353591adbf62fe04
|
| 8 |
+
stage_b,structured_tutorial (Cosmopedia v2 + FinePhrase tutorial),structured_tutorial,HuggingFaceTB/smollm-corpus + HuggingFaceFW/finephrase,3ba9d605774198c5868892d7a8deda78031a781f + 78cf4a5ed0099214979c094c963e699c19163838,cosmopedia-v2 + tutorial,odc-by-1.0 (both),343750000,343750175,175,530450,11.4583,HuggingFaceTB/smollm-corpus + HuggingFaceFW/finephrase,3ba9d605774198c5868892d7a8deda78031a781f + 78cf4a5ed0099214979c094c963e699c19163838
|
| 9 |
+
stage_b,Python-Edu,python_gate_c_full,common-pile/stackv2_edu_filtered,c354dbe88469a1153e97c6a63ac50591849654de,default,per-record metadata.license (Software Heritage permissive subset),250000000,250000383,383,347145,8.3333,NOT_RECORDED,NOT_RECORDED
|
| 10 |
+
stage_b,Wikipedia (FineWiki EN),finewiki_en,HuggingFaceFW/finewiki,8bd13e72e6a002407649b3e898535f42ceb1aeb9,en,cc-by-sa-4.0,171875000,171877052,2052,189465,5.7292,HuggingFaceFW/finewiki,8bd13e72e6a002407649b3e898535f42ceb1aeb9
|
| 11 |
+
stage_b,PES2O,pes2o,allenai/dolmino-mix-1124,a319f19eef1e257417b11ea8c30da266ae175557,pes2o,odc-by-1.0,171875000,171875364,364,625468,5.7292,allenai/dolmino-mix-1124,c58ab4b6ff990115e1ff3121953754ee2bc29501
|
| 12 |
+
stage_b,StackExchange,stackexchange,allenai/dolmino-mix-1124,a319f19eef1e257417b11ea8c30da266ae175557,stackexchange,cc-by-sa,171875000,171875353,353,336025,5.7292,allenai/dolmino-mix-1124,c58ab4b6ff990115e1ff3121953754ee2bc29501
|
| 13 |
+
stage_a,TOTAL,,,,,,10000000000,10000003234,,9989324,100.0000,,
|
| 14 |
+
stage_b,TOTAL,,,,,,3000000000,3000004240,,3766407,100.0000,,
|
| 15 |
+
both,GRAND TOTAL,,,,,,13000000000,13000007474,,13755731,,,
|
tables/PUBLIC_BENCHMARK_RESULTS.csv
ADDED
|
@@ -0,0 +1,7 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
result_version,protocol_version,model,task,split,dataset,dataset_revision,documents,expected_documents,candidate_sequences,acc_correct,acc,acc_pct,acc_norm_correct,acc_norm,acc_norm_pct,status,historical_or_author_reported,runner,evidence_id
|
| 2 |
+
FP32_V2,FP32_V2,PetitGPT-alpha075,arc_easy,test,allenai/ai2_arc,210d026faf9955653af8916fad021475a3f00453,2376,2376,9501,1372,0.5774410774410774,57.74,1244,0.5235690235690236,52.36,COMPLETE,null,native_protocol_compatible_evaluator_NOT_installed_lm_eval,EVID-BENCH-01
|
| 3 |
+
FP32_V2,FP32_V2,PetitGPT-alpha075,piqa,validation,baber/piqa,142f6d7367fd9877f0fb3b5734ea6a545f54cdd1,1838,1838,3676,1167,0.6349292709466812,63.49,1145,0.6229597388465724,62.30,COMPLETE,null,native_protocol_compatible_evaluator_NOT_installed_lm_eval,EVID-BENCH-01
|
| 4 |
+
FP32_V2,FP32_V2,SmolLM-135M-Instruct,arc_easy,test,allenai/ai2_arc,210d026faf9955653af8916fad021475a3f00453,2376,2376,9501,1170,0.49242424242424243,49.24,1033,0.43476430976430974,43.48,COMPLETE,null,native_protocol_compatible_evaluator_NOT_installed_lm_eval,EVID-BENCH-01
|
| 5 |
+
FP32_V2,FP32_V2,SmolLM-135M-Instruct,piqa,validation,baber/piqa,142f6d7367fd9877f0fb3b5734ea6a545f54cdd1,1838,1838,3676,1233,0.6708378672470077,67.08,1236,0.6724700761697497,67.25,COMPLETE,null,native_protocol_compatible_evaluator_NOT_installed_lm_eval,EVID-BENCH-01
|
| 6 |
+
FP32_V2,FP32_V2,SmolLM2-135M-Instruct,arc_easy,test,allenai/ai2_arc,210d026faf9955653af8916fad021475a3f00453,2376,2376,9501,1283,0.539983164983165,54.00,1160,0.4882154882154882,48.82,COMPLETE,null,native_protocol_compatible_evaluator_NOT_installed_lm_eval,EVID-BENCH-01
|
| 7 |
+
FP32_V2,FP32_V2,SmolLM2-135M-Instruct,piqa,validation,baber/piqa,142f6d7367fd9877f0fb3b5734ea6a545f54cdd1,1838,1838,3676,1226,0.6670293797606094,66.70,1227,0.6675734494015234,66.76,COMPLETE,null,native_protocol_compatible_evaluator_NOT_installed_lm_eval,EVID-BENCH-01
|
tokenizer.json
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|