Instructions to use EldanRing/Winnow-E4B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use EldanRing/Winnow-E4B with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf EldanRing/Winnow-E4B:BF16 # Run inference directly in the terminal: llama cli -hf EldanRing/Winnow-E4B:BF16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf EldanRing/Winnow-E4B:BF16 # Run inference directly in the terminal: llama cli -hf EldanRing/Winnow-E4B:BF16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf EldanRing/Winnow-E4B:BF16 # Run inference directly in the terminal: ./llama-cli -hf EldanRing/Winnow-E4B:BF16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf EldanRing/Winnow-E4B:BF16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf EldanRing/Winnow-E4B:BF16
Use Docker
docker model run hf.co/EldanRing/Winnow-E4B:BF16
- LM Studio
- Jan
- vLLM
How to use EldanRing/Winnow-E4B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "EldanRing/Winnow-E4B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "EldanRing/Winnow-E4B", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/EldanRing/Winnow-E4B:BF16
- Ollama
How to use EldanRing/Winnow-E4B with Ollama:
ollama run hf.co/EldanRing/Winnow-E4B:BF16
- Unsloth Desktop
- Docker Model Runner
How to use EldanRing/Winnow-E4B with Docker Model Runner:
docker model run hf.co/EldanRing/Winnow-E4B:BF16
- Lemonade
How to use EldanRing/Winnow-E4B with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull EldanRing/Winnow-E4B:BF16
Run and chat with the model
lemonade run user.Winnow-E4B-BF16
List all available models
lemonade list
- Atomic Chat
Add verified NVFP4 and matching MTP assistant assets
Browse files- .gitattributes +1 -0
- SHA256SUMS +1 -0
- docs/assistants/LICENSE-APACHE-2.0.txt +202 -0
- docs/assistants/NOTICE.txt +2 -0
- docs/assistants/README.md +7 -0
- docs/assistants/UPSTREAM-README.md +589 -0
- docs/assistants/provenance.json +54 -0
- gguf/Gemma-4-E4B-IT-Assistant-BF16.gguf +3 -0
- release-manifest.json +25 -1
.gitattributes
CHANGED
|
@@ -39,3 +39,4 @@ gguf/mmproj-Winnow-E4B.gguf filter=lfs diff=lfs merge=lfs -text
|
|
| 39 |
assets/e4b-profiles.png filter=lfs diff=lfs merge=lfs -text
|
| 40 |
assets/e4b-quality.png filter=lfs diff=lfs merge=lfs -text
|
| 41 |
assets/e4b-speed-memory.png filter=lfs diff=lfs merge=lfs -text
|
|
|
|
|
|
| 39 |
assets/e4b-profiles.png filter=lfs diff=lfs merge=lfs -text
|
| 40 |
assets/e4b-quality.png filter=lfs diff=lfs merge=lfs -text
|
| 41 |
assets/e4b-speed-memory.png filter=lfs diff=lfs merge=lfs -text
|
| 42 |
+
gguf/Gemma-4-E4B-IT-Assistant-BF16.gguf filter=lfs diff=lfs merge=lfs -text
|
SHA256SUMS
CHANGED
|
@@ -1,3 +1,4 @@
|
|
| 1 |
840e3f50e5a9c218727f44e121d1b37cc9e2c3b318c8eb422ba6ef2e27b618a2 gguf/Winnow-E4B-Q8_0.gguf
|
| 2 |
53c3e504c061f6a7ce84cb6240d427ac6c630850a2cf93f3c2ee980c8136e7e7 gguf/Winnow-E4B-BF16.gguf
|
| 3 |
ddf46c21d7078e95338cfc22306b19b276a29a5ad089023449dd54d4b6170a51 gguf/mmproj-Winnow-E4B.gguf
|
|
|
|
|
|
| 1 |
840e3f50e5a9c218727f44e121d1b37cc9e2c3b318c8eb422ba6ef2e27b618a2 gguf/Winnow-E4B-Q8_0.gguf
|
| 2 |
53c3e504c061f6a7ce84cb6240d427ac6c630850a2cf93f3c2ee980c8136e7e7 gguf/Winnow-E4B-BF16.gguf
|
| 3 |
ddf46c21d7078e95338cfc22306b19b276a29a5ad089023449dd54d4b6170a51 gguf/mmproj-Winnow-E4B.gguf
|
| 4 |
+
4e3c9d335b248ced9bd3e0584547efca30dbf6cfd0debb810613f811a575162a gguf/Gemma-4-E4B-IT-Assistant-BF16.gguf
|
docs/assistants/LICENSE-APACHE-2.0.txt
ADDED
|
@@ -0,0 +1,202 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
|
| 2 |
+
Apache License
|
| 3 |
+
Version 2.0, January 2004
|
| 4 |
+
http://www.apache.org/licenses/
|
| 5 |
+
|
| 6 |
+
TERMS AND CONDITIONS FOR USE, REPRODUCTION, AND DISTRIBUTION
|
| 7 |
+
|
| 8 |
+
1. Definitions.
|
| 9 |
+
|
| 10 |
+
"License" shall mean the terms and conditions for use, reproduction,
|
| 11 |
+
and distribution as defined by Sections 1 through 9 of this document.
|
| 12 |
+
|
| 13 |
+
"Licensor" shall mean the copyright owner or entity authorized by
|
| 14 |
+
the copyright owner that is granting the License.
|
| 15 |
+
|
| 16 |
+
"Legal Entity" shall mean the union of the acting entity and all
|
| 17 |
+
other entities that control, are controlled by, or are under common
|
| 18 |
+
control with that entity. For the purposes of this definition,
|
| 19 |
+
"control" means (i) the power, direct or indirect, to cause the
|
| 20 |
+
direction or management of such entity, whether by contract or
|
| 21 |
+
otherwise, or (ii) ownership of fifty percent (50%) or more of the
|
| 22 |
+
outstanding shares, or (iii) beneficial ownership of such entity.
|
| 23 |
+
|
| 24 |
+
"You" (or "Your") shall mean an individual or Legal Entity
|
| 25 |
+
exercising permissions granted by this License.
|
| 26 |
+
|
| 27 |
+
"Source" form shall mean the preferred form for making modifications,
|
| 28 |
+
including but not limited to software source code, documentation
|
| 29 |
+
source, and configuration files.
|
| 30 |
+
|
| 31 |
+
"Object" form shall mean any form resulting from mechanical
|
| 32 |
+
transformation or translation of a Source form, including but
|
| 33 |
+
not limited to compiled object code, generated documentation,
|
| 34 |
+
and conversions to other media types.
|
| 35 |
+
|
| 36 |
+
"Work" shall mean the work of authorship, whether in Source or
|
| 37 |
+
Object form, made available under the License, as indicated by a
|
| 38 |
+
copyright notice that is included in or attached to the work
|
| 39 |
+
(an example is provided in the Appendix below).
|
| 40 |
+
|
| 41 |
+
"Derivative Works" shall mean any work, whether in Source or Object
|
| 42 |
+
form, that is based on (or derived from) the Work and for which the
|
| 43 |
+
editorial revisions, annotations, elaborations, or other modifications
|
| 44 |
+
represent, as a whole, an original work of authorship. For the purposes
|
| 45 |
+
of this License, Derivative Works shall not include works that remain
|
| 46 |
+
separable from, or merely link (or bind by name) to the interfaces of,
|
| 47 |
+
the Work and Derivative Works thereof.
|
| 48 |
+
|
| 49 |
+
"Contribution" shall mean any work of authorship, including
|
| 50 |
+
the original version of the Work and any modifications or additions
|
| 51 |
+
to that Work or Derivative Works thereof, that is intentionally
|
| 52 |
+
submitted to Licensor for inclusion in the Work by the copyright owner
|
| 53 |
+
or by an individual or Legal Entity authorized to submit on behalf of
|
| 54 |
+
the copyright owner. For the purposes of this definition, "submitted"
|
| 55 |
+
means any form of electronic, verbal, or written communication sent
|
| 56 |
+
to the Licensor or its representatives, including but not limited to
|
| 57 |
+
communication on electronic mailing lists, source code control systems,
|
| 58 |
+
and issue tracking systems that are managed by, or on behalf of, the
|
| 59 |
+
Licensor for the purpose of discussing and improving the Work, but
|
| 60 |
+
excluding communication that is conspicuously marked or otherwise
|
| 61 |
+
designated in writing by the copyright owner as "Not a Contribution."
|
| 62 |
+
|
| 63 |
+
"Contributor" shall mean Licensor and any individual or Legal Entity
|
| 64 |
+
on behalf of whom a Contribution has been received by Licensor and
|
| 65 |
+
subsequently incorporated within the Work.
|
| 66 |
+
|
| 67 |
+
2. Grant of Copyright License. Subject to the terms and conditions of
|
| 68 |
+
this License, each Contributor hereby grants to You a perpetual,
|
| 69 |
+
worldwide, non-exclusive, no-charge, royalty-free, irrevocable
|
| 70 |
+
copyright license to reproduce, prepare Derivative Works of,
|
| 71 |
+
publicly display, publicly perform, sublicense, and distribute the
|
| 72 |
+
Work and such Derivative Works in Source or Object form.
|
| 73 |
+
|
| 74 |
+
3. Grant of Patent License. Subject to the terms and conditions of
|
| 75 |
+
this License, each Contributor hereby grants to You a perpetual,
|
| 76 |
+
worldwide, non-exclusive, no-charge, royalty-free, irrevocable
|
| 77 |
+
(except as stated in this section) patent license to make, have made,
|
| 78 |
+
use, offer to sell, sell, import, and otherwise transfer the Work,
|
| 79 |
+
where such license applies only to those patent claims licensable
|
| 80 |
+
by such Contributor that are necessarily infringed by their
|
| 81 |
+
Contribution(s) alone or by combination of their Contribution(s)
|
| 82 |
+
with the Work to which such Contribution(s) was submitted. If You
|
| 83 |
+
institute patent litigation against any entity (including a
|
| 84 |
+
cross-claim or counterclaim in a lawsuit) alleging that the Work
|
| 85 |
+
or a Contribution incorporated within the Work constitutes direct
|
| 86 |
+
or contributory patent infringement, then any patent licenses
|
| 87 |
+
granted to You under this License for that Work shall terminate
|
| 88 |
+
as of the date such litigation is filed.
|
| 89 |
+
|
| 90 |
+
4. Redistribution. You may reproduce and distribute copies of the
|
| 91 |
+
Work or Derivative Works thereof in any medium, with or without
|
| 92 |
+
modifications, and in Source or Object form, provided that You
|
| 93 |
+
meet the following conditions:
|
| 94 |
+
|
| 95 |
+
(a) You must give any other recipients of the Work or
|
| 96 |
+
Derivative Works a copy of this License; and
|
| 97 |
+
|
| 98 |
+
(b) You must cause any modified files to carry prominent notices
|
| 99 |
+
stating that You changed the files; and
|
| 100 |
+
|
| 101 |
+
(c) You must retain, in the Source form of any Derivative Works
|
| 102 |
+
that You distribute, all copyright, patent, trademark, and
|
| 103 |
+
attribution notices from the Source form of the Work,
|
| 104 |
+
excluding those notices that do not pertain to any part of
|
| 105 |
+
the Derivative Works; and
|
| 106 |
+
|
| 107 |
+
(d) If the Work includes a "NOTICE" text file as part of its
|
| 108 |
+
distribution, then any Derivative Works that You distribute must
|
| 109 |
+
include a readable copy of the attribution notices contained
|
| 110 |
+
within such NOTICE file, excluding those notices that do not
|
| 111 |
+
pertain to any part of the Derivative Works, in at least one
|
| 112 |
+
of the following places: within a NOTICE text file distributed
|
| 113 |
+
as part of the Derivative Works; within the Source form or
|
| 114 |
+
documentation, if provided along with the Derivative Works; or,
|
| 115 |
+
within a display generated by the Derivative Works, if and
|
| 116 |
+
wherever such third-party notices normally appear. The contents
|
| 117 |
+
of the NOTICE file are for informational purposes only and
|
| 118 |
+
do not modify the License. You may add Your own attribution
|
| 119 |
+
notices within Derivative Works that You distribute, alongside
|
| 120 |
+
or as an addendum to the NOTICE text from the Work, provided
|
| 121 |
+
that such additional attribution notices cannot be construed
|
| 122 |
+
as modifying the License.
|
| 123 |
+
|
| 124 |
+
You may add Your own copyright statement to Your modifications and
|
| 125 |
+
may provide additional or different license terms and conditions
|
| 126 |
+
for use, reproduction, or distribution of Your modifications, or
|
| 127 |
+
for any such Derivative Works as a whole, provided Your use,
|
| 128 |
+
reproduction, and distribution of the Work otherwise complies with
|
| 129 |
+
the conditions stated in this License.
|
| 130 |
+
|
| 131 |
+
5. Submission of Contributions. Unless You explicitly state otherwise,
|
| 132 |
+
any Contribution intentionally submitted for inclusion in the Work
|
| 133 |
+
by You to the Licensor shall be under the terms and conditions of
|
| 134 |
+
this License, without any additional terms or conditions.
|
| 135 |
+
Notwithstanding the above, nothing herein shall supersede or modify
|
| 136 |
+
the terms of any separate license agreement you may have executed
|
| 137 |
+
with Licensor regarding such Contributions.
|
| 138 |
+
|
| 139 |
+
6. Trademarks. This License does not grant permission to use the trade
|
| 140 |
+
names, trademarks, service marks, or product names of the Licensor,
|
| 141 |
+
except as required for reasonable and customary use in describing the
|
| 142 |
+
origin of the Work and reproducing the content of the NOTICE file.
|
| 143 |
+
|
| 144 |
+
7. Disclaimer of Warranty. Unless required by applicable law or
|
| 145 |
+
agreed to in writing, Licensor provides the Work (and each
|
| 146 |
+
Contributor provides its Contributions) on an "AS IS" BASIS,
|
| 147 |
+
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or
|
| 148 |
+
implied, including, without limitation, any warranties or conditions
|
| 149 |
+
of TITLE, NON-INFRINGEMENT, MERCHANTABILITY, or FITNESS FOR A
|
| 150 |
+
PARTICULAR PURPOSE. You are solely responsible for determining the
|
| 151 |
+
appropriateness of using or redistributing the Work and assume any
|
| 152 |
+
risks associated with Your exercise of permissions under this License.
|
| 153 |
+
|
| 154 |
+
8. Limitation of Liability. In no event and under no legal theory,
|
| 155 |
+
whether in tort (including negligence), contract, or otherwise,
|
| 156 |
+
unless required by applicable law (such as deliberate and grossly
|
| 157 |
+
negligent acts) or agreed to in writing, shall any Contributor be
|
| 158 |
+
liable to You for damages, including any direct, indirect, special,
|
| 159 |
+
incidental, or consequential damages of any character arising as a
|
| 160 |
+
result of this License or out of the use or inability to use the
|
| 161 |
+
Work (including but not limited to damages for loss of goodwill,
|
| 162 |
+
work stoppage, computer failure or malfunction, or any and all
|
| 163 |
+
other commercial damages or losses), even if such Contributor
|
| 164 |
+
has been advised of the possibility of such damages.
|
| 165 |
+
|
| 166 |
+
9. Accepting Warranty or Additional Liability. While redistributing
|
| 167 |
+
the Work or Derivative Works thereof, You may choose to offer,
|
| 168 |
+
and charge a fee for, acceptance of support, warranty, indemnity,
|
| 169 |
+
or other liability obligations and/or rights consistent with this
|
| 170 |
+
License. However, in accepting such obligations, You may act only
|
| 171 |
+
on Your own behalf and on Your sole responsibility, not on behalf
|
| 172 |
+
of any other Contributor, and only if You agree to indemnify,
|
| 173 |
+
defend, and hold each Contributor harmless for any liability
|
| 174 |
+
incurred by, or claims asserted against, such Contributor by reason
|
| 175 |
+
of your accepting any such warranty or additional liability.
|
| 176 |
+
|
| 177 |
+
END OF TERMS AND CONDITIONS
|
| 178 |
+
|
| 179 |
+
APPENDIX: How to apply the Apache License to your work.
|
| 180 |
+
|
| 181 |
+
To apply the Apache License to your work, attach the following
|
| 182 |
+
boilerplate notice, with the fields enclosed by brackets "[]"
|
| 183 |
+
replaced with your own identifying information. (Don't include
|
| 184 |
+
the brackets!) The text should be enclosed in the appropriate
|
| 185 |
+
comment syntax for the file format. We also recommend that a
|
| 186 |
+
file or class name and description of purpose be included on the
|
| 187 |
+
same "printed page" as the copyright notice for easier
|
| 188 |
+
identification within third-party archives.
|
| 189 |
+
|
| 190 |
+
Copyright [yyyy] [name of copyright owner]
|
| 191 |
+
|
| 192 |
+
Licensed under the Apache License, Version 2.0 (the "License");
|
| 193 |
+
you may not use this file except in compliance with the License.
|
| 194 |
+
You may obtain a copy of the License at
|
| 195 |
+
|
| 196 |
+
http://www.apache.org/licenses/LICENSE-2.0
|
| 197 |
+
|
| 198 |
+
Unless required by applicable law or agreed to in writing, software
|
| 199 |
+
distributed under the License is distributed on an "AS IS" BASIS,
|
| 200 |
+
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
| 201 |
+
See the License for the specific language governing permissions and
|
| 202 |
+
limitations under the License.
|
docs/assistants/NOTICE.txt
ADDED
|
@@ -0,0 +1,2 @@
|
|
|
|
|
|
|
|
|
|
| 1 |
+
Official Gemma 4 MTP assistant models by Google DeepMind, Apache License 2.0.
|
| 2 |
+
CPU-converted BF16 GGUF files preserve the exact pinned upstream identities recorded in manifest.json. Conversion changes serialization and performs no training. Upstream cards and sanitized provenance are included. Google does not endorse Winnow. Upstream pinned repository listings contained no separate NOTICE file.
|
docs/assistants/README.md
ADDED
|
@@ -0,0 +1,7 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# E4B MTP assistant
|
| 2 |
+
|
| 3 |
+
`Gemma-4-E4B-IT-Assistant-BF16.gguf` is a BF16 GGUF conversion of Google's official [google/gemma-4-E4B-it-assistant](https://huggingface.co/google/gemma-4-E4B-it-assistant) at revision `8d0031ea8c2109e2b1e86bb9368a4539b537f80a`. Conversion changes serialization and performs no training.
|
| 4 |
+
|
| 5 |
+
File size: 171766688 bytes. SHA256: `4e3c9d335b248ced9bd3e0584547efca30dbf6cfd0debb810613f811a575162a`. The matching Winnow launcher verifies the complete file before loading; other assistant files are not silently substituted.
|
| 6 |
+
|
| 7 |
+
The original [upstream model card](https://huggingface.co/google/gemma-4-E4B-it-assistant/blob/8d0031ea8c2109e2b1e86bb9368a4539b537f80a/README.md) and Apache-2.0 license apply. See LICENSE-APACHE-2.0.txt, NOTICE.txt and provenance.json in this directory for attribution and exact conversion provenance.
|
docs/assistants/UPSTREAM-README.md
ADDED
|
@@ -0,0 +1,589 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
library_name: transformers
|
| 3 |
+
license: apache-2.0
|
| 4 |
+
license_link: https://ai.google.dev/gemma/docs/gemma_4_license
|
| 5 |
+
pipeline_tag: any-to-any
|
| 6 |
+
---
|
| 7 |
+
|
| 8 |
+
<div align="center">
|
| 9 |
+
<img src=https://ai.google.dev/gemma/images/gemma4_banner.png>
|
| 10 |
+
</div>
|
| 11 |
+
|
| 12 |
+
|
| 13 |
+
<p align="center">
|
| 14 |
+
<a href="https://huggingface.co/collections/google/gemma-4" target="_blank">Hugging Face</a> |
|
| 15 |
+
<a href="https://github.com/google-gemma" target="_blank">GitHub</a> |
|
| 16 |
+
<a href="https://ai.google.dev/gemma/docs/mtp/mtp" target="_blank">MTP Documentation</a> |
|
| 17 |
+
<a href="https://arxiv.org/abs/2607.02770" target="_blank">Technical Report</a>
|
| 18 |
+
<br>
|
| 19 |
+
<b>License</b>: <a href="https://ai.google.dev/gemma/docs/gemma_4_license" target="_blank">Apache 2.0</a> | <b>Authors</b>: <a href="https://deepmind.google/models/gemma/" target="_blank">Google DeepMind</a>
|
| 20 |
+
</p>
|
| 21 |
+
|
| 22 |
+
> [!Note]
|
| 23 |
+
> This model card is for the Multi-Token Prediction (MTP) drafters for the Gemma 4 models. MTP is implemented by extending the base model with a smaller, faster draft model. When used in a Speculative Decoding pipeline, the draft model predicts several tokens ahead, which the target model then verifies in parallel. This results in significant decoding speedups (up to 3x) while guaranteeing the exact same quality as standard generation, making these checkpoints perfect for low-latency and on-device applications.
|
| 24 |
+
|
| 25 |
+
Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input (with audio supported on E2B, E4B, and 12B) and generating text output. This release includes open-weights models in both pre-trained and instruction-tuned variants. Gemma 4 features a context window of up to 256K tokens and maintains multilingual support in over 140 languages.
|
| 26 |
+
|
| 27 |
+
Featuring both Dense and Mixture-of-Experts (MoE) architectures, Gemma 4 is well-suited for tasks like text generation, coding, and reasoning. The models are available in five distinct sizes: **E2B**, **E4B**, **12B**, **26B A4B**, and **31B**. Their diverse sizes make them deployable in environments ranging from high-end phones to laptops and servers, democratizing access to state-of-the-art AI.
|
| 28 |
+
|
| 29 |
+
Gemma 4 introduces key **capability and architectural advancements**:
|
| 30 |
+
|
| 31 |
+
* **Reasoning** – All models in the family are designed as highly capable reasoners, with configurable thinking modes.
|
| 32 |
+
|
| 33 |
+
* **Extended Multimodalities** – Processes Text, Image with variable aspect ratio and resolution support (all models), Video, and Audio (featured natively on the E2B, E4B, and 12B models).
|
| 34 |
+
|
| 35 |
+
* **Diverse & Efficient Architectures** – Offers Dense and Mixture-of-Experts (MoE) variants of different sizes for scalable deployment.
|
| 36 |
+
|
| 37 |
+
* **Optimized for On-Device** – Smaller models are specifically designed for efficient local execution on laptops and mobile devices.
|
| 38 |
+
|
| 39 |
+
* **Increased Context Window** – The small models feature a 128K context window, while the medium models support 256K.
|
| 40 |
+
|
| 41 |
+
* **Enhanced Coding & Agentic Capabilities** – Achieves notable improvements in coding benchmarks alongside native function-calling support, powering highly capable autonomous agents.
|
| 42 |
+
|
| 43 |
+
* **Native System Prompt Support** – Gemma 4 introduces native support for the `system` role, enabling more structured and controllable conversations.
|
| 44 |
+
|
| 45 |
+
## **Models Overview**
|
| 46 |
+
|
| 47 |
+
Gemma 4 models are designed to deliver frontier-level performance at each size, targeting deployment scenarios from mobile and edge devices (E2B, E4B) to consumer GPUs and workstations (12B, 26B A4B, 31B). They are well-suited for reasoning, agentic workflows, coding, and multimodal understanding.
|
| 48 |
+
|
| 49 |
+
The models employ a hybrid attention mechanism that interleaves local sliding window attention with full global attention, ensuring the final layer is always global. This hybrid design delivers the processing speed and low memory footprint of a lightweight model without sacrificing the deep awareness required for complex, long-context tasks. To optimize memory for long contexts, global layers feature unified Keys and Values, and apply Proportional RoPE (p-RoPE).
|
| 50 |
+
|
| 51 |
+
### Dense Models
|
| 52 |
+
|
| 53 |
+
| Property | E2B | E4B | 12B Unified | 31B Dense |
|
| 54 |
+
| :---- | :---- | :---- | :---- | :---- |
|
| 55 |
+
| **Total Parameters** | 2.3B effective <br> (5.1B with embeddings) | 4.5B effective <br> (8B with embeddings) | 11.95B | 30.7B |
|
| 56 |
+
| **Layers** | 35 | 42 | 48 | 60 |
|
| 57 |
+
| **Sliding Window** | 512 tokens | 512 tokens | 1024 tokens | 1024 tokens |
|
| 58 |
+
| **Context Length** | 128K tokens | 128K tokens | 256K tokens | 256K tokens |
|
| 59 |
+
| **Vocabulary Size** | 262K | 262K | 262K | 262K |
|
| 60 |
+
| **Supported Modalities** | Text, Image, Audio | Text, Image, Audio | Text, Image, Audio | Text, Image |
|
| 61 |
+
| **Vision Encoder Parameters** | *~150M* | *~150M* | - | *~550M* |
|
| 62 |
+
| **Audio Encoder Parameters** | *~300M* | *~300M* | - | No Audio |
|
| 63 |
+
|
| 64 |
+
The "E" in E2B and E4B stands for "effective" parameters. The smaller models incorporate Per-Layer Embeddings (PLE) to maximize parameter efficiency in on-device deployments. Rather than adding more layers or parameters to the model, PLE gives each decoder layer its own small embedding for every token. These embedding tables are large but are only used for quick lookups, which is why the effective parameter count is much smaller than the total.
|
| 65 |
+
|
| 66 |
+
The "Unified" in Gemma 4 12B Unified refers to its encoder-free architecture. Other Gemma 4 models use dedicated encoders to process multimodal data before passing it to the LLM. Gemma 4 12B eliminates these encoders entirely, projecting raw image patches and audio waveforms directly into the LLM's embedding space through lightweight linear layers. This unified approach means all modalities flow straight into a single decoder-only transformer, reducing multimodal latency and allowing the entire model to be fine-tuned in one pass.
|
| 67 |
+
|
| 68 |
+
### Mixture-of-Experts (MoE) Model
|
| 69 |
+
|
| 70 |
+
| Property | 26B A4B MoE |
|
| 71 |
+
| :---- | :---- |
|
| 72 |
+
| **Total Parameters** | 25.2B |
|
| 73 |
+
| **Active Parameters** | 3.8B |
|
| 74 |
+
| **Layers** | 30 |
|
| 75 |
+
| **Sliding Window** | 1024 tokens |
|
| 76 |
+
| **Context Length** | 256K tokens |
|
| 77 |
+
| **Vocabulary Size** | 262K |
|
| 78 |
+
| **Expert Count** | 8 active / 128 total and 1 shared |
|
| 79 |
+
| **Supported Modalities** | Text, Image |
|
| 80 |
+
| **Vision Encoder Parameters** | *~550M* |
|
| 81 |
+
|
| 82 |
+
The "A" in 26B A4B stands for "active parameters" in contrast to the total number of parameters the model contains. By only activating a 4B subset of parameters during inference, the Mixture-of-Experts model runs much faster than its 26B total might suggest. This makes it an excellent choice for fast inference compared to the dense 31B model since it runs almost as fast as a 4B-parameter model.
|
| 83 |
+
|
| 84 |
+
## **Benchmark Results**
|
| 85 |
+
|
| 86 |
+
These models were evaluated against a large collection of different datasets and metrics to cover different aspects of text generation. Evaluation results marked in the table are for instruction-tuned models.
|
| 87 |
+
|
| 88 |
+
| | Gemma 4 31B | Gemma 4 26B A4B | Gemma 4 12B Unified | Gemma 4 E4B | Gemma 4 E2B | Gemma 3 27B (no think) |
|
| 89 |
+
| :---- | :---- | :---- | :---- | :---- | :---- | :---- |
|
| 90 |
+
| MMLU Pro | 85.2% | 82.6% | 77.2% | 69.4% | 60.0% | 67.6% |
|
| 91 |
+
| AIME 2026 no tools | 89.2% | 88.3% | 77.5% | 42.5% | 37.5% | 20.8% |
|
| 92 |
+
| LiveCodeBench v6 | 80.0% | 77.1% | 72.0% | 52.0% | 44.0% | 29.1% |
|
| 93 |
+
| Codeforces ELO | 2150 | 1718 | 1659 | 940 | 633 | 110 |
|
| 94 |
+
| GPQA Diamond | 84.3% | 82.3% | 78.8% | 58.6% | 43.4% | 42.4% |
|
| 95 |
+
| Tau2 (average over 3) | 76.9% | 68.2% | 69.0% | 42.2% | 24.5% | 16.2% |
|
| 96 |
+
| HLE no tools | 19.5% | 8.7% | 5.2% | - | - | - |
|
| 97 |
+
| HLE with search | 26.5% | 17.2% | - | - | - | - |
|
| 98 |
+
| BigBench Extra Hard | 74.4% | 64.8% | 53.0% | 33.1% | 21.9% | 19.3% |
|
| 99 |
+
| MMMLU | 88.4% | 86.3% | 83.4% | 76.6% | 67.4% | 70.7% |
|
| 100 |
+
| **Vision** | | | | | | |
|
| 101 |
+
| MMMU Pro | 76.9% | 73.8% | 69.1% | 52.6% | 44.2% | 49.7% |
|
| 102 |
+
| OmniDocBench 1.5 (average edit distance, lower is better) | 0.131 | 0.149 | 0.164 | 0.181 | 0.290 | 0.365 |
|
| 103 |
+
| MATH-Vision | 85.6% | 82.4% | 79.7% | 59.5% | 52.4% | 46.0% |
|
| 104 |
+
| MedXPertQA MM | 61.3% | 58.1% | 48.7% | 28.7% | 23.5% | - |
|
| 105 |
+
| **Audio** | | | | | | |
|
| 106 |
+
| CoVoST | - | - | 38.5<sup>*</sup> | 35.54 | 33.47 | - |
|
| 107 |
+
| FLEURS (lower is better) | - | - | 0.069<sup>*</sup> | 0.08 | 0.09 | - |
|
| 108 |
+
| **Long Context** | | | | | | |
|
| 109 |
+
| MRCR v2 8 needle 128k (average) | 66.4% | 44.1% | 43.4% | 25.4% | 19.1% | 13.5% |
|
| 110 |
+
|
| 111 |
+
<sup>*</sup>Excluding Chinese language.
|
| 112 |
+
|
| 113 |
+
## **Core Capabilities**
|
| 114 |
+
|
| 115 |
+
Gemma 4 models handle a broad range of tasks across text, vision, and audio. Key capabilities include:
|
| 116 |
+
|
| 117 |
+
* **Thinking** – Built-in reasoning mode that lets the model think step-by-step before answering.
|
| 118 |
+
* **Long Context** – Context windows of up to 128K tokens (E2B/E4B) and 256K tokens (12B, 26B A4B/31B).
|
| 119 |
+
* **Image Understanding** – Object detection, Document/PDF parsing, screen and UI understanding, chart comprehension, OCR (including multilingual), handwriting recognition, and pointing. Images can be processed at variable aspect ratios and resolutions.
|
| 120 |
+
* **Video Understanding** – Analyze video by processing sequences of frames.
|
| 121 |
+
* **Interleaved Multimodal Input** – Freely mix text and images in any order within a single prompt.
|
| 122 |
+
* **Function Calling** – Native support for structured tool use, enabling agentic workflows.
|
| 123 |
+
* **Coding** – Code generation, completion, and correction.
|
| 124 |
+
* **Multilingual** – Out-of-the-box support for 35+ languages, pre-trained on 140+ languages.
|
| 125 |
+
* **Audio** (E2B, E4B, and 12B only) – Automatic speech recognition (ASR) and speech-to-translated-text translation across multiple languages.
|
| 126 |
+
|
| 127 |
+
|
| 128 |
+
## Getting Started
|
| 129 |
+
|
| 130 |
+
You can use all Gemma 4 models with the latest version of Transformers. To get started, install the necessary dependencies in your environment:
|
| 131 |
+
|
| 132 |
+
`pip install -U transformers torch accelerate`
|
| 133 |
+
|
| 134 |
+
Once you have everything installed, you can proceed to load the target model and the assistant model with the code below:
|
| 135 |
+
|
| 136 |
+
```python
|
| 137 |
+
from transformers import AutoProcessor, AutoModelForCausalLM
|
| 138 |
+
|
| 139 |
+
TARGET_MODEL_ID = "google/gemma-4-E4B-it"
|
| 140 |
+
ASSISTANT_MODEL_ID = "google/gemma-4-E4B-it-assistant"
|
| 141 |
+
|
| 142 |
+
# Target Model
|
| 143 |
+
processor = AutoProcessor.from_pretrained(TARGET_MODEL_ID)
|
| 144 |
+
target_model = AutoModelForCausalLM.from_pretrained(
|
| 145 |
+
TARGET_MODEL_ID,
|
| 146 |
+
dtype="auto",
|
| 147 |
+
device_map="auto",
|
| 148 |
+
|
| 149 |
+
)
|
| 150 |
+
|
| 151 |
+
# Assistant Model (the drafter)
|
| 152 |
+
assistant_model = AutoModelForCausalLM.from_pretrained(
|
| 153 |
+
ASSISTANT_MODEL_ID,
|
| 154 |
+
dtype="auto",
|
| 155 |
+
device_map="auto",
|
| 156 |
+
)
|
| 157 |
+
```
|
| 158 |
+
|
| 159 |
+
Once the models are loaded, you can start generating output:
|
| 160 |
+
|
| 161 |
+
```python
|
| 162 |
+
# Prompt
|
| 163 |
+
messages = [
|
| 164 |
+
{"role": "system", "content": "You are a helpful assistant."},
|
| 165 |
+
{"role": "user", "content": "Write a short joke about saving RAM."},
|
| 166 |
+
]
|
| 167 |
+
|
| 168 |
+
# Process input
|
| 169 |
+
inputs = processor.apply_chat_template(
|
| 170 |
+
messages,
|
| 171 |
+
tokenize=True,
|
| 172 |
+
return_dict=True,
|
| 173 |
+
return_tensors="pt",
|
| 174 |
+
add_generation_prompt=True,
|
| 175 |
+
enable_thinking=False
|
| 176 |
+
).to(model.device)
|
| 177 |
+
input_len = inputs["input_ids"].shape[-1]
|
| 178 |
+
|
| 179 |
+
# Generate output
|
| 180 |
+
outputs = target_model.generate(
|
| 181 |
+
**inputs,
|
| 182 |
+
assistant_model=assistant_model,
|
| 183 |
+
max_new_tokens=256,
|
| 184 |
+
)
|
| 185 |
+
response = processor.decode(outputs[0][input_len:], skip_special_tokens=False)
|
| 186 |
+
|
| 187 |
+
# Parse output
|
| 188 |
+
processor.parse_response(response)
|
| 189 |
+
```
|
| 190 |
+
|
| 191 |
+
To enable reasoning, set `enable_thinking=True` and the `parse_response` function will take care of parsing the thinking output.
|
| 192 |
+
|
| 193 |
+
Below, you will also find snippets for processing audio (E2B, E4B, and 12B only), images, and video alongside text:
|
| 194 |
+
|
| 195 |
+
<details>
|
| 196 |
+
<summary>Code for processing Audio</summary>
|
| 197 |
+
|
| 198 |
+
Make sure to install the following packages:
|
| 199 |
+
|
| 200 |
+
|
| 201 |
+
`pip install -U transformers torch torchvision librosa accelerate`
|
| 202 |
+
|
| 203 |
+
Once you have everything installed, you can proceed to load the target model and the assistant model with the code below:
|
| 204 |
+
|
| 205 |
+
```python
|
| 206 |
+
import torch
|
| 207 |
+
from transformers import AutoProcessor, AutoModelForCausalLM, AutoModelForMultimodalLM
|
| 208 |
+
|
| 209 |
+
TARGET_MODEL_ID = "google/gemma-4-E4B-it"
|
| 210 |
+
ASSISTANT_MODEL_ID = "google/gemma-4-E4B-it-assistant"
|
| 211 |
+
|
| 212 |
+
# Target Model
|
| 213 |
+
processor = AutoProcessor.from_pretrained(TARGET_MODEL_ID)
|
| 214 |
+
target_model = AutoModelForMultimodalLM.from_pretrained(
|
| 215 |
+
TARGET_MODEL_ID,
|
| 216 |
+
torch_dtype=torch.bfloat16,
|
| 217 |
+
device_map="auto",
|
| 218 |
+
|
| 219 |
+
)
|
| 220 |
+
|
| 221 |
+
# Assistant Model (the drafter)
|
| 222 |
+
assistant_model = AutoModelForCausalLM.from_pretrained(
|
| 223 |
+
ASSISTANT_MODEL_ID,
|
| 224 |
+
torch_dtype=torch.bfloat16,
|
| 225 |
+
device_map="auto",
|
| 226 |
+
)
|
| 227 |
+
```
|
| 228 |
+
|
| 229 |
+
Once the model is loaded, you can start generating output by directly referencing the audio URL in the prompt:
|
| 230 |
+
|
| 231 |
+
|
| 232 |
+
```python
|
| 233 |
+
# Prompt - add audio after text
|
| 234 |
+
messages = [
|
| 235 |
+
{
|
| 236 |
+
"role": "user",
|
| 237 |
+
"content": [
|
| 238 |
+
{"type": "text", "text": "Transcribe the following speech segment in its original language. Follow these specific instructions for formatting the answer:\n* Only output the transcription, with no newlines.\n* When transcribing numbers, write the digits, i.e. write 1.7 and not one point seven, and write 3 instead of three."},
|
| 239 |
+
{"type": "audio", "audio": "https://github.com/google-gemma/cookbook/raw/refs/heads/main/apps/sample-data/journal1.wav"},
|
| 240 |
+
]
|
| 241 |
+
}
|
| 242 |
+
]
|
| 243 |
+
|
| 244 |
+
# Process input
|
| 245 |
+
text = processor.apply_chat_template(
|
| 246 |
+
messages,
|
| 247 |
+
tokenize=False,
|
| 248 |
+
add_generation_prompt=True,
|
| 249 |
+
)
|
| 250 |
+
inputs = processor(text=text, return_tensors="pt").to(target_model.device)
|
| 251 |
+
input_len = inputs["input_ids"].shape[-1]
|
| 252 |
+
|
| 253 |
+
# Generate output
|
| 254 |
+
outputs = target_model.generate(
|
| 255 |
+
**inputs,
|
| 256 |
+
assistant_model=assistant_model,
|
| 257 |
+
max_new_tokens=256,
|
| 258 |
+
)
|
| 259 |
+
response = processor.decode(outputs[0][input_len:], skip_special_tokens=False)
|
| 260 |
+
|
| 261 |
+
# Parse output
|
| 262 |
+
processor.parse_response(response)
|
| 263 |
+
```
|
| 264 |
+
|
| 265 |
+
</details>
|
| 266 |
+
|
| 267 |
+
<details>
|
| 268 |
+
<summary>Code for processing Images</summary>
|
| 269 |
+
|
| 270 |
+
Make sure to install the following packages:
|
| 271 |
+
|
| 272 |
+
|
| 273 |
+
`pip install -U transformers torch torchvision accelerate`
|
| 274 |
+
|
| 275 |
+
Once you have everything installed, you can proceed to load the target model and the assistant model with the code below:
|
| 276 |
+
|
| 277 |
+
```python
|
| 278 |
+
import torch
|
| 279 |
+
from transformers import AutoProcessor, AutoModelForCausalLM, AutoModelForMultimodalLM
|
| 280 |
+
|
| 281 |
+
TARGET_MODEL_ID = "google/gemma-4-E4B-it"
|
| 282 |
+
ASSISTANT_MODEL_ID = "google/gemma-4-E4B-it-assistant"
|
| 283 |
+
|
| 284 |
+
# Target Model
|
| 285 |
+
processor = AutoProcessor.from_pretrained(TARGET_MODEL_ID)
|
| 286 |
+
target_model = AutoModelForMultimodalLM.from_pretrained(
|
| 287 |
+
TARGET_MODEL_ID,
|
| 288 |
+
torch_dtype=torch.bfloat16,
|
| 289 |
+
device_map="auto",
|
| 290 |
+
|
| 291 |
+
)
|
| 292 |
+
|
| 293 |
+
# Assistant Model (the drafter)
|
| 294 |
+
assistant_model = AutoModelForCausalLM.from_pretrained(
|
| 295 |
+
ASSISTANT_MODEL_ID,
|
| 296 |
+
torch_dtype=torch.bfloat16,
|
| 297 |
+
device_map="auto",
|
| 298 |
+
)
|
| 299 |
+
```
|
| 300 |
+
|
| 301 |
+
Once the model is loaded, you can start generating output by directly referencing the image URL in the prompt:
|
| 302 |
+
|
| 303 |
+
|
| 304 |
+
```python
|
| 305 |
+
# Prompt - add image before text
|
| 306 |
+
messages = [
|
| 307 |
+
{
|
| 308 |
+
"role": "user", "content": [
|
| 309 |
+
{"type": "image", "url": "https://raw.githubusercontent.com/google-gemma/cookbook/refs/heads/main/apps/sample-data/GoldenGate.png"},
|
| 310 |
+
{"type": "text", "text": "What is shown in this image?"}
|
| 311 |
+
]
|
| 312 |
+
}
|
| 313 |
+
]
|
| 314 |
+
|
| 315 |
+
# Process input
|
| 316 |
+
inputs = processor.apply_chat_template(
|
| 317 |
+
messages,
|
| 318 |
+
tokenize=True,
|
| 319 |
+
return_dict=True,
|
| 320 |
+
return_tensors="pt",
|
| 321 |
+
add_generation_prompt=True,
|
| 322 |
+
).to(target_model.device)
|
| 323 |
+
input_len = inputs["input_ids"].shape[-1]
|
| 324 |
+
|
| 325 |
+
# Generate output
|
| 326 |
+
outputs = target_model.generate(**inputs, max_new_tokens=512)
|
| 327 |
+
response = processor.decode(outputs[0][input_len:], skip_special_tokens=False)
|
| 328 |
+
|
| 329 |
+
# Parse output
|
| 330 |
+
processor.parse_response(response)
|
| 331 |
+
```
|
| 332 |
+
|
| 333 |
+
</details>
|
| 334 |
+
|
| 335 |
+
|
| 336 |
+
<details>
|
| 337 |
+
<summary>Code for processing Videos</summary>
|
| 338 |
+
|
| 339 |
+
Make sure to install the following packages:
|
| 340 |
+
|
| 341 |
+
`pip install -U transformers torch torchvision librosa accelerate`
|
| 342 |
+
|
| 343 |
+
Once you have everything installed, you can proceed to load the target model and the assistant model with the code below:
|
| 344 |
+
|
| 345 |
+
```python
|
| 346 |
+
import torch
|
| 347 |
+
from transformers import AutoProcessor, AutoModelForCausalLM, AutoModelForMultimodalLM
|
| 348 |
+
|
| 349 |
+
TARGET_MODEL_ID = "google/gemma-4-E4B-it"
|
| 350 |
+
ASSISTANT_MODEL_ID = "google/gemma-4-E4B-it-assistant"
|
| 351 |
+
|
| 352 |
+
# Target Model
|
| 353 |
+
processor = AutoProcessor.from_pretrained(TARGET_MODEL_ID)
|
| 354 |
+
target_model = AutoModelForMultimodalLM.from_pretrained(
|
| 355 |
+
TARGET_MODEL_ID,
|
| 356 |
+
torch_dtype=torch.bfloat16,
|
| 357 |
+
device_map="auto",
|
| 358 |
+
|
| 359 |
+
)
|
| 360 |
+
|
| 361 |
+
# Assistant Model (the drafter)
|
| 362 |
+
assistant_model = AutoModelForCausalLM.from_pretrained(
|
| 363 |
+
ASSISTANT_MODEL_ID,
|
| 364 |
+
torch_dtype=torch.bfloat16,
|
| 365 |
+
device_map="auto",
|
| 366 |
+
)
|
| 367 |
+
```
|
| 368 |
+
|
| 369 |
+
Once the model is loaded, you can start generating output by directly referencing the video URL in the prompt:
|
| 370 |
+
|
| 371 |
+
|
| 372 |
+
```python
|
| 373 |
+
# Prompt - add video before text
|
| 374 |
+
messages = [
|
| 375 |
+
{
|
| 376 |
+
'role': 'user',
|
| 377 |
+
'content': [
|
| 378 |
+
{"type": "video", "video": "https://github.com/bebechien/gemma/raw/refs/heads/main/videos/ForBiggerBlazes.mp4"},
|
| 379 |
+
{'type': 'text', 'text': 'Describe this video.'}
|
| 380 |
+
]
|
| 381 |
+
}
|
| 382 |
+
]
|
| 383 |
+
|
| 384 |
+
# Process input
|
| 385 |
+
inputs = processor.apply_chat_template(
|
| 386 |
+
messages,
|
| 387 |
+
tokenize=True,
|
| 388 |
+
return_dict=True,
|
| 389 |
+
return_tensors="pt",
|
| 390 |
+
add_generation_prompt=True,
|
| 391 |
+
).to(target_model.device)
|
| 392 |
+
input_len = inputs["input_ids"].shape[-1]
|
| 393 |
+
|
| 394 |
+
# Generate output
|
| 395 |
+
outputs = target_model.generate(**inputs, max_new_tokens=512)
|
| 396 |
+
response = processor.decode(outputs[0][input_len:], skip_special_tokens=False)
|
| 397 |
+
|
| 398 |
+
# Parse output
|
| 399 |
+
processor.parse_response(response)
|
| 400 |
+
```
|
| 401 |
+
|
| 402 |
+
</details>
|
| 403 |
+
|
| 404 |
+
|
| 405 |
+
|
| 406 |
+
|
| 407 |
+
## **Best Practices**
|
| 408 |
+
|
| 409 |
+
For the best performance, use these configurations and best practices:
|
| 410 |
+
|
| 411 |
+
### 1. Sampling Parameters
|
| 412 |
+
|
| 413 |
+
Use the following standardized sampling configuration across all use cases:
|
| 414 |
+
|
| 415 |
+
* `temperature=1.0`
|
| 416 |
+
* `top_p=0.95`
|
| 417 |
+
* `top_k=64`
|
| 418 |
+
|
| 419 |
+
### 2. Thinking Mode Configuration
|
| 420 |
+
|
| 421 |
+
Compared to Gemma 3, the models use standard `system`, `assistant`, and `user` roles. To properly manage the thinking process, use the following control tokens:
|
| 422 |
+
|
| 423 |
+
* **Trigger Thinking:** Thinking is enabled by including the `<|think|>` token at the start of the system prompt. To disable thinking, remove the token.
|
| 424 |
+
* **Standard Generation:** When thinking is enabled, the model will output its internal reasoning followed by the final answer using this structure:
|
| 425 |
+
`<|channel>thought\n`**[Internal reasoning]**`<channel|>`
|
| 426 |
+
* **Disabled Thinking Behavior:** For all models except for the E2B and E4B variants, if thinking is disabled, the model will still generate the tags but with an empty thought block:
|
| 427 |
+
`<|channel>thought\n<channel|>`**[Final answer]**
|
| 428 |
+
|
| 429 |
+
> [!Note]
|
| 430 |
+
> Note that many libraries like Transformers and llama.cpp handle the complexities of the chat template for you.
|
| 431 |
+
|
| 432 |
+
### 3. Multi-Turn Conversations
|
| 433 |
+
|
| 434 |
+
* **No Thinking Content in History**: In multi-turn conversations, the historical model output should only include the final response. Thoughts from previous model turns must *not be added* before the next user turn begins, with the exception of tool call turns where thinking content should be preserved.
|
| 435 |
+
|
| 436 |
+
### 4. Modality order
|
| 437 |
+
|
| 438 |
+
For optimal performance with multimodal inputs, place:
|
| 439 |
+
|
| 440 |
+
* Image content **before** the text in your prompt.
|
| 441 |
+
* Audio content **after** the text in your prompt.
|
| 442 |
+
|
| 443 |
+
### 5. Variable Image Resolution
|
| 444 |
+
|
| 445 |
+
Aside from variable aspect ratios, Gemma 4 supports variable image resolution through a configurable visual token budget, which controls how many tokens are used to represent an image. A higher token budget preserves more visual detail at the cost of additional compute, while a lower budget enables faster inference for tasks that don't require fine-grained understanding.
|
| 446 |
+
|
| 447 |
+
* The supported token budgets are: **70**, **140**, **280**, **560**, and **1120**.
|
| 448 |
+
* Use *lower budgets* for classification, captioning, or video understanding, where faster inference and processing many frames outweigh fine-grained detail.
|
| 449 |
+
* Use *higher budgets* for tasks like OCR, document parsing, or reading small text.
|
| 450 |
+
|
| 451 |
+
### 6. Audio
|
| 452 |
+
|
| 453 |
+
Use the following prompt structures for audio processing:
|
| 454 |
+
|
| 455 |
+
* **Audio Speech Recognition (ASR)**
|
| 456 |
+
|
| 457 |
+
```text
|
| 458 |
+
Transcribe the following speech segment in {LANGUAGE} into {LANGUAGE} text.
|
| 459 |
+
|
| 460 |
+
Follow these specific instructions for formatting the answer:
|
| 461 |
+
* Only output the transcription, with no newlines.
|
| 462 |
+
* When transcribing numbers, write the digits, i.e. write 1.7 and not one point seven, and write 3 instead of three.
|
| 463 |
+
```
|
| 464 |
+
|
| 465 |
+
* **Automatic Speech Translation (AST)**
|
| 466 |
+
|
| 467 |
+
```text
|
| 468 |
+
Transcribe the following speech segment in {SOURCE_LANGUAGE}, then translate it into {TARGET_LANGUAGE}.
|
| 469 |
+
When formatting the answer, first output the transcription in {SOURCE_LANGUAGE}, then one newline, then output the string '{TARGET_LANGUAGE}: ', then the translation in {TARGET_LANGUAGE}.
|
| 470 |
+
```
|
| 471 |
+
|
| 472 |
+
### 7. Audio and Video Length
|
| 473 |
+
|
| 474 |
+
All models support image inputs and can process videos as frames whereas the E2B, E4B, and 12B models also support audio inputs. Audio supports a maximum length of 30 seconds. Video supports a maximum of 60 seconds assuming the images are processed at one frame per second.
|
| 475 |
+
|
| 476 |
+
## **Model Data**
|
| 477 |
+
|
| 478 |
+
Data used for model training and how the data was processed.
|
| 479 |
+
|
| 480 |
+
### **Training Dataset**
|
| 481 |
+
|
| 482 |
+
Our pre-training dataset is a large-scale, diverse collection of data encompassing a wide range of domains and modalities, which includes web documents, code, images, audio, with a cutoff date of January 2025. Here are the key components:
|
| 483 |
+
|
| 484 |
+
* **Web Documents**: A diverse collection of web text ensures the model is exposed to a broad range of linguistic styles, topics, and vocabulary. The training dataset includes content in over 140 languages.
|
| 485 |
+
* **Code**: Exposing the model to code helps it to learn the syntax and patterns of programming languages, which improves its ability to generate code and understand code-related questions.
|
| 486 |
+
* **Mathematics**: Training on mathematical text helps the model learn logical reasoning, symbolic representation, and to address mathematical queries.
|
| 487 |
+
* **Images**: A wide range of images enables the model to perform image analysis and visual data extraction tasks.
|
| 488 |
+
|
| 489 |
+
The combination of these diverse data sources is crucial for training a powerful multimodal model that can handle a wide variety of different tasks and data formats.
|
| 490 |
+
|
| 491 |
+
### **Data Preprocessing**
|
| 492 |
+
|
| 493 |
+
Here are the key data cleaning and filtering methods applied to the training data:
|
| 494 |
+
|
| 495 |
+
* **CSAM Filtering**: Rigorous CSAM (Child Sexual Abuse Material) filtering was applied at multiple stages in the data preparation process to ensure the exclusion of harmful and illegal content.
|
| 496 |
+
* **Sensitive Data Filtering**: As part of making Gemma pre-trained models safe and reliable, automated techniques were used to filter out certain personal information and other sensitive data from training sets.
|
| 497 |
+
* **Additional methods**: Filtering based on content quality and safety in line with [our policies](https://ai.google/static/documents/ai-responsibility-update-published-february-2025.pdf).
|
| 498 |
+
|
| 499 |
+
## **Ethics and Safety**
|
| 500 |
+
|
| 501 |
+
As open models become central to enterprise infrastructure, provenance and security are paramount. Developed by Google DeepMind, Gemma 4 undergoes the same rigorous safety evaluations as our proprietary Gemini models.
|
| 502 |
+
|
| 503 |
+
### **Evaluation Approach**
|
| 504 |
+
|
| 505 |
+
Gemma 4 models were developed in partnership with internal safety and responsible AI teams. A range of automated as well as human evaluations were conducted to help improve model safety. These evaluations align with [Google’s AI principles](https://ai.google/principles/), as well as safety policies, which aim to prevent our generative AI models from generating harmful content, including:
|
| 506 |
+
|
| 507 |
+
* Content related to child sexual abuse material and exploitation
|
| 508 |
+
* Dangerous content (e.g., promoting suicide, or instructing in activities that could cause real-world harm)
|
| 509 |
+
* Sexually explicit content
|
| 510 |
+
* Hate speech (e.g., dehumanizing members of protected groups)
|
| 511 |
+
* Harassment (e.g., encouraging violence against people)
|
| 512 |
+
|
| 513 |
+
### **Evaluation Results**
|
| 514 |
+
|
| 515 |
+
For all areas of safety testing, we saw major improvements in all categories of content safety relative to previous Gemma models. Overall, Gemma 4 models significantly outperform Gemma 3 and 3n models in improving safety, while keeping unjustified refusals low. All testing was conducted without safety filters to evaluate the model capabilities and behaviors. For both text-to-text and image-to-text, and across all model sizes, the model produced minimal policy violations, and showed significant improvements over previous Gemma models' performance.
|
| 516 |
+
|
| 517 |
+
## **Usage and Limitations**
|
| 518 |
+
|
| 519 |
+
These models have certain limitations that users should be aware of.
|
| 520 |
+
|
| 521 |
+
### **Intended Usage**
|
| 522 |
+
|
| 523 |
+
Multimodal models (capable of processing vision, language, and/or audio) have a wide range of applications across various industries and domains. The following list of potential uses is not comprehensive. The purpose of this list is to provide contextual information about the possible use-cases that the model creators considered as part of model training and development.
|
| 524 |
+
|
| 525 |
+
* **Content Creation and Communication**
|
| 526 |
+
* **Text Generation**: These models can be used to generate creative text formats such as poems, scripts, code, marketing copy, and email drafts.
|
| 527 |
+
* **Chatbots and Conversational AI**: Power conversational interfaces for customer service, virtual assistants, or interactive applications.
|
| 528 |
+
* **Text Summarization**: Generate concise summaries of a text corpus, research papers, or reports.
|
| 529 |
+
* **Image Data Extraction**: These models can be used to extract, interpret, and summarize visual data for text communications.
|
| 530 |
+
* **Audio Processing and Interaction**: The E2B, E4B, and 12B models can analyze and interpret audio inputs, enabling voice-driven interactions and transcriptions.
|
| 531 |
+
* **Research and Education**
|
| 532 |
+
* **Natural Language Processing (NLP) and VLM Research**: These models can serve as a foundation for researchers to experiment with VLM and NLP techniques, develop algorithms, and contribute to the advancement of the field.
|
| 533 |
+
* **Language Learning Tools**: Support interactive language learning experiences, aiding in grammar correction or providing writing practice.
|
| 534 |
+
* **Knowledge Exploration**: Assist researchers in exploring large bodies of text by generating summaries or answering questions about specific topics.
|
| 535 |
+
|
| 536 |
+
### **Limitations**
|
| 537 |
+
|
| 538 |
+
* **Training Data**
|
| 539 |
+
* The quality and diversity of the training data significantly influence the model's capabilities. Biases or gaps in the training data can lead to limitations in the model's responses.
|
| 540 |
+
* The scope of the training dataset determines the subject areas the model can handle effectively.
|
| 541 |
+
* **Context and Task Complexity**
|
| 542 |
+
* Models perform well on tasks that can be framed with clear prompts and instructions. Open-ended or highly complex tasks might be challenging.
|
| 543 |
+
* A model's performance can be influenced by the amount of context provided (longer context generally leads to better outputs, up to a certain point).
|
| 544 |
+
* **Language Ambiguity and Nuance**
|
| 545 |
+
* Natural language is inherently complex. Models might struggle to grasp subtle nuances, sarcasm, or figurative language.
|
| 546 |
+
* **Factual Accuracy**
|
| 547 |
+
* Models generate responses based on information they learned from their training datasets, but they are not knowledge bases. They may generate incorrect or outdated factual statements.
|
| 548 |
+
* **Common Sense**
|
| 549 |
+
* Models rely on statistical patterns in language. They might lack the ability to apply common sense reasoning in certain situations.
|
| 550 |
+
|
| 551 |
+
### **Ethical Considerations and Risks**
|
| 552 |
+
|
| 553 |
+
The development of vision-language models (VLMs) raises several ethical concerns. In creating an open model, we have carefully considered the following:
|
| 554 |
+
|
| 555 |
+
* **Bias and Fairness**
|
| 556 |
+
* VLMs trained on large-scale, real-world text and image data can reflect socio-cultural biases embedded in the training material. Gemma 4 models underwent careful scrutiny, input data pre-processing, and post-training evaluations as reported in this card to help mitigate the risk of these biases.
|
| 557 |
+
* **Misinformation and Misuse**
|
| 558 |
+
* VLMs can be misused to generate text that is false, misleading, or harmful.
|
| 559 |
+
* Guidelines are provided for responsible use with the model, see the [Responsible Generative AI Toolkit](https://ai.google.dev/responsible).
|
| 560 |
+
* **Transparency and Accountability**
|
| 561 |
+
* This model card summarizes details on the models' architecture, capabilities, limitations, and evaluation processes.
|
| 562 |
+
* A responsibly developed open model offers the opportunity to share innovation by making VLM technology accessible to developers and researchers across the AI ecosystem.
|
| 563 |
+
|
| 564 |
+
**Risks identified and mitigations**:
|
| 565 |
+
|
| 566 |
+
* **Generation of harmful content**: Mechanisms and guidelines for content safety are essential. Developers are encouraged to exercise caution and implement appropriate content safety safeguards based on their specific product policies and application use cases.
|
| 567 |
+
* **Misuse for malicious purposes**: Technical limitations and developer and end-user education can help mitigate against malicious applications of VLMs. Educational resources and reporting mechanisms for users to flag misuse are provided.
|
| 568 |
+
* **Privacy violations**: Models were trained on data filtered for removal of certain personal information and other sensitive data. Developers are encouraged to adhere to privacy regulations with privacy-preserving techniques.
|
| 569 |
+
* **Perpetuation of biases**: It's encouraged to perform continuous monitoring (using evaluation metrics, human review) and the exploration of de-biasing techniques during model training, fine-tuning, and other use cases.
|
| 570 |
+
|
| 571 |
+
### **Benefits**
|
| 572 |
+
|
| 573 |
+
At the time of release, this family of models provides high-performance open vision-language model implementations designed from the ground up for responsible AI development compared to similarly sized models.
|
| 574 |
+
|
| 575 |
+
## **Citation**
|
| 576 |
+
|
| 577 |
+
If you find our work helpful, please consider citing it:
|
| 578 |
+
|
| 579 |
+
```bibtex
|
| 580 |
+
@misc{gemmateam2026gemma4,
|
| 581 |
+
title={Gemma 4 Technical Report},
|
| 582 |
+
author={Gemma Team},
|
| 583 |
+
year={2026},
|
| 584 |
+
eprint={2607.02770},
|
| 585 |
+
archivePrefix={arXiv},
|
| 586 |
+
primaryClass={cs.CL},
|
| 587 |
+
url={https://arxiv.org/abs/2607.02770},
|
| 588 |
+
}
|
| 589 |
+
```
|
docs/assistants/provenance.json
ADDED
|
@@ -0,0 +1,54 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"schema_version": 1,
|
| 3 |
+
"upstream": {
|
| 4 |
+
"repository": "google/gemma-4-E4B-it-assistant",
|
| 5 |
+
"revision": "8d0031ea8c2109e2b1e86bb9368a4539b537f80a",
|
| 6 |
+
"license": "Apache-2.0",
|
| 7 |
+
"license_source": "https://ai.google.dev/gemma/apache_2",
|
| 8 |
+
"gated": false,
|
| 9 |
+
"files": [
|
| 10 |
+
{
|
| 11 |
+
"file": "config.json",
|
| 12 |
+
"bytes": 2315,
|
| 13 |
+
"sha256": "13cdf5418e5a89cf951717a953cd11cef4e5c2743c4767e9e5dd72802954f822",
|
| 14 |
+
"url": "https://huggingface.co/google/gemma-4-E4B-it-assistant/resolve/8d0031ea8c2109e2b1e86bb9368a4539b537f80a/config.json"
|
| 15 |
+
},
|
| 16 |
+
{
|
| 17 |
+
"file": "generation_config.json",
|
| 18 |
+
"bytes": 307,
|
| 19 |
+
"sha256": "8e58004dc0e2407b63410b190bb8470efbdcfeb71533f1770e09c20abe193a6f",
|
| 20 |
+
"url": "https://huggingface.co/google/gemma-4-E4B-it-assistant/resolve/8d0031ea8c2109e2b1e86bb9368a4539b537f80a/generation_config.json"
|
| 21 |
+
},
|
| 22 |
+
{
|
| 23 |
+
"file": "tokenizer_config.json",
|
| 24 |
+
"bytes": 822,
|
| 25 |
+
"sha256": "089594a3924fcfd4cb1c596a7906fbf476193519e5198f780912eed02b177e42",
|
| 26 |
+
"url": "https://huggingface.co/google/gemma-4-E4B-it-assistant/resolve/8d0031ea8c2109e2b1e86bb9368a4539b537f80a/tokenizer_config.json"
|
| 27 |
+
},
|
| 28 |
+
{
|
| 29 |
+
"file": "tokenizer.json",
|
| 30 |
+
"bytes": 32169440,
|
| 31 |
+
"sha256": "75a6583c1a418e2bbd79c60d95d28e0f5bf549ad3f2990b5bdb5238c6c2bf70c",
|
| 32 |
+
"url": "https://huggingface.co/google/gemma-4-E4B-it-assistant/resolve/8d0031ea8c2109e2b1e86bb9368a4539b537f80a/tokenizer.json"
|
| 33 |
+
},
|
| 34 |
+
{
|
| 35 |
+
"file": "model.safetensors",
|
| 36 |
+
"bytes": 159138208,
|
| 37 |
+
"sha256": "12875062fc25c51e8fa9b62abd2de7ad48b7d63f8559d5d604fbd5a3d6bcff16",
|
| 38 |
+
"url": "https://huggingface.co/google/gemma-4-E4B-it-assistant/resolve/8d0031ea8c2109e2b1e86bb9368a4539b537f80a/model.safetensors"
|
| 39 |
+
}
|
| 40 |
+
]
|
| 41 |
+
},
|
| 42 |
+
"conversion": {
|
| 43 |
+
"output_file": "Gemma-4-E4B-IT-Assistant-BF16.gguf",
|
| 44 |
+
"bytes": 171766688,
|
| 45 |
+
"sha256": "4e3c9d335b248ced9bd3e0584547efca30dbf6cfd0debb810613f811a575162a",
|
| 46 |
+
"CPU_only": true,
|
| 47 |
+
"output_type": "BF16 GGUF",
|
| 48 |
+
"converter_repository": "https://github.com/ggml-org/llama.cpp",
|
| 49 |
+
"converter_revision": "911f6cdc8ab8a530b2bee09ee61471a6f3178eeb",
|
| 50 |
+
"converter_sha256": "e9a1da876330bbce9687541ab31736542a01b4ac43c6686126514a50f122fb7f",
|
| 51 |
+
"gemma_converter_sha256": "872f3fea7496cefdcd39308ad02e2d30c9f507f68a18f8b4f6962995b09350f5",
|
| 52 |
+
"changed": "Serialization format only; no training."
|
| 53 |
+
}
|
| 54 |
+
}
|
gguf/Gemma-4-E4B-IT-Assistant-BF16.gguf
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:4e3c9d335b248ced9bd3e0584547efca30dbf6cfd0debb810613f811a575162a
|
| 3 |
+
size 171766688
|
release-manifest.json
CHANGED
|
@@ -4,7 +4,6 @@
|
|
| 4 |
"base_model": "google/gemma-4-E4B-it",
|
| 5 |
"base_revision": "ee0ef6023621cff504d758262d4e04895a5af4a2",
|
| 6 |
"license": "apache-2.0",
|
| 7 |
-
"selected_checkpoint": "targeted:438",
|
| 8 |
"training_method": "rank-32 alpha-64 LoRA on language tensors; FP32 adapter merge",
|
| 9 |
"converter_revision": "911f6cdc8ab8a530b2bee09ee61471a6f3178eeb",
|
| 10 |
"inference_source_commit": "77d1458",
|
|
@@ -23,6 +22,15 @@
|
|
| 23 |
"path": "gguf/mmproj-Winnow-E4B.gguf",
|
| 24 |
"bytes": 990372672,
|
| 25 |
"sha256": "ddf46c21d7078e95338cfc22306b19b276a29a5ad089023449dd54d4b6170a51"
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 26 |
}
|
| 27 |
},
|
| 28 |
"decision_calibration": {
|
|
@@ -67,5 +75,21 @@
|
|
| 67 |
"kv_cache": "q8_0",
|
| 68 |
"decision_parallel": 4,
|
| 69 |
"scope": "synthetic-image operational smoke test; not broad vision or long-context quality"
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 70 |
}
|
| 71 |
}
|
|
|
|
| 4 |
"base_model": "google/gemma-4-E4B-it",
|
| 5 |
"base_revision": "ee0ef6023621cff504d758262d4e04895a5af4a2",
|
| 6 |
"license": "apache-2.0",
|
|
|
|
| 7 |
"training_method": "rank-32 alpha-64 LoRA on language tensors; FP32 adapter merge",
|
| 8 |
"converter_revision": "911f6cdc8ab8a530b2bee09ee61471a6f3178eeb",
|
| 9 |
"inference_source_commit": "77d1458",
|
|
|
|
| 22 |
"path": "gguf/mmproj-Winnow-E4B.gguf",
|
| 23 |
"bytes": 990372672,
|
| 24 |
"sha256": "ddf46c21d7078e95338cfc22306b19b276a29a5ad089023449dd54d4b6170a51"
|
| 25 |
+
},
|
| 26 |
+
"mtp_assistant_bf16": {
|
| 27 |
+
"path": "gguf/Gemma-4-E4B-IT-Assistant-BF16.gguf",
|
| 28 |
+
"sha256": "4e3c9d335b248ced9bd3e0584547efca30dbf6cfd0debb810613f811a575162a",
|
| 29 |
+
"bytes": 171766688,
|
| 30 |
+
"source_model": "google/gemma-4-E4B-it-assistant",
|
| 31 |
+
"source_revision": "8d0031ea8c2109e2b1e86bb9368a4539b537f80a",
|
| 32 |
+
"license": "apache-2.0",
|
| 33 |
+
"documentation": "docs/assistants/README.md"
|
| 34 |
}
|
| 35 |
},
|
| 36 |
"decision_calibration": {
|
|
|
|
| 75 |
"kv_cache": "q8_0",
|
| 76 |
"decision_parallel": 4,
|
| 77 |
"scope": "synthetic-image operational smoke test; not broad vision or long-context quality"
|
| 78 |
+
},
|
| 79 |
+
"training_pipeline_released": false,
|
| 80 |
+
"optional_inference": {
|
| 81 |
+
"repository": "https://github.com/EldanRing/winnow-inference",
|
| 82 |
+
"runtime_lock_sha256": "9560b74c9fd736c80880c2733569075e066e7e8f49a7b465c8a4589604c4f65e",
|
| 83 |
+
"context": 8192,
|
| 84 |
+
"native_parallel": 4,
|
| 85 |
+
"chat_parallel": 1,
|
| 86 |
+
"cache": "q8_0",
|
| 87 |
+
"scope": "Linux/CUDA singleGPU tested RTX5070Ti16GB; experimental adaptive calibration text only",
|
| 88 |
+
"supported_presets": [
|
| 89 |
+
"e4b-q8-vision8k-mtp"
|
| 90 |
+
],
|
| 91 |
+
"adaptive_policies": [
|
| 92 |
+
"e4b-calibrated50-v1"
|
| 93 |
+
]
|
| 94 |
}
|
| 95 |
}
|