# Optional E2B MTP assistant Direct decisions and ordinary chat need only the Winnow-E2B target. MTP additionally uses the bundled BF16 GGUF conversion of Google's official [Gemma 4 E2B IT assistant](https://huggingface.co/google/gemma-4-E2B-it-assistant/tree/2d874ef7d29f9a30599a1e4b3c1cbc9595f005df). No conversion is required to use this download. ## Download the evaluated assistant Download the Q8 target and its matching assistant from the model repository: ```sh huggingface-cli download EldanRing/Winnow-E2B \ gguf/Winnow-E2B-Q8_0.gguf \ gguf/Gemma-4-E2B-IT-Assistant-BF16.gguf \ --local-dir Winnow-E2B ``` Verify both files against [SHA256SUMS](../SHA256SUMS). An E4B or 12B assistant is not interchangeable with the E2B assistant. Add `gguf/mmproj-Winnow-E2B.gguf` for image input. The assistant's BF16 precision is distinct from the optional `Winnow-E2B-BF16.gguf` target. Reported MTP4 results use Q8 target weights with BF16 assistant weights and shared F16 target K/V cache. ## Provenance The bundled assistant is the evaluated CPU conversion from upstream revision `2d874ef7d29f9a30599a1e4b3c1cbc9595f005df`, made with `convert_hf_to_gguf.py` at pinned [llama.cpp revision](https://github.com/ggml-org/llama.cpp/tree/911f6cdc8ab8a530b2bee09ee61471a6f3178eeb) and BF16 output. Conversion changes serialization and performs no training. A different conversion does not inherit these measurements. | Artifact | Bytes | SHA-256 | |---|---:|---| | Upstream `model.safetensors` | 157,565,344 | `93682eb1c97639d18f007704dc880bd74cbe530adaf7b1bb561213863fdad2a6` | | Evaluated BF16 GGUF | 170,194,016 | `0772dcc50761a47a2da87151d3105ad2b82994d9dd1dae3ad3c5c0cd26086b16` | The assistant is by Google DeepMind under Apache License 2.0. The [assistant attribution directory](assistants/README.md) includes the license, notice, upstream model card, and conversion provenance. Target, projector, base revision, and assistant identities are recorded in the [release manifest](../release-manifest.json).