Instructions to use webmp3/Sakura-EmbeddingGemma-2-AutoRound with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use webmp3/Sakura-EmbeddingGemma-2-AutoRound with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("webmp3/Sakura-EmbeddingGemma-2-AutoRound") sentences = [ "The weather is lovely today.", "It's so sunny outside!", "He drove to the stadium." ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [3, 3] - Transformers
How to use webmp3/Sakura-EmbeddingGemma-2-AutoRound with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("feature-extraction", model="webmp3/Sakura-EmbeddingGemma-2-AutoRound")# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModel processor = AutoProcessor.from_pretrained("webmp3/Sakura-EmbeddingGemma-2-AutoRound") model = AutoModel.from_pretrained("webmp3/Sakura-EmbeddingGemma-2-AutoRound", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Sakura EmbeddingGemma 2 — AutoRound W4A16
Community Quantization. Not official from Google.
Quantized and evaluated by Sakura (webmp3) using Intel AutoRound optimization on top of Google's officialembeddinggemma-2architecture.
Overview
This repository provides an optimized AutoRound W4A16 (4-bit weights, 16-bit activations) quantization of google/embeddinggemma-2.
It is packaged as a standard Hugging Face repository containing packed Safetensors weights that run natively and out of the box with sentence-transformers and transformers.
- Upstream Model:
google/embeddinggemma-2 - Pinned Upstream Commit:
914f7f89142e33e77833254d9c9b90c3cef7303b - Base Architecture:
EmbeddingGemma2ForSequenceClassification/EmbeddingGemma2TextModel - Model Size: 1.24 GB total directory (packed
model.safetensors: 1,236 MB / 1.21 GiB) - Verification Status: Public Hub access verified; local release smoke test passed; packaged for Transformers & SentenceTransformers.
- Quantization Framework: AutoRound (v0.8+)
- Quantization Scheme: W4A16 Symmetric (
group_size=128, 100 tuning iterations, sign gradient descent with Hessian approximation) - Release Decision: RELEASE GO (Early community release based on empirical fidelity retention)
Architectural Breakdown: What is W4 vs. What Remains BF16
We explicitly do not claim "Full INT4/Q4". High-fidelity embedding models require careful treatment of sensitive components:
| Component | Precision | Details |
|---|---|---|
| Text Backbone Linear Layers | W4A16 | 216 Linear layers in language_model.layers.0 through language_model.layers.23 (q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj). Group size 128, symmetric. |
| Token Embeddings | BF16 | language_model.embed_tokens (256,000 vocab) preserved in bfloat16 to avoid semantic vocabulary collapse. |
| Embedding Projection Head | BF16 | language_model.embedding_projection (768d output head) preserved in bfloat16 to preserve precise directional geometry. |
| Normalization Layers | BF16 | All RMSNorm and LayerNorm modules preserved in bfloat16 to prevent activation scale clipping. |
| Vision & Audio Towers | BF16 | Multimodal encoders (vision_tower, audio_tower) and multimodal projection heads are preserved unquantized. |
Fidelity & Benchmark Results
All evaluations were conducted against an unquantized bfloat16 reference across 100 query/document retrieval pairs (English and German) and 50 code retrieval pairs.
Release Gate Status
The candidate was evaluated against strict quality criteria:
- Empirical Gate Result: RELEASE GO / Early Community Release
- Gate Context: The initial theoretical target of $\ge 0.99$ Mean Cosine Similarity and $\ge 95%$ Top-5 Retrieval Agreement was not fully reached in every dimension (achieved: 0.9878 Mean Cosine, 87.6% Top-5 Agreement). However, with 99.0% Top-1 Agreement, 100.0% Recall@5, and 0.9781 Spearman Correlation, the model demonstrates outstanding retrieval fidelity and semantic stability without a single NaN or Inf anomaly.
MRL (Matryoshka Representation Learning) Dimensions
EmbeddingGemma 2 supports dimension truncation followed by L2-renormalization. The W4A16 model demonstrates robust stability across all standard MRL truncations:
| MRL Dimension | Mean Cosine Sim | Min Cosine Sim | Spearman Rank Corr | Mean L2 Drift | Top-1 Agreement | Top-5 Agreement | Recall@5 | nDCG@10 |
|---|---|---|---|---|---|---|---|---|
| 768d (Full) | 0.9878 | 0.9671 | 0.9781 | 0.1557 | 99.0 % | 87.6 % | 100.0 % | 0.9655 |
| 512d | 0.9882 | 0.9675 | 0.9782 | 0.1531 | 99.0 % | 88.4 % | 100.0 % | 0.9583 |
| 256d | 0.9894 | 0.9691 | 0.9774 | 0.1451 | 100.0 % | 86.2 % | 100.0 % | 0.9595 |
| 128d | 0.9922 | 0.9734 | 0.9744 | 0.1246 | 100.0 % | 84.8 % | 100.0 % | 0.9482 |
Language & Domain Subsets (768d)
| Subset | Mean Cosine Sim | Top-1 Agreement | Top-5 Agreement | Recall@5 |
|---|---|---|---|---|
| English Queries & Docs | 0.9890 | 100.0 % | 88.0 % | 100.0 % |
| German Queries & Docs | 0.9867 | 98.0 % | 87.2 % | 100.0 % |
| Code Retrieval | 0.9859 | 100.0 % | 88.0 % | 100.0 % |
Runtime & Community Format Comparison
| Distribution / Format | Runtime Compatibility | Direct Transformers / ST Support | Notes |
|---|---|---|---|
| Sakura AutoRound W4A16 (This repo) | Python, PyTorch, Transformers, SentenceTransformers | Yes (Plug-and-play) | Runs directly in existing Python AI pipelines without needing custom binary builds. |
GGUF Community Releases (unsloth, ggml-org) |
llama.cpp |
No (requires llama-server or bindings) | During this release test, the downloaded community GGUF could not be loaded with the older local llama.cpp build used for evaluation (unknown model architecture). Current llama.cpp versions include dedicated Gemma embedding support, so this should not be interpreted as a general llama.cpp limitation. |
ONNX Community Releases (onnx-community) |
ONNX Runtime / Transformers.js | No (ONNX graph format) | Modular multi-graph export tailored for WebGPU/JavaScript execution. |
| Sakura AutoRound GGUF (Companion) | llama.cpp |
Via llama-server / bindings | GGUF runtime variant: webmp3/Sakura-EmbeddingGemma-2-AutoRound-GGUF |
Quickstart & Usage
1. With SentenceTransformers (Recommended)
from sentence_transformers import SentenceTransformer
import torch
# Load the quantized model
model = SentenceTransformer(
"webmp3/Sakura-EmbeddingGemma-2-AutoRound",
model_kwargs={"torch_dtype": torch.bfloat16}
)
# Text Retrieval Query (using official prompt_name)
query = "What is quantum entanglement?"
query_embedding = model.encode(query, prompt_name="SearchQuery")
# Documents (unprompted)
docs = [
"Quantum entanglement is a phenomenon where particles remain connected regardless of distance.",
"The recipe for chocolate chip cookies requires flour, butter, and sugar."
]
doc_embeddings = model.encode(docs)
# Compute similarity
similarities = model.similarity(query_embedding, doc_embeddings)
print("Similarities:", similarities)
2. Matryoshka Dimension Truncation (MRL)
To reduce memory and storage footprint, simply slice the vector and re-normalize:
import torch
import torch.nn.functional as F
# 128-dimensional embedding
full_embedding = model.encode(["Example sentence"], convert_to_tensor=True)
mrl_128 = full_embedding[:, :128]
mrl_128_normalized = F.normalize(mrl_128, p=2, dim=-1)
print("128d shape:", mrl_128_normalized.shape)
3. With Hugging Face Transformers
from transformers import AutoModel, AutoTokenizer
import torch
model_id = "webmp3/Sakura-EmbeddingGemma-2-AutoRound"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModel.from_pretrained(model_id, torch_dtype=torch.bfloat16)
inputs = tokenizer(["SearchQuery: What is machine learning?"], return_tensors="pt")
with torch.no_grad():
outputs = model(**inputs)
Limitations & Honest Disclosure
- Text Precision: The 216 linear layers of the text backbone are quantized to 4-bit weights (
torch_zp/ GPTQ-style packed format). While memory footprint is significantly reduced (~1.24 GB vs ~5.5 GB for full BF16), low-bit weights inherently involve quantization approximations. - Multimodal Towers: The vision and audio towers remain in 16-bit precision. If you do not use vision or audio inputs, memory consumption can be minimized by only loading the text language model.
- Execution Device: Optimal execution requires modern CPU (AVX-512 / VNNI) or GPU environments supporting accelerated bfloat16 and packed int4 kernels.
SHA256 Checksums
Every file in this release has been cryptographically verified:
ed5b25820992aef4a31b99a216fa369f1845f0b5001165866d82c81128c4cbc3 model.safetensors
ee5befff1a18299a2acdf77e2895cde534fcfd73d2793b7a12df644be0e6acdd config.json
79bac98844a054c881016641b7ad2b97539c9a43c3bdd26205823e8fe1c85a15 quantization_config.json
031e56a498d33c349ab489a21885bcfe25b4fcba841149dc99e1e90d4a7c28f5 config_sentence_transformers.json
b1bcd9f2dce3ae863b359e87d0710b5dbc3314a59ecb4e2f97c7778fc8e4b228 sentence_bert_config.json
3d02572a0455b832de67fb8e63a54981bc7e8b46e337c95e917bd8122a533bfd modules.json
ea2ae257e901064abdd98dceb19f2b0da06af600bed15e0f99f5c85c37ee9d78 preprocessor_config.json
168f6a08522f3ce5dea596d94d003af2fd691742d4f41fe1f9d8cce76bfbf69c processor_config.json
4d777ef5bdc1aa36227abdfb77c3e49e7b9c892d16e1b6bda41c393504828be4 tokenizer.json
17bd5d6e9364ca49a534e1502076593317c298d4a663623091ed45388f004874 tokenizer_config.json
4b852efc0b9960283e735363331e6f325b33bc74bdbaa076f595bc4e9b94d85e chat_template.jinja
8759bdf7c77efc7df7723f64856a593c8943b71ee38baf2a88771fbaf78438f9 1_Pooling/config.json
cdb09dfca347a56aa2d691744e38d5ad3c7cbc2834e7181272b9a15328b82524 2_Normalize/config.json
Citation & Acknowledgements
- Upstream Model by Google: google/embeddinggemma-2
- AutoRound framework by Intel: Intel AutoRound
- Quantization & Benchmark Pipeline by Sakura (
webmp3).
- Downloads last month
- -
Model tree for webmp3/Sakura-EmbeddingGemma-2-AutoRound
Base model
google/embeddinggemma-2
from sentence_transformers import SentenceTransformer model = SentenceTransformer("webmp3/Sakura-EmbeddingGemma-2-AutoRound") sentences = [ "The weather is lovely today.", "It's so sunny outside!", "He drove to the stadium." ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [3, 3]