How to use from the
Use from the
sentence-transformers library
from sentence_transformers import SentenceTransformer

model = SentenceTransformer("webmp3/Sakura-EmbeddingGemma-2-AutoRound")

sentences = [
    "The weather is lovely today.",
    "It's so sunny outside!",
    "He drove to the stadium."
]
embeddings = model.encode(sentences)

similarities = model.similarity(embeddings, embeddings)
print(similarities.shape)
# [3, 3]

Sakura EmbeddingGemma 2 — AutoRound W4A16

Community Quantization. Not official from Google.
Quantized and evaluated by Sakura (webmp3) using Intel AutoRound optimization on top of Google's official embeddinggemma-2 architecture.


Overview

This repository provides an optimized AutoRound W4A16 (4-bit weights, 16-bit activations) quantization of google/embeddinggemma-2.

It is packaged as a standard Hugging Face repository containing packed Safetensors weights that run natively and out of the box with sentence-transformers and transformers.

  • Upstream Model: google/embeddinggemma-2
  • Pinned Upstream Commit: 914f7f89142e33e77833254d9c9b90c3cef7303b
  • Base Architecture: EmbeddingGemma2ForSequenceClassification / EmbeddingGemma2TextModel
  • Model Size: 1.24 GB total directory (packed model.safetensors: 1,236 MB / 1.21 GiB)
  • Verification Status: Public Hub access verified; local release smoke test passed; packaged for Transformers & SentenceTransformers.
  • Quantization Framework: AutoRound (v0.8+)
  • Quantization Scheme: W4A16 Symmetric (group_size=128, 100 tuning iterations, sign gradient descent with Hessian approximation)
  • Release Decision: RELEASE GO (Early community release based on empirical fidelity retention)

Architectural Breakdown: What is W4 vs. What Remains BF16

We explicitly do not claim "Full INT4/Q4". High-fidelity embedding models require careful treatment of sensitive components:

Component Precision Details
Text Backbone Linear Layers W4A16 216 Linear layers in language_model.layers.0 through language_model.layers.23 (q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj). Group size 128, symmetric.
Token Embeddings BF16 language_model.embed_tokens (256,000 vocab) preserved in bfloat16 to avoid semantic vocabulary collapse.
Embedding Projection Head BF16 language_model.embedding_projection (768d output head) preserved in bfloat16 to preserve precise directional geometry.
Normalization Layers BF16 All RMSNorm and LayerNorm modules preserved in bfloat16 to prevent activation scale clipping.
Vision & Audio Towers BF16 Multimodal encoders (vision_tower, audio_tower) and multimodal projection heads are preserved unquantized.

Fidelity & Benchmark Results

All evaluations were conducted against an unquantized bfloat16 reference across 100 query/document retrieval pairs (English and German) and 50 code retrieval pairs.

Release Gate Status

The candidate was evaluated against strict quality criteria:

  • Empirical Gate Result: RELEASE GO / Early Community Release
  • Gate Context: The initial theoretical target of $\ge 0.99$ Mean Cosine Similarity and $\ge 95%$ Top-5 Retrieval Agreement was not fully reached in every dimension (achieved: 0.9878 Mean Cosine, 87.6% Top-5 Agreement). However, with 99.0% Top-1 Agreement, 100.0% Recall@5, and 0.9781 Spearman Correlation, the model demonstrates outstanding retrieval fidelity and semantic stability without a single NaN or Inf anomaly.

MRL (Matryoshka Representation Learning) Dimensions

EmbeddingGemma 2 supports dimension truncation followed by L2-renormalization. The W4A16 model demonstrates robust stability across all standard MRL truncations:

MRL Dimension Mean Cosine Sim Min Cosine Sim Spearman Rank Corr Mean L2 Drift Top-1 Agreement Top-5 Agreement Recall@5 nDCG@10
768d (Full) 0.9878 0.9671 0.9781 0.1557 99.0 % 87.6 % 100.0 % 0.9655
512d 0.9882 0.9675 0.9782 0.1531 99.0 % 88.4 % 100.0 % 0.9583
256d 0.9894 0.9691 0.9774 0.1451 100.0 % 86.2 % 100.0 % 0.9595
128d 0.9922 0.9734 0.9744 0.1246 100.0 % 84.8 % 100.0 % 0.9482

Language & Domain Subsets (768d)

Subset Mean Cosine Sim Top-1 Agreement Top-5 Agreement Recall@5
English Queries & Docs 0.9890 100.0 % 88.0 % 100.0 %
German Queries & Docs 0.9867 98.0 % 87.2 % 100.0 %
Code Retrieval 0.9859 100.0 % 88.0 % 100.0 %

Runtime & Community Format Comparison

Distribution / Format Runtime Compatibility Direct Transformers / ST Support Notes
Sakura AutoRound W4A16 (This repo) Python, PyTorch, Transformers, SentenceTransformers Yes (Plug-and-play) Runs directly in existing Python AI pipelines without needing custom binary builds.
GGUF Community Releases (unsloth, ggml-org) llama.cpp No (requires llama-server or bindings) During this release test, the downloaded community GGUF could not be loaded with the older local llama.cpp build used for evaluation (unknown model architecture). Current llama.cpp versions include dedicated Gemma embedding support, so this should not be interpreted as a general llama.cpp limitation.
ONNX Community Releases (onnx-community) ONNX Runtime / Transformers.js No (ONNX graph format) Modular multi-graph export tailored for WebGPU/JavaScript execution.
Sakura AutoRound GGUF (Companion) llama.cpp Via llama-server / bindings GGUF runtime variant: webmp3/Sakura-EmbeddingGemma-2-AutoRound-GGUF

Quickstart & Usage

1. With SentenceTransformers (Recommended)

from sentence_transformers import SentenceTransformer
import torch

# Load the quantized model
model = SentenceTransformer(
    "webmp3/Sakura-EmbeddingGemma-2-AutoRound",
    model_kwargs={"torch_dtype": torch.bfloat16}
)

# Text Retrieval Query (using official prompt_name)
query = "What is quantum entanglement?"
query_embedding = model.encode(query, prompt_name="SearchQuery")

# Documents (unprompted)
docs = [
    "Quantum entanglement is a phenomenon where particles remain connected regardless of distance.",
    "The recipe for chocolate chip cookies requires flour, butter, and sugar."
]
doc_embeddings = model.encode(docs)

# Compute similarity
similarities = model.similarity(query_embedding, doc_embeddings)
print("Similarities:", similarities)

2. Matryoshka Dimension Truncation (MRL)

To reduce memory and storage footprint, simply slice the vector and re-normalize:

import torch
import torch.nn.functional as F

# 128-dimensional embedding
full_embedding = model.encode(["Example sentence"], convert_to_tensor=True)
mrl_128 = full_embedding[:, :128]
mrl_128_normalized = F.normalize(mrl_128, p=2, dim=-1)
print("128d shape:", mrl_128_normalized.shape)

3. With Hugging Face Transformers

from transformers import AutoModel, AutoTokenizer
import torch

model_id = "webmp3/Sakura-EmbeddingGemma-2-AutoRound"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModel.from_pretrained(model_id, torch_dtype=torch.bfloat16)

inputs = tokenizer(["SearchQuery: What is machine learning?"], return_tensors="pt")
with torch.no_grad():
    outputs = model(**inputs)

Limitations & Honest Disclosure

  1. Text Precision: The 216 linear layers of the text backbone are quantized to 4-bit weights (torch_zp / GPTQ-style packed format). While memory footprint is significantly reduced (~1.24 GB vs ~5.5 GB for full BF16), low-bit weights inherently involve quantization approximations.
  2. Multimodal Towers: The vision and audio towers remain in 16-bit precision. If you do not use vision or audio inputs, memory consumption can be minimized by only loading the text language model.
  3. Execution Device: Optimal execution requires modern CPU (AVX-512 / VNNI) or GPU environments supporting accelerated bfloat16 and packed int4 kernels.

SHA256 Checksums

Every file in this release has been cryptographically verified:

ed5b25820992aef4a31b99a216fa369f1845f0b5001165866d82c81128c4cbc3  model.safetensors
ee5befff1a18299a2acdf77e2895cde534fcfd73d2793b7a12df644be0e6acdd  config.json
79bac98844a054c881016641b7ad2b97539c9a43c3bdd26205823e8fe1c85a15  quantization_config.json
031e56a498d33c349ab489a21885bcfe25b4fcba841149dc99e1e90d4a7c28f5  config_sentence_transformers.json
b1bcd9f2dce3ae863b359e87d0710b5dbc3314a59ecb4e2f97c7778fc8e4b228  sentence_bert_config.json
3d02572a0455b832de67fb8e63a54981bc7e8b46e337c95e917bd8122a533bfd  modules.json
ea2ae257e901064abdd98dceb19f2b0da06af600bed15e0f99f5c85c37ee9d78  preprocessor_config.json
168f6a08522f3ce5dea596d94d003af2fd691742d4f41fe1f9d8cce76bfbf69c  processor_config.json
4d777ef5bdc1aa36227abdfb77c3e49e7b9c892d16e1b6bda41c393504828be4  tokenizer.json
17bd5d6e9364ca49a534e1502076593317c298d4a663623091ed45388f004874  tokenizer_config.json
4b852efc0b9960283e735363331e6f325b33bc74bdbaa076f595bc4e9b94d85e  chat_template.jinja
8759bdf7c77efc7df7723f64856a593c8943b71ee38baf2a88771fbaf78438f9  1_Pooling/config.json
cdb09dfca347a56aa2d691744e38d5ad3c7cbc2834e7181272b9a15328b82524  2_Normalize/config.json

Citation & Acknowledgements

Downloads last month
-
Safetensors
Model size
0.6B params
Tensor type
I32
·
BF16
·
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for webmp3/Sakura-EmbeddingGemma-2-AutoRound

Quantized
(19)
this model