Instructions to use EldanRing/Winnow-E2B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use EldanRing/Winnow-E2B with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf EldanRing/Winnow-E2B:BF16 # Run inference directly in the terminal: llama cli -hf EldanRing/Winnow-E2B:BF16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf EldanRing/Winnow-E2B:BF16 # Run inference directly in the terminal: llama cli -hf EldanRing/Winnow-E2B:BF16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf EldanRing/Winnow-E2B:BF16 # Run inference directly in the terminal: ./llama-cli -hf EldanRing/Winnow-E2B:BF16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf EldanRing/Winnow-E2B:BF16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf EldanRing/Winnow-E2B:BF16
Use Docker
docker model run hf.co/EldanRing/Winnow-E2B:BF16
- LM Studio
- Jan
- vLLM
How to use EldanRing/Winnow-E2B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "EldanRing/Winnow-E2B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "EldanRing/Winnow-E2B", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/EldanRing/Winnow-E2B:BF16
- Ollama
How to use EldanRing/Winnow-E2B with Ollama:
ollama run hf.co/EldanRing/Winnow-E2B:BF16
- Unsloth Desktop
- Docker Model Runner
How to use EldanRing/Winnow-E2B with Docker Model Runner:
docker model run hf.co/EldanRing/Winnow-E2B:BF16
- Lemonade
How to use EldanRing/Winnow-E2B with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull EldanRing/Winnow-E2B:BF16
Run and chat with the model
lemonade run user.Winnow-E2B-BF16
List all available models
lemonade list
- Atomic Chat
Download docs/EVALUATION.md from EldanRing/Winnow-E2B: direct link, hf CLI and curl.
- Browser
- Download file 7.46 kB
-
https://huggingface.co/EldanRing/Winnow-E2B/resolve/main/docs/EVALUATION.md
- Command line
-
hf download hf://EldanRing/Winnow-E2B/docs/EVALUATION.md
-
curl -L -o EVALUATION.md https://huggingface.co/EldanRing/Winnow-E2B/resolve/main/docs/EVALUATION.md
Winnow-E2B evaluation
Scope
The completed October 6, 2026 evaluation uses the Q8 target and official BF16 MTP assistant identified in the release manifest, on an RTX 5070 Ti. It evaluates the adaptive reasoning path at source revision ee6bd37d34ae35d2e69ebb0c4b0957a10c727357.
JevBench and Kev use verified source labels. Typed measures synthetic hard-teacher agreement: 2,000 decisions from 400 five-question groups. All three panels were previously exposed during development. They are regression and transfer evidence, not independent estimates of generalization. JevBench means public-subset accuracy, not its official composite leaderboard score. No tuning on these labels occurred in this campaign.
E2B method
The frozen adaptive rule routes when raw direct maximum probability is below 0.99, scores at temperature 1, and combines direct and augmented probabilities 50:50. Only completed natural-EOS generations are rescored. A failed reasoning attempt retains direct probabilities and stays in the denominator.
The text-only run used the e2b-q8-text8k-mtp preset: an 8,192-position context, F16 target KV shared with the pinned MTP assistant, MTP draft length 4, uncapped generation at temperature 0, and seed 314159. The retained draft-cache launch flags were Q8_0, but the assistant attention cache shares the target's F16 K/V tensors; this is not a separately measured Q8 attention-cache pool. The reasoning prompt serializes options with native labels while preserving input escaping and structured-state ownership. This is an actual integrated-path run, not the earlier selective replay.
All 3,277 cases completed with zero request errors and exact direct-vector parity against frozen references (maximum difference 0; tolerance 1e-10). Routing selected 131 Jev, 676 Kev and 1,777 Typed cases. Every selected generation completed except two Kev attempts that reached the 8K context limit; those cases retained direct scoring. Partial text was not rescored or blended.
Comparison with E4B
The chart uses the same frozen cases, source labels and answer order. Shared settings are Q8 targets, F16 target KV shared with the pinned MTP assistant, 8K context, one chat slot and MTP4. Model-specific prompts, samplers, augmentation paths and adaptive policies differ. Winnow-12B's published campaigns remain separate and are not included in this matched comparison.
| Panel | E2B direct β adaptive reasoning | E2B change | E4B direct β adaptive reasoning | E4B change |
|---|---|---|---|---|
| JevBench, 231 | 175 β 202 | +11.69 pp | 183 β 203 | +8.66 pp |
| Kev v9 clean, 1,046 | 728 β 851 | +11.76 pp | 762 β 842 | +7.65 pp |
| Typed teacher agreement, 2,000 | 1,237 β 1,366 | +6.45 pp | 1,446 β 1,439 | β0.35 pp |
Changes are adaptive minus the same model's direct result, in percentage points. Chart percentages are computed from exact counts, with two decimal places. The card and graph both round percentages to two decimals.
E4B used the live e4b-calibrated75-g95-v1 path: raw maxP below 0.95, direct/augmented temperatures 1.2041180007310734 / 3.4209273427377678, and 25:75 direct/augmented weighting. Its brief reasoning prompt, released sampler, 75-second deadline, and batch 2048 differ from E2B's native-label prompt, backend temperature sampler, 180-second deadline and batch 1024. Both use ubatch 1024. F16 KV was outside E4B's measured release policy profile; calibration was held fixed, not revalidated here. The comparison does not isolate model size or the effect of reasoning alone.
E4B corrected/regressed 22/2 Jev answers, 114/34 Kev answers, and 129/136 Typed agreements. Its Typed NLL, Brier, and rating MAE also worsened slightly. The graph uses a downward cap for the seven-agreement loss. These fresh E4B counts differ from its older public Q8-KV comparisons; the historical release tables remain separate.
Latency
| Panel | Clean cases | Mean | Median | p95 | p99 |
|---|---|---|---|---|---|
| JevBench | 231 | 1.356 s | 0.887 s | 4.099 s | 6.057 s |
| Kev v9 clean | 1,046 | 1.212 s | 1.029 s | 3.522 s | 6.895 s |
| Typed | 1,984 | 1.951 s | 2.095 s | 3.005 s | 3.744 s |
These are serial adaptive reasoning pipeline times, including failed reasoning attempts. Sixteen Typed cases marked by the GPU-overlap evidence are excluded from timing only; all 2,000 remain in quality scores. Scored-case pipeline cost across all cases was 5,482.2 seconds. Startup, diagnostic replays, controller overhead and idle recovery gaps are outside those times. The full launch-to-cleanup span was 169.3 minutes and must not be presented as summed benchmark latency.
The earlier E4B live comparison used different output lengths, prompts and timing boundaries. It does not establish intrinsic relative speed, equal-output MTP speedup or concurrent service throughput. Native backend allocation figures are not continuously measured peak GPU VRAM.
Earlier paths
| Panel | Earlier numbered options | Historical F16/backend replay | Current adaptive reasoning path |
|---|---|---|---|
| JevBench, 231 | 195 | 203 | 202 |
| Kev v9 clean, 1,046 | 845 | 851 | 851 |
| Typed teacher agreement, 2,000 | 1,342 | 1,365 | 1,366 |
The current path restores most historical reasoning behavior but is not identical: it loses one Jev answer, matches Kev's winners, and gains one Typed agreement relative to the historical path. Native labels, safe escaping and state augmentation can change outcomes. These older paths are historical comparisons, not current release scores or independent calibration studies.
Disjoint holdout
An earlier separately frozen 288-case holdout contained 96 LogiQA2 choices, 96 PAWS Boolean decisions, and 96 HelpSteer2 consensus ratings. The 0.99/50% rule scored 158/288 versus 155/288 for the transferred 0.80/50% rule: nine fixes and six regressions. Its paired 95% change interval was β1.74 to +3.82 percentage points; NLL and Brier change intervals also included zero. Mean serial replay time rose from 1.066 to 1.677 seconds.
This holdout was disjoint from selection, but was not wholly independent external validation: 74 HelpSteer2 cases had previously been scored by other models, and public sources may have appeared in base-model pretraining. The small inconclusive gain does not establish a general benefit. It is not a holdout result for the current path.
Vision
This campaign is text only and adds no new vision or always-mode evidence. Vision evidence is limited to an earlier 180-decision procedural-image panel, with four context-limit fallbacks. The full text comparison above does not establish photo, screenshot, document, or image-conditioned reasoning quality.
Evidence identity
The machine-readable comparison contains exact chart counts and receipt hashes. The retained analysis, freeze, and completion receipt identify the evaluated source, artifacts, fallbacks and timing exclusions. Complete case logs remain in the private review evidence bundle, separate from the model package.
The final result log SHA-256 is db3d7cb81972ad2909a75aca659b97015d9e96a770f04620275b79b9fbdbd80d. The chart builder verifies the analysis, freeze and completion identities alongside the retained E4B receipt. Producing these figures performed no additional inference or fitting.