Request access to the BottleCap AI model
This repository is publicly accessible, but you have to accept the conditions to access its files and content.
Request access below. Add your company if you're evaluating this for work — we have enterprise versions that go further, and we'll make sure you hear about them first. We'll also send you new ThinkingCap releases and early access before they're public.
Tell us how you plan to use the model and we can help you get the most out of it — there's a short form for that too.
Log in or Sign Up to review the conditions and access this model content.
ThinkingCap-Qwen3.8-27B-FP8
FP8 (e4m3) block-wise weights, 128x128 blocks, dynamic activation scales; loads in vLLM and SGLang, native on Hopper and Blackwell.
Built from bottlecapai/ThinkingCap-Qwen3.8-27B (bf16).
Vision tower, MTP head, lm_head and the GDN in_proj_a / in_proj_b projections stay bf16 (quantization_config.ignore in config.json). Serve with vLLM 0.29 (--trust-remote-code not needed).
Expected performance
Paired comparison with the bf16 source on the full quantization plan: both builds answer the same questions with the same seeds — RealWorldQA 765 questions × 2 seeds (images), GPQA-Diamond 198 × 4, MMLU-Pro 1,500 × 1 (a fixed slice), IFBench 300 × 2, AA-LCR 100 × 1 (long-document prompts, graded by Gemma-4-26B-A4B-it with thinking off). Thinking at the chat template's default reasoning effort (xhigh), sampled decoding (temperature 1.0, top_p 0.95, top_k 20, min_p 0.0), 65,536-token generation cap. Both builds served by vLLM 0.29.0 on an H200.
| benchmark (questions × seeds) | accuracy %, bf16 → FP8 | Δ accuracy, pp [95% CI] | tokens mean / median / p95, bf16 → FP8 | Δ mean tokens [95% CI] | Δ median tokens |
|---|---|---|---|---|---|
| RealWorldQA (765 × 2) | 83.1 → 83.6 | +0.5 [−1.0, +2.0] | 488 / 112 / 1,912 → 550 / 113 / 2,722 | +12.8% [−4.3, +32.8] | +0.9% |
| GPQA-Diamond (198 × 4) | 88.0 → 87.8 | −0.3 [−2.2, +1.7] | 7,115 / 1,031 / 37,459 → 7,084 / 1,220 / 38,326 | −0.4% [−8.3, +7.7] | +18.4% |
| MMLU-Pro (1,500 × 1) | 84.1 → 84.1 | 0.0 [−1.3, +1.3] | 1,436 / 166 / 7,255 → 1,323 / 176 / 6,977 | −7.9% [−19.2, +6.0] | +6.0% |
| IFBench (300 × 2) | 79.7 → 79.3 | −0.3 [−3.6, +2.9] | 4,531 / 1,822 / 20,630 → 4,283 / 1,644 / 18,008 | −5.5% [−13.1, +3.0] | −9.8% |
| AA-LCR (100 × 1) | 81.0 → 82.0 | +1.0 [−5.5, +7.5] | 1,718 / 844 / 4,937 → 1,506 / 855 / 4,593 | −12.4% [−30.8, +10.0] | +1.2% |
Δ accuracy is FP8 minus bf16 on the same answers; its interval treats the question as the unit (seeds averaged per question first). Tokens are completion tokens (reasoning plus answer).
Throughput — one RTX PRO 6000 Blackwell, vLLM 0.29.0, synthetic prompts of 1,024 tokens with 512 generated (last column: 32,768-token prompts, 128 generated), end-of-sequence ignored, prefix caching off, --max-num-seqs 64. Aggregate output tokens/s, median time to first token in ms in parentheses.
| build | 1 request | 16 concurrent | 64 concurrent | 4 concurrent, 32k-token prompts |
|---|---|---|---|---|
| bf16 | 26.2 (158) | 341 (1,867) | 844 (3,627) | 18.6 (9,678) |
| FP8 | 44.7 (107) | 540 (1,227) | 1,227 (2,369) | 27.0 (7,668) |
Decode speed and MTP self-speculative decoding (MMLU-Pro) — 32 questions × 1 seed, vLLM 0.26 on one H200, 16 concurrent requests
| config | median tokens | tok/s | s / task | MTP speedup | accept_len (max 4) |
|---|---|---|---|---|---|
| Qwen3.8-27B base · standard | 484 | 51.3 | 8.9 | 1.00× | — |
| Qwen3.8-27B base · MTP | 502 | 91.0 | 4.2 | 1.77× | 2.59 |
| ThinkingCap-Qwen3.8-27B bf16 · standard | 232 | 50.5 | 3.7 | 1.00× | — |
| ThinkingCap-Qwen3.8-27B bf16 · MTP | 216 | 85.7 | 2.0 | 1.70× | 2.60 |
| FP8 · standard | 216 | 62.9 | 2.8 | 1.00× | — |
| FP8 · MTP | 252 | 112.1 | 2.3 | 1.78× | 2.55 |
Where to find us
Need even more efficiency? The open release is production-ready. Our enterprise versions go further — fewer thinking tokens still, tuned to your workload, at matched accuracy on your own tasks. Built for AI labs, inference providers and enterprises running models at scale. Deployed on your infrastructure, or in the cloud and region you choose. Talk to our team
License
ThinkingCap: PolyForm Small Business 1.0.0 + BottleCap personal-use grant (see LICENSE).
Upstream Qwen materials: Apache-2.0 (see NOTICE).
Commercial license: contact BottleCap AI.
Citation
If you use this model, please cite:
@misc{ThinkingCap-Qwen3.8-27B,
title = {bottlecapai/ThinkingCap-Qwen3.8-27B},
author = {Osusky, Adam and Lindauer, Jan and Jirkovsky, Adam and Mihal, Filip and Platek, Ondrej and Herel, David and Ihnatchenko, Luka and Bartek, Vojtech and Jirak, Jiri and Kubista, Daniel and Krus, Frantisek and Mikolov, Tomas},
year = {2026},
}
- Downloads last month
- 2,513