Request access to the BottleCap AI model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Request access below. Add your company if you're evaluating this for work — we have enterprise versions that go further, and we'll make sure you hear about them first. We'll also send you new ThinkingCap releases and early access before they're public.

Tell us how you plan to use the model and we can help you get the most out of it — there's a short form for that too.

Log in or Sign Up to review the conditions and access this model content.

ThinkingCap-Qwen3.8-27B-FP8

FP8 (e4m3) block-wise weights, 128x128 blocks, dynamic activation scales; loads in vLLM and SGLang, native on Hopper and Blackwell.

Built from bottlecapai/ThinkingCap-Qwen3.8-27B (bf16). Vision tower, MTP head, lm_head and the GDN in_proj_a / in_proj_b projections stay bf16 (quantization_config.ignore in config.json). Serve with vLLM 0.29 (--trust-remote-code not needed).

Expected performance

Paired comparison with the bf16 source on the full quantization plan: both builds answer the same questions with the same seeds — RealWorldQA 765 questions × 2 seeds (images), GPQA-Diamond 198 × 4, MMLU-Pro 1,500 × 1 (a fixed slice), IFBench 300 × 2, AA-LCR 100 × 1 (long-document prompts, graded by Gemma-4-26B-A4B-it with thinking off). Thinking at the chat template's default reasoning effort (xhigh), sampled decoding (temperature 1.0, top_p 0.95, top_k 20, min_p 0.0), 65,536-token generation cap. Both builds served by vLLM 0.29.0 on an H200.

benchmark (questions × seeds) accuracy %, bf16 → FP8 Δ accuracy, pp [95% CI] tokens mean / median / p95, bf16 → FP8 Δ mean tokens [95% CI] Δ median tokens
RealWorldQA (765 × 2) 83.1 → 83.6 +0.5 [−1.0, +2.0] 488 / 112 / 1,912 → 550 / 113 / 2,722 +12.8% [−4.3, +32.8] +0.9%
GPQA-Diamond (198 × 4) 88.0 → 87.8 −0.3 [−2.2, +1.7] 7,115 / 1,031 / 37,459 → 7,084 / 1,220 / 38,326 −0.4% [−8.3, +7.7] +18.4%
MMLU-Pro (1,500 × 1) 84.1 → 84.1 0.0 [−1.3, +1.3] 1,436 / 166 / 7,255 → 1,323 / 176 / 6,977 −7.9% [−19.2, +6.0] +6.0%
IFBench (300 × 2) 79.7 → 79.3 −0.3 [−3.6, +2.9] 4,531 / 1,822 / 20,630 → 4,283 / 1,644 / 18,008 −5.5% [−13.1, +3.0] −9.8%
AA-LCR (100 × 1) 81.0 → 82.0 +1.0 [−5.5, +7.5] 1,718 / 844 / 4,937 → 1,506 / 855 / 4,593 −12.4% [−30.8, +10.0] +1.2%

Δ accuracy is FP8 minus bf16 on the same answers; its interval treats the question as the unit (seeds averaged per question first). Tokens are completion tokens (reasoning plus answer).

Throughput — one RTX PRO 6000 Blackwell, vLLM 0.29.0, synthetic prompts of 1,024 tokens with 512 generated (last column: 32,768-token prompts, 128 generated), end-of-sequence ignored, prefix caching off, --max-num-seqs 64. Aggregate output tokens/s, median time to first token in ms in parentheses.

build 1 request 16 concurrent 64 concurrent 4 concurrent, 32k-token prompts
bf16 26.2 (158) 341 (1,867) 844 (3,627) 18.6 (9,678)
FP8 44.7 (107) 540 (1,227) 1,227 (2,369) 27.0 (7,668)

Decode speed and MTP self-speculative decoding (MMLU-Pro) — 32 questions × 1 seed, vLLM 0.26 on one H200, 16 concurrent requests

config median tokens tok/s s / task MTP speedup accept_len (max 4)
Qwen3.8-27B base · standard 484 51.3 8.9 1.00× —
Qwen3.8-27B base · MTP 502 91.0 4.2 1.77× 2.59
ThinkingCap-Qwen3.8-27B bf16 · standard 232 50.5 3.7 1.00× —
ThinkingCap-Qwen3.8-27B bf16 · MTP 216 85.7 2.0 1.70× 2.60
FP8 · standard 216 62.9 2.8 1.00× —
FP8 · MTP 252 112.1 2.3 1.78× 2.55

Where to find us

Website LinkedIn Instagram X

Need even more efficiency? The open release is production-ready. Our enterprise versions go further — fewer thinking tokens still, tuned to your workload, at matched accuracy on your own tasks. Built for AI labs, inference providers and enterprises running models at scale. Deployed on your infrastructure, or in the cloud and region you choose. Talk to our team

License

ThinkingCap: PolyForm Small Business 1.0.0 + BottleCap personal-use grant (see LICENSE).

Upstream Qwen materials: Apache-2.0 (see NOTICE).

Commercial license: contact BottleCap AI.

Citation

If you use this model, please cite:

@misc{ThinkingCap-Qwen3.8-27B,
  title     = {bottlecapai/ThinkingCap-Qwen3.8-27B},
  author    = {Osusky, Adam and Lindauer, Jan and Jirkovsky, Adam and Mihal, Filip and Platek, Ondrej and Herel, David and Ihnatchenko, Luka and Bartek, Vojtech and Jirak, Jiri and Kubista, Daniel and Krus, Frantisek and Mikolov, Tomas},
  year      = {2026},
}
Downloads last month
2,513
Safetensors
Model size
28B params
Tensor type
BF16
·
F8_E4M3
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for bottlecapai/ThinkingCap-Qwen3.8-27B-FP8

Base model

Qwen/Qwen3.8-27B
Quantized
(33)
this model

Spaces using bottlecapai/ThinkingCap-Qwen3.8-27B-FP8 2

Collection including bottlecapai/ThinkingCap-Qwen3.8-27B-FP8