Request access to the BottleCap AI model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Please fill out the form below to request access. Name, Company Name, and Company Email are optional — just enter NA if you'd prefer not to share. Your email may be used to send you information about BottleCap AI model updates and early access before public release.

Log in or Sign Up to review the conditions and access this model content.

ThinkingCap — BottleCap AI

bottlecapai/ThinkingCap-Qwen3.6-27B-GGUF

GGUF / llama.cpp quantizations of bottlecapai/ThinkingCap-Qwen3.6-27B — capability of Qwen3.6-27B with 50% less thinking tokens on average, achieved by finetuning Qwen3.6-27B (Qwen Team, 2026) while preserving the original answer quality and style.

➡️ Full model description, evaluation results (multi-seed, statistically tested), recommended sampling params, and citation: see the main model card at bottlecapai/ThinkingCap-Qwen3.6-27B.

About GGUF and quantization

GGUF is a single-file model format for running LLMs locally with llama.cpp and compatible runtimes (Ollama, LM Studio, …). The quantized variants below store weights at reduced precision — e.g. ≈4.7 bits per weight for Q4_K_M instead of the 16-bit f16 source — cutting download size and memory severalfold at a small, measured quality cost.

Files

File Quant Size
ThinkingCap-Qwen3.6-27B-Q4_K_M.gguf Q4_K_M 15.7 GB
ThinkingCap-Qwen3.6-27B-Q6_K.gguf Q6_K 20.9 GB
ThinkingCap-Qwen3.6-27B-Q8_0.gguf Q8_0 27.1 GB
ThinkingCap-Qwen3.6-27B-f16.gguf f16 50.9 GB
mmproj-ThinkingCap-Qwen3.6-27B-f16.gguf mmproj (vision) 0.9 GB

f16 is the unquantized source; Q8_0 is near-lossless; Q6_K is a slightly smaller near-lossless option; Q4_K_M is the recommended size/quality balance for most local setups.

Usage (llama.cpp)

# pull a specific quant straight from the Hub and chat
llama-cli -hf bottlecapai/ThinkingCap-Qwen3.6-27B-GGUF:Q4_K_M -p "Hi"

# or download one file and run it
huggingface-cli download bottlecapai/ThinkingCap-Qwen3.6-27B-GGUF ThinkingCap-Qwen3.6-27B-Q4_K_M.gguf --local-dir .
llama-cli -m ThinkingCap-Qwen3.6-27B-Q4_K_M.gguf -p "Hi"

Speculative decoding (MTP)

llama.cpp can run MTP (multi-token-prediction) self-speculative decoding on these GGUFs for a decode speed-up — no separate draft model needed. Add --spec-type draft-mtp when serving:

llama-server -hf bottlecapai/ThinkingCap-Qwen3.6-27B-GGUF:Q4_K_M --spec-type draft-mtp

Set the draft length with --spec-draft-n-max (e.g. 4). Requires a recent llama.cpp build with MTP support.

Vision (image input)

ThinkingCap is a vision-language model. Image input needs the multimodal projector mmproj-ThinkingCap-Qwen3.6-27B-f16.gguf (in this repo) loaded alongside a text GGUF — the single f16 mmproj pairs with any of the quants above.

  • LM Studio / Jan / Ollama, …: download the mmproj-*.gguf from this repo; LM Studio auto-detects it and enables the image (🖼️) button.
  • llama.cpp CLI:
huggingface-cli download bottlecapai/ThinkingCap-Qwen3.6-27B-GGUF \
  ThinkingCap-Qwen3.6-27B-Q4_K_M.gguf mmproj-ThinkingCap-Qwen3.6-27B-f16.gguf --local-dir .
llama-mtmd-cli -m ThinkingCap-Qwen3.6-27B-Q4_K_M.gguf \
  --mmproj mmproj-ThinkingCap-Qwen3.6-27B-f16.gguf --image photo.jpg -p "Describe this image."
  • llama-server: add --mmproj mmproj-ThinkingCap-Qwen3.6-27B-f16.gguf to expose an OpenAI-compatible vision endpoint.

Expected performance

Measured on our serving harness (llama.cpp, CUDA llama-server) over MMLU-Pro (reasoning) and RealWorldQA (vision), N=200/dataset × 3 seeds, sampled decoding (temp 1.0 / top_p 0.95 / top_k 20). acc is the mean ± 95% CI across seeds; task s / tok/s are measured per request during the eval at batch size 16 (a throughput setting — a single local user decoding one request at a time will see meaningfully higher tok/s, but the per-row comparison is unaffected). For the full multi-benchmark evaluation see the main model card.

Accuracy is statistically identical across f16, the quantized variants, and the base model — every 95% CI overlaps (MMLU-Pro ≈0.89–0.91, RealWorldQA ≈0.78–0.82). Quantization is near-lossless, and the finetune matches base accuracy. What these GGUFs buy you is brevity → speed: the finetune reasons ~2× shorter than base (≈1000 vs ≈2100 MMLU-Pro tokens; ≈330 vs ≈700 RealWorldQA), so at equal accuracy it finishes each task far faster. MTP self-speculative decoding (--spec-type draft-mtp, n=4) accepts ≈3.3–3.7 tokens per verify step on top. The bolded row — Q4_K_M + MTP — is the fastest per task and the recommended size/quality balance; Q8_0 and Q6_K are near-lossless alternatives at larger sizes. For reference we also list unsloth's Dynamic GGUFs of the base model (UD-*): same llama.cpp path and same accuracy, but base-model quants reason ≈2× longer (none of the finetune's token savings), so they finish only ~1.2–2.2× faster than base-std vs our ~2.6–3.9×.

median tokens = median completion length; task s = measured end-to-end wall-clock per request (bs=16); tok/s = per-request decode rate under that load; speedup = base·bf16·std task s ÷ row task s — same llama.cpp path for every row, so the within-table comparison is apples-to-apples.

MMLU-Pro (reasoning)

config acc (mean ± 95% CI) median tokens tok/s task s speedup accept_len (n=4)
Qwen3.6-27B base (bf16 GGUF) · standard 0.908 ± 0.007 2098 16.0 138.1 1.00×
f16 · standard 0.908 ± 0.038 1027 16.0 67.0 2.06×
f16 · MTP 0.903 ± 0.007 1045 25.9 43.3 3.19× 3.69
Q8_0 · standard 0.897 ± 0.044 990 21.0 49.6 2.78×
Q8_0 · MTP 0.900 ± 0.012 1043 31.3 36.7 3.77× 3.69
Q6_K · standard 0.908 ± 0.007 1011 22.1 49.2 2.81×
Q6_K · MTP 0.910 ± 0.012 992 30.4 36.6 3.77× 3.69
Q4_K_M · standard 0.903 ± 0.019 996 25.5 42.3 3.27×
Q4_K_M · MTP 0.910 ± 0.012 1054 32.8 35.5 3.89× 3.65
unsloth UD-Q8_K_XL (base) · standard 0.892 ± 0.014 2139 19.6 112.9 1.22×
unsloth UD-Q8_K_XL (base) · MTP 0.905 ± 0.000 2133 31.2 72.3 1.91× 3.66
unsloth UD-Q4_K_XL (base) · standard 0.900 ± 0.025 2172 26.0 86.3 1.60×
unsloth UD-Q4_K_XL (base) · MTP 0.903 ± 0.014 2094 35.4 63.4 2.18× 3.61

RealWorldQA (vision)

config acc (mean ± 95% CI) median tokens tok/s task s speedup accept_len (n=4)
Qwen3.6-27B base (bf16 GGUF) · standard 0.797 ± 0.019 700 15.1 50.9 1.00×
f16 · standard 0.815 ± 0.045 340 13.4 29.3 1.74×
f16 · MTP 0.818 ± 0.031 331 16.1 26.6 1.92× 3.29
Q8_0 · standard 0.815 ± 0.054 318 16.5 23.8 2.14×
Q8_0 · MTP 0.797 ± 0.019 335 21.2 21.0 2.42× 3.32
Q6_K · standard 0.813 ± 0.007 359 17.0 25.0 2.04×
Q6_K · MTP 0.817 ± 0.031 355 20.4 23.3 2.19× 3.32
Q4_K_M · standard 0.808 ± 0.014 334 19.2 21.7 2.35×
Q4_K_M · MTP 0.810 ± 0.045 315 21.3 19.8 2.57× 3.26
unsloth UD-Q8_K_XL (base) · standard 0.795 ± 0.000 662 18.2 45.0 1.13×
unsloth UD-Q8_K_XL (base) · MTP 0.782 ± 0.073 672 25.6 35.5 1.44× 3.43
unsloth UD-Q4_K_XL (base) · standard 0.792 ± 0.031 727 23.5 40.2 1.27×
unsloth UD-Q4_K_XL (base) · MTP 0.785 ± 0.081 760 29.1 35.3 1.44× 3.38
Downloads last month
32,289
GGUF
Model size
27B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

4-bit

6-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for bottlecapai/ThinkingCap-Qwen3.6-27B-GGUF

Base model

Qwen/Qwen3.6-27B
Quantized
(51)
this model