Text Generation
NInfer
qwen.83
qwen3_8
nvfp4
blackwell
mtp
speculative-decoding
vocabulary-prune
conversational
Instructions to use Schestex/ThinkingCap-Qwen3.8-27B-NInfer with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- NInfer
How to use Schestex/ThinkingCap-Qwen3.8-27B-NInfer with NInfer:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
Reducing the vocabulary to free up more VRAM and potentially improve decoding
#3
by Schestex - opened
... sooon ๐
The first pruning was successful and is working like a charm.
The second and final pruning was successful too and is working like a charm.
The next step will be to lower the voca proposal slightly and find the exact settings for it in the Ninfer kernel! Possibly including a scaling range so that, by selecting a range, you always have the appropriate working range without needing future changes.