4096 to 8192

#1
by JNK333 - opened

this is a CoT model, and at lower quants other than Q8 it rambles alot more i dont think 4096 is the best, i hope a 8192 version can come out because 8192 gives the model enough space for very strong questions or multiple questions

image
it also doesnt work at all on my S24Ultra

LiteRT Community (FKA TFLite) org

Hi β€” just letting you know this model is now available to download in Box, an open-source, fully offline on-device AI app for Android (Apache-2.0, forked from Google's AI Edge Gallery). Box runs LiteRT / LiteRT-LM alongside llama.cpp, so .litertlm models run natively with GPU/NPU acceleration where the hardware supports it.

It appears in the in-app model browser with attribution and its original licence, and downloads directly from this repo β€” nothing is mirrored or re-hosted.

Thanks for publishing it. If you'd prefer it not be included, or want the description or credits changed, just say and I'll sort it.

https://github.com/jegly/Box

Sign up or log in to comment