ManniX PRO

ManniX-ITA

32 6 28

https://github.com/mann1x

mann1x

AI & ML interests

None yet

Recent Activity

new activity 15 days ago

ManniX-ITA/Qwen3.6-27B-Omnimerge-v4-MTP-GGUF:My experience

new activity 16 days ago

ManniX-ITA/gemma-4-A4B-98e-v7-coder-it-GGUF:Gets stuck in loops using llama.cpp

posted an update 17 days ago

--- 🚀 Gemma-4-A4B 98e v7-coder cohort — loop-fixed re-release. Two 20.8B MoE coders (4B-active), fresh-map prunes of Gemma 4 26B-A4B, 30/128 experts dropped per layer. The headline isn't a benchmark: the agentic loop is gone at the weights, not papered over by the sampler. 🔧 How: at prune time we force-keep the 46 agentic_eog experts a loop-protection signal flags as load-bearing for clean multi-turn termination (+ shared-FFN α=1.2). Result: 0 loops across 48 seeds on every published tier. 📊 Q6_K · llama.cpp · greedy · same host (from summary.json): ⚖️ v7-coder (fkbroad code3/lcb2) — balanced coder: LCB-med-55 98.18, HumanEval 98.17, HE+ 92.07, AIME 80.0, MATH-500 95.0, GSM8K 91, IFEval 92, MultiPL-E 89.7, ARC 92.2. ⚡ v7-coderx (code4/lcb3) — code-maximal: all-hard LCB-77 85.71 (cohort-best; 128e 79.22, v7-coder 84.42), HE+ 93.29, GSM8K 93, MATH-500 95.0, AIME 76.67. Whole budget on code. 🎯 Both land near GPQA ~51 — graduate science is the budget axis, neither is a science model. Pick v7-coder for the broad LCB-medium + HumanEval lead; v7-coderx for the all-hard slice and HE+. 🧪 The harness we used to prove the fix is now an omk tool: agentic-loop-harness replays a frozen agentic conversation across a sampler×seed matrix and reports a fail-rate per chat-template, so you can isolate a loop to one variable. Model-agnostic — any OpenAI-compatible server. The version we shared with Google: https://huggingface.co/google/gemma-4-12B-it/discussions/41#6a3926720abc934d03fd85c0 📦 Each ships bf16 · GGUF (+ CD-* + imatrix + mmproj vision) · NVFP4A16 (~13 GB) · Ollama. 🔗 https://huggingface.co/ManniX-ITA/gemma-4-A4B-98e-v7-coder-it (+ -it-GGUF, -NVFP4A16) · https://ollama.com/mannix/gemma4-98e-v7-coder 🔗 https://huggingface.co/ManniX-ITA/gemma-4-A4B-98e-v7-coderx-it (+ -it-GGUF, -NVFP4A16) · https://ollama.com/mannix/gemma4-98e-v7-coderx 🔧 https://github.com/mann1x/omnimergekit/tree/main/tools/agentic-loop-harness

View all activity

Organizations

None yet

New activity in ManniX-ITA/Qwen3.6-27B-Omnimerge-v4-MTP-GGUF 15 days ago

My experience

👍 2

#1 opened about 1 month ago by

tooltd

New activity in ManniX-ITA/gemma-4-A4B-98e-v7-coder-it-GGUF 16 days ago

Gets stuck in loops using llama.cpp

#1 opened about 1 month ago by

DPS-900

posted an update 17 days ago

Post

223

---
🚀 Gemma-4-A4B 98e v7-coder cohort — loop-fixed re-release. Two 20.8B MoE coders (4B-active), fresh-map prunes of Gemma 4 26B-A4B, 30/128 experts dropped per layer. The headline isn't a benchmark: the agentic loop is
gone at the weights, not papered over by the sampler.

🔧 How: at prune time we force-keep the 46 agentic_eog experts a loop-protection signal flags as load-bearing for clean multi-turn termination (+ shared-FFN α=1.2). Result: 0 loops across 48 seeds on every published
tier.

📊 Q6_K · llama.cpp · greedy · same host (from summary.json):

⚖️ v7-coder (fkbroad code3/lcb2) — balanced coder: LCB-med-55 98.18, HumanEval 98.17, HE+ 92.07, AIME 80.0, MATH-500 95.0, GSM8K 91, IFEval 92, MultiPL-E 89.7, ARC 92.2.

⚡ v7-coderx (code4/lcb3) — code-maximal: all-hard LCB-77 85.71 (cohort-best; 128e 79.22, v7-coder 84.42), HE+ 93.29, GSM8K 93, MATH-500 95.0, AIME 76.67. Whole budget on code.

🎯 Both land near GPQA ~51 — graduate science is the budget axis, neither is a science model. Pick v7-coder for the broad LCB-medium + HumanEval lead; v7-coderx for the all-hard slice and HE+.

🧪 The harness we used to prove the fix is now an omk tool: agentic-loop-harness replays a frozen agentic conversation across a sampler×seed matrix and reports a fail-rate per chat-template, so you can isolate a loop
to one variable. Model-agnostic — any OpenAI-compatible server. The version we shared with Google: google/gemma-4-12B-it#41

📦 Each ships bf16 · GGUF (+ CD-* + imatrix + mmproj vision) · NVFP4A16 (~13 GB) · Ollama.
🔗 ManniX-ITA/gemma-4-A4B-98e-v7-coder-it (+ -it-GGUF, -NVFP4A16) · https://ollama.com/mannix/gemma4-98e-v7-coder
🔗 ManniX-ITA/gemma-4-A4B-98e-v7-coderx-it (+ -it-GGUF, -NVFP4A16) · https://ollama.com/mannix/gemma4-98e-v7-coderx
🔧 https://github.com/mann1x/omnimergekit/tree/main/tools/agentic-loop-harness

updated 6 models 18 days ago

Repeated pauses in OpenCode

#2 opened 29 days ago by

mohkamfer

New activity in google/gemma-4-12B-it 19 days ago

gemma-4-12B-it: deterministic "thought\n thought\n …" degenerate loop on long agent prompts (~60% reproducer, 4-bit, repros across temperatures)

#41 opened 23 days ago by

Raullen

New activity in google/gemma-4-12B-it about 1 month ago

Chat template may re-inject prior-turn reasoning during multi-turn tool use → repetition loops

#38 opened about 1 month ago by

ManniX-ITA

updated 2 models about 1 month ago

ManniX-ITA/gemma-4-A4B-98e-v6-coder-it-GGUF

20B • Updated Jun 10 • 5.75k • 4

ManniX-ITA/gemma-4-A4B-98e-v6-coder-it

20B • Updated Jun 10 • 12 • 1

New activity in bartowski/google_gemma-4-26B-A4B-it-GGUF about 1 month ago

Chat template may re-inject prior-turn reasoning during multi-turn tool use → repetition loops

#6 opened about 1 month ago by

ManniX-ITA

New activity in unsloth/gemma-4-26B-A4B-it-GGUF about 1 month ago

Chat template may re-inject prior-turn reasoning during multi-turn tool use → repetition loops

#43 opened about 1 month ago by

ManniX-ITA

New activity in unsloth/gemma-4-26B-A4B-it about 1 month ago

Chat template may re-inject prior-turn reasoning during multi-turn tool use → repetition loops

#2 opened about 1 month ago by

ManniX-ITA

New activity in google/gemma-4-E2B-it about 1 month ago

Chat template may re-inject prior-turn reasoning during multi-turn tool use → repetition loops

#36 opened about 1 month ago by

ManniX-ITA

New activity in google/gemma-4-E4B-it about 1 month ago

Chat template may re-inject prior-turn reasoning during multi-turn tool use → repetition loops

#37 opened about 1 month ago by

ManniX-ITA

New activity in google/gemma-4-31B-it about 1 month ago

Chat template may re-inject prior-turn reasoning during multi-turn tool use → repetition loops

#119 opened about 1 month ago by

ManniX-ITA

ManniX PRO

AI & ML interests

Recent Activity

Organizations

ManniX-ITA's activity

My experience

Gets stuck in loops using llama.cpp

Repeated pauses in OpenCode

gemma-4-12B-it: deterministic "thought\n thought\n …" degenerate loop on long agent prompts (~60% reproducer, 4-bit, repros across temperatures)

Chat template may re-inject prior-turn reasoning during multi-turn tool use → repetition loops

Chat template may re-inject prior-turn reasoning during multi-turn tool use → repetition loops

Chat template may re-inject prior-turn reasoning during multi-turn tool use → repetition loops

Chat template may re-inject prior-turn reasoning during multi-turn tool use → repetition loops

Chat template may re-inject prior-turn reasoning during multi-turn tool use → repetition loops

Chat template may re-inject prior-turn reasoning during multi-turn tool use → repetition loops

Chat template may re-inject prior-turn reasoning during multi-turn tool use → repetition loops