Reducing the vocabulary to free up more VRAM and potentially improve decoding

#3
by Schestex - opened

... sooon ๐Ÿ˜€

The first pruning was successful and is working like a charm.

The second and final pruning was successful too and is working like a charm.

The next step will be to lower the voca proposal slightly and find the exact settings for it in the Ninfer kernel! Possibly including a scaling range so that, by selecting a range, you always have the appropriate working range without needing future changes.

Proposal target @ the moment 1126400

Workload Kernel A Kernel B Diff
mixed64 116.315 116.395 +0.069 %
mixed110 107.745 107.830 +0.079 %

no change so far for:

  • Acceptance
  • Rounds
  • Fallbacks

Proposal target @ the moment 1126400

Workload Kernel A Kernel B Diff
mixed64 116.315 116.395 +0.069 %
mixed110 107.745 107.830 +0.079 %

no change so far for:

  • Acceptance
  • Rounds
  • Fallbacks

Sign up or log in to comment