File size: 2,083 Bytes
fdaebc9
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
# Runtime profiles and compatibility

These named presets were tested on Linux/CUDA with RTX5070Ti16GB. Observed memory is profile-specific; it is not a universal GPU-fit guarantee.

| Tested preset | Context | Batch / microbatch | Vision | Observed peak device memory |
|---|---:|---|---|---:|
| `12b-nvfp4-vision8k-mtp` |8,192 |2,048 /1,024 |yes |11,773 MiB |
| `12b-q8-text8k-mtp` |8,192 |512 /256 |no |14,927–15,108 MiB |

All use q8_0 target/draft KV, four native branches, one chat slot, full GPU residency, AUTO memory and MTP depth4. These are measured defaults. The unified presets `q8`, `nv4` and `e4b` accept context, cache, batch and native-branch overrides. MTP requires one chat slot and auto memory. Custom settings do not inherit the measured calibration or performance claims; `winnow presets` lists advisory memory estimates. The recorded build used CUDA13.3/SM120. Install the server prerequisites and select the appropriate supported build settings.

## Supported combinations

Q8 vision plus MTP is available with explicit flags, but its measured 8K configuration exceeded 16 GB. Use Q8 MTP through the text preset, NVFP4 for the tested combined vision+MTP preset, or direct Q8 vision without MTP. A larger-memory Q8 combined configuration has not been verified. BF16 has no validated MTP/adaptive preset.

MTP and vision require additional VRAM. Quantization, context, batch and concurrency change memory use. Larger context/concurrency or different hardware is not established by these measurements. Direct vision and adaptive text serving are separate profiles; stop one before starting the other. Adaptive input supports one named question and a text state, with no image calibration claim.

## Measurement scope

Memory peaks are sampled device usage and may miss transients. They are not sustained-capacity measurements. These historical measurements retain their original runtime provenance; later compact parity checks are not new broad memory/quality evaluations. Timing methods and direct64Kvision measurements are in [BENCHMARKS.md](BENCHMARKS.md).