Warm-started Checkpoints Collection A collection of three models trained on the Nemotron Post Training Dataset for reasoning tasks with IVON ⢠4 items ⢠Updated 13 days ago
Parameter Exploration for RLVR via Variational Learning Paper ⢠2608.09805 ⢠Published 28 days ago ⢠7
Parameter Exploration for RLVR via Variational Learning Paper ⢠2608.09805 ⢠Published 28 days ago ⢠7
3PO Models Collection 3PO family methods trained on DapoMath-17k using Olmo3-IVON-SFT-7B and Qwen2.5Math-IVON-SFT-7B ⢠10 items ⢠Updated 25 days ago
Parameter Exploration for RLVR via Variational Learning Paper ⢠2608.09805 ⢠Published 28 days ago ⢠7