Same bytes, closer to the original: two lines of AutoRound we had wrong FINAL-Bench • about 19 hours ago • 12
Trained 210M text-to-image model from scratch on one GPU: what actually mattered ivanmikhnenkov • 5 days ago • 6
A Guide to Reinforcement Learning Post-Training for LLMs: PPO, DPO, GRPO, and Beyond karina-zadorozhny • Jan 19 • 48
Open Yap 1K: 1,000 hours of full-duplex natural conversation, free for commercial use TheAgenticDataCompany • 11 days ago • 13
Same bytes, closer to the original: two lines of AutoRound we had wrong FINAL-Bench • about 19 hours ago • 12
Trained 210M text-to-image model from scratch on one GPU: what actually mattered ivanmikhnenkov • 5 days ago • 6
A Guide to Reinforcement Learning Post-Training for LLMs: PPO, DPO, GRPO, and Beyond karina-zadorozhny • Jan 19 • 48
Open Yap 1K: 1,000 hours of full-duplex natural conversation, free for commercial use TheAgenticDataCompany • 11 days ago • 13