Read It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image Generation
Paper • 2607.11886 • Published • 83
None defined yet.
MISA: Mixture of Indexer Sparse Attention for Long-Context LLM Inference
VideoZeroBench: Probing the Limits of Video MLLMs with Spatio-Temporal Evidence Verification