ToolArtist: Tool-Using Unified Multimodal Models for Agentic Image Generation Paper • 2608.04436 • Published 3 days ago • 51
MPIE-Bench: Benchmarking Anatomically Plausible Multi-Person Interaction Editing Paper • 2607.27616 • Published 9 days ago • 38
VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System Paper • 2607.27380 • Published 10 days ago • 70
Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention Sparsification Paper • 2607.24027 • Published 12 days ago • 37
JarvisHub: An Open Harness for Canvas-Native Multimodal Creative Agents Paper • 2607.23588 • Published 13 days ago • 125
From Proprietary to Open-Source: Bridging the Distribution Gap via Multi-Agent Protocol Distillation in Agentic Search Paper • 2607.24280 • Published 12 days ago • 82
Streaming Multi-Agent Autoregressive Diffusion Model with World State Registers Paper • 2607.21594 • Published 16 days ago • 16
SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation Paper • 2607.21553 • Published 16 days ago • 39
Apple-π: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence Paper • 2607.16401 • Published 22 days ago • 44
DeepSearch-World: Self-Distillation for Deep Search Agents in a Verifiable Environment Paper • 2607.07820 • Published about 1 month ago • 92
HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enchancement Paper • 2607.18217 • Published 19 days ago • 61
RESOURCE2SKILL: Distilling Executable Agent Skills from Human-Created Multimodal Resources Paper • 2606.29538 • Published 23 days ago • 142