SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation Paper • 2608.17426 • Published 4 days ago • 155
Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination Paper • 2608.14391 • Published 8 days ago • 277
AVA-Encoder: Towards Agent-Native Video Representation Learning Paper • 2608.12313 • Published 10 days ago • 40
OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution Paper • 2608.00677 • Published 21 days ago • 261
PaDoc: Layout-Grounded Parallel Decoding for Document Parsing Paper • 2608.06146 • Published 16 days ago • 24
When Memory Lies: An Empirical Study of Spatial Memory Staleness in VLM Agents Paper • 2608.04574 • Published 17 days ago • 16
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Paper • 2608.02023 • Published 19 days ago • 156
DeepVoyager-VL: Incentivizing Vision-in-the-Loop Search for Long-Horizon Multimodal Agents Paper • 2608.01827 • Published 19 days ago • 18