Adaptive Reward Routing: Dynamic Multi-Reward Optimization for Joint Audio-Video Diffusion via Forward-Process RL Paper • 2609.37200 • Published 8 days ago • 138
Just MLPs: Efficient Visual State Reconstruction for Multimodal Language Models Paper • 2609.34972 • Published 9 days ago • 30
Think Before You Score: Thinking Reward Model for Visual Generation Paper • 2609.37372 • Published 8 days ago • 101
RewardVerse: Rubric-Guided Policy Optimization for Video Reward Modeling Paper • 2609.22947 • Published 18 days ago • 43
All modalities are equal, but video is more equal: Closing the Cross-Attention Gap in Joint Video Generation Paper • 2609.27901 • Published 14 days ago • 24
SpeakerMem-R1: Speaker-Centered Dual-Track Memory for Multi-Party Dialogue Paper • 2609.26780 • Published 15 days ago • 103
🎲 Decision 1.0 Collection Towards Open Decision Foundation Models • 7 items • Updated 6 days ago • 23
onPanda: Efficient Annotation of On-Policy Alignment Data for LLMs and Agents via Token-Level Correction Paper • 2609.24983 • Published 16 days ago • 55
MoME: Mixture-of-Memory Embeddings for Context-Aware Sparse Lookup Paper • 2609.15126 • Published 23 days ago • 12
VideoGen-Agent: Reinforcing Video Generation Agents Paper • 2609.24997 • Published 16 days ago • 73
WorldCrafter: Consistent Video World Model with Implicit 3D-aware Memory Paper • 2609.24984 • Published 16 days ago • 157
RecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-Use Agents Paper • 2609.22000 • Published 19 days ago • 79
Designer-RSI: Evolving Procedural Memory from User Traffic for Agentic Graphic Design Paper • 2609.22086 • Published 19 days ago • 34
Transferring the Intelligence of VLMs to Robotic Control Paper • 2609.22966 • Published 18 days ago • 120
ActionPiece: Rethinking Action Tokenization for Autoregressive Vision-Language-Action Models Paper • 2609.18487 • Published 21 days ago • 48
SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness Paper • 2609.20519 • Published 20 days ago • 139