RealCompanion: Benchmarking Human Understanding from Reasoning over Longitudinal Real-World Conversations Paper • 2610.01780 • Published 5 days ago • 136
4DCodeBench: Benchmarking Agents on Inverse Graphics of Dynamic Scenes Paper • 2610.03715 • Published 4 days ago • 9
LVMT: Video Mask Transformer for Long-term Video Segmentation Paper • 2609.34895 • Published 7 days ago • 11
Scaling Properties of Same-Family On-Policy Distillation Paper • 2609.32722 • Published 10 days ago • 322
PrismQuant: Optimal Null-Space Rotations for Grouped Quantizers Paper • 2609.32429 • Published 10 days ago • 5
AutoDataBench: Can Agents Write the Data That Feeds the Self-Improvement Loop? Paper • 2609.35025 • Published 8 days ago • 7
WISE-ATTA: When to Ask for Labels in Budgeted Active Test-Time Adaptation Paper • 2609.37687 • Published 7 days ago • 6