AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution Traces Paper • 2608.23041 • Published 17 days ago • 64
AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution Traces Paper • 2608.23041 • Published 17 days ago • 64
Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents Paper • 2510.02837 • Published May 25
Personalize-then-Store: Benchmarking and Learning Personalized Memory for Long-horizon Agents Paper • 2605.25535 • Published May 25 • 46
Rethinking Failure Attribution in Multi-Agent Systems: A Multi-Perspective Benchmark and Evaluation Paper • 2603.25001 • Published Mar 26
Self-EvolveRec: Self-Evolving Recommender Systems with LLM-based Directional Feedback Paper • 2602.12612 • Published Feb 13 • 4
R1-ACT: Efficient Reasoning Model Safety Alignment by Activating Safety Knowledge Paper • 2508.00324 • Published Aug 1, 2025
Why Far Looks Up: Probing Spatial Representation in Vision-Language Models Paper • 2605.30161 • Published May 28 • 60
Personalize-then-Store: Benchmarking and Learning Personalized Memory for Long-horizon Agents Paper • 2605.25535 • Published May 25 • 46
Is Safety Standard Same for Everyone? User-Specific Safety Evaluation of Large Language Models Paper • 2502.15086 • Published Feb 20, 2025 • 16
Is Safety Standard Same for Everyone? User-Specific Safety Evaluation of Large Language Models Paper • 2502.15086 • Published Feb 20, 2025 • 16