Agent Trajectories Collection Agent trajectories from the AI Safety and Alignment Group’s benchmarks, released for analysing agent behaviour, capabilities, and safety. • 5 items • Updated 1 day ago
AI Safety Evaluation Datasets Collection Datasets and viewer resources for studying evaluation awareness, instrumental behaviour, sabotage, and monitoring in AI agents. • 4 items • Updated 1 day ago
ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D Paper • 2607.19321 • Published 25 days ago
AI Safety Evaluation Datasets Collection Datasets and viewer resources for studying evaluation awareness, instrumental behaviour, sabotage, and monitoring in AI agents. • 4 items • Updated 1 day ago
Agent Trajectories Collection Agent trajectories from the AI Safety and Alignment Group’s benchmarks, released for analysing agent behaviour, capabilities, and safety. • 5 items • Updated 1 day ago