arXiv 2507.06134
OpenAgentSafety: A Comprehensive Framework for Evaluating Real-World AI Agent Safety
By Sanidhya Vijayvargiya, Aditya Bharat Soni, et al.
Published 2025-07-08
Citation lineage
Review the prior work and downstream research connected to this paper.
Recent advances in AI agents capable of solving complex, everyday tasks, from scheduling to customer service, have enabled deployment in real-world settings, but their possibilities for unsafe behavior demands rigorous evaluation. While prior benchmarks have attempted to assess agent safety, most fall short by relying on simulated environments, narrow task domains, or unrealistic tool abstractions. We introduce Open…