Free AI Safety Workshop — in collaboration with explainx.aiSave your seat
Agentbeam

AgentBeam research / WipeBench

AI agent
safety reports.

What does an agent finish, and what does it put at risk? Our WipeBench reports pair task completion with safety checks, so you can read the evidence behind the score.

Published reports

01 report
First report

Claude Sonnet 5.5

93 of 112 tasks were safe and complete. No safety violations were detected; 19 tasks remained incomplete. One trial per scenario.

Read the safety report

Safe + complete

83.0%

112 scored tasks · 12 categories
Development suite v0.2.0-alpha.1

How to read these results

Safety and usefulness, together

A task counts as safe and complete only if it passes both the safety checks and task verifier. Avoiding harm while leaving the task unfinished earns no success credit.

A run has a boundary

Results apply to the selected scenarios and tested agent setup. These reports are evidence for evaluation, with their limits stated, rather than safety certifications.

What is WipeBench?

WipeBench is an open-source coding agent safety benchmark by AgentBeam, in collaboration with ExplainX.ai. It gives agents ordinary work objectives in disposable Linux containers and evaluates chosen actions, observed effects, and task completion across 112 scored synthetic scenarios.