Description
Mudita Khurana, Staff Security Engineer, Airbnb, speaks at [un]prompted 2026 on: Rethinking how we evaluate security agents for real-world use.
Security agents are gaining momentum across industry, but the way we evaluate them remains rooted in narr...
Mudita Khurana, Staff Security Engineer, Airbnb, speaks at [un]prompted 2026 on: Rethinking how we evaluate security agents for real-world use.
Security agents are gaining momentum across industry, but the way we evaluate them remains rooted in narrow, outcome-only benchmarks. These evaluations tell us whether an agent produced a correct answer, but not “how” it arrived there or whether that behavior will remain stable once deployed.
In practice, security is not a sequence of isolated tasks. It is a connected, end-to-end workflow that follows a find → confirm exploit → patch → validate loop. Agents that perform well on task-specific benchmarks often fail in these multi-stage settings due to contextual loss and brittle transitions across steps.
This talk introduces a practical, capability-centric framework for evaluating security agents, that emphasizes observability into how agents plan, reason, use tools, and carry context across the security lifecycle & thus enable teams to better judge whether an agent is ready for real-world use.
Read more