The application of large language models to security work presents a distinctive combination of characteristics, including rapid gains in capability, a fast-expanding set of autonomous tools, and limited public evidence of how these tools behave in practice. This underscores the need for structured, objective assessments that move beyond capability claims and isolated benchmarks. A meaningful evaluation requires observing an autonomous source code auditor end-to-end, documenting where it performs effectively, where it encounters limitations, and the operational costs and trade-offs associated with its use.

The currently published series focuses on Project Acuario, a multi-agent source code security auditor developed as a research and development initiative and examined through a single, fully instrumented assessment. The findings will be documented across a five-article series. 

pdf-iconIntroduction & Approach
pdf-iconContainment & sandboxing
pdf-iconCountering false positives & biasness