DooDooLamb News

Show HN: A benchmark for AI agent guardrails that caught my own plugin

Brief published August 19, 2026 ยท Original source published August 17, 2026

Original reporting by couldbeme_ at github.com.

Automated brief. Verify important details at the original source.

Show HN: A benchmark for AI agent guardrails that caught my own plugin

What happened

A developer released Holdline, an open-source benchmark designed to evaluate guardrails for AI agents that perform write actions. The tool scores any guard expressed as a function that takes commitments and an action and returns a block decision. Metrics include catch rate, false-block rate, class-balanced kappa, and a dedicated injection-attack class. The project was prompted by the author discovering their own plugin failed the benchmark, suggesting the test surface catches real gaps that developers may miss during self-review.

Why it matters

Write-guard evaluation has lacked a neutral, standardized tool. Builders shipping AI agents with file, database, or API write access currently have no common baseline to compare guard implementations. Holdline proposes a concrete scoring interface and a balanced metric set that penalizes both missed blocks and over-blocking, which are competing failure modes in production systems.

What to watch

Whether the benchmark gains adoption beyond the author's own tooling, and whether the injection-attack class proves comprehensive enough to cover real adversarial patterns seen in deployed agents.

Original source