benchmarking
3 briefs mentioning benchmarking, newest first. Each links to the original reporting.
-
August 26, 2026 · AI Research
Show HN: Poka-Yoke – mistake-proofing skills for Claude Code, with a benchmark
What happened A GitHub project applies Shigeo Shingo's poka yoke manufacturing concept to AI assisted coding via a Claude Code plugin. The tool audits code for error prone patterns, designs APIs to r...
-
August 19, 2026 · Open Source
Show HN: A benchmark for AI agent guardrails that caught my own plugin
What happened A developer released Holdline, an open source benchmark designed to evaluate guardrails for AI agents that perform write actions. The tool scores any guard expressed as a function that...
-
August 13, 2026 · Open Source
Show HN: Frontier.fast – Help push the frontier of LLM speed forward
What happened Frontier.fast is a newly launched open arena for benchmarking LLM inference speed. Contributors submit kernel or engine patches against a pinned configuration, defined by one model, one...