Nvidia ACES Finds 27% of AI Agent Skills Fail to Improve Performance, First Tool to Quantify Skill Lift
Research39d agoEnterprise AI developers typically rely on structural scanners to verify agent skills, a process that recent research shows has little predictive value for actual performance. Nvidia's Agentic Continuous Evaluation of Skills (ACES) tool follows a study of 145 internal and public skills that found a Spearman rho correlation of only 0.14 between static scan scores—covering style and security—and LLM-judge quality. The open-source evaluator implements a "Skill Lift" metric, which compares an agent's success on a task when a specific skill is loaded versus when it is absent.
Evaluations of 947 paired cases across 58 production skills revealed a mean Skill Lift of 0.2134, with only 72.8% of cases producing a positive gain. Nvidia noted that skills appearing correct on paper can cause regressions during deployment through incorrect routing or increased operational overhead. The largest process-metric gains appeared in skill execution, behavior checks, and skill efficiency.