Enterprise AI developers typically rely on structural scanners to verify agent skills, a process that recent research shows has little predictive value for actual performance. Nvidia's Agentic Continuous Evaluation of Skills (ACES) tool follows a study of 145 internal and public skills that found a Spearman rho correlation of only 0.14 between static scan scores—covering style and security—and LLM-judge quality. The open-source evaluator implements a "Skill Lift" metric, which compares an agent's success on a task when a specific skill is loaded versus when it is absent.

Evaluations of 947 paired cases across 58 production skills revealed a mean Skill Lift of 0.2134, with only 72.8% of cases producing a positive gain. Nvidia noted that skills appearing correct on paper can cause regressions during deployment through incorrect routing or increased operational overhead. The largest process-metric gains appeared in skill execution, behavior checks, and skill efficiency.

Sign in to suggest edits

Key sources

  1. SUPPORT@omarsar0“the largest process-metric gains appear in skill execution, behavior check, and skill efficiency”x.com
  2. SUPPORT@theultronai“A skill can look perfectly fine on paper... and still make the agent worse in practice”x.com
Markdown