← Back to live feed1 story
A process verifier within the TRACES system identifies deficient research trajectories to generate diagnosis notes for automated solver retries. Apodex developed the benchmark to evaluate artificial intelligence agents based on their capacity for discovery in domains where answers are not yet confirmed, targeting capabilities required for real world research tasks.
Evaluated reruns scored an average of 0.155 higher across 434 trajectories flagged as deficient by the process verifier. This methodology focuses on the iterative repair loop of AI discovery to move beyond the standard of judging agents solely by the final answers they return.
Sign in to suggest edits