← Back to live feed1 story
Google DeepMind and AVERI tested Gemini 2.5 Flash-Lite using a secure enclave to keep both model weights and test prompts private. The evaluation utilized the AILuminate safety benchmark from MLCommons in a double-blind process where neither the developer nor the evaluator had access to the other's sensitive assets.
The project, which also involved OpenMined and MLCommons, seeks to eliminate benchmark contamination, a problem where AI models are trained on the very data used to test them. This production-scale test follows a 2024 pilot involving GPT-2 and the AI Security Institute, and is part of an AVERI strategy to establish open source auditing standards for frontier models.
Sign in to suggest edits