Thursday, Oct 1, 2026
← Back to live feed · 1 stories across 1 day A new system for creating synthetic training sets from text descriptions has been released to help developers post-train AI models without relying on pre-existing data. Developed by Adaption AI, the tool called Invent a Dataset outperforms GPT-5.6, Claude Opus 5, and Gemini 3.1 Pro across 8 task types with 17% higher quality and 19% greater sample diversity. The diversity gap increases as the dataset grows, reaching a 37% lead at 20K samples with 0.0% duplicates. The project, which involved researchers Sarah Hooker and Shivam Singh, addresses the constraints of a zero data regime where no high-quality curation is available for a specific capability. Invent a Dataset is a prompt based system that converts a simple capability description into large scale datasets to enable adaptive specialization in specific domains.
Key sources
- SOURCEmarketbrief.now