A new reinforcement learning approach on synthetic tasks allows a small language model to perform complex software engineering without a larger teacher. The Fr…

Sign in to suggest edits
Markdown