I’ve seen great results getting an LLM to write thorough and detailed prompts for another LLM. It’s great to see this area explored in a much more thorough and systematic way.
For the past few years, humans have been doing “prompt engineering” to coax the best performance out of different LLMs. In this work, we explored what happens if we train an AI to do that job instead. Link to our #ICLR2026 paper: arxiv.org/abs/2512.04388 Thread: