uh oh huh
Instead of actually generating CoT, this paper uses a hypernetwork to predict small specific updates to the LLM's bias parameters, approximating the effect of thinking without generating the trace. On some benchmarks, it matches thinking with only 7% of the latency alphaxiv.org/abs/2610.03039