The gains are consistent across environment, action model (Qwen3-8B and GPT-5.2), insight source
(DeepSeek-R1, GPT-5.2, prebuilt skills), corpus granularity, and retrieval budget. InsightEmb also exceeds an
ALFWorld-trained in-domain retriever despite using no ALFWorld data, and transfer is bidirectional:
an ALFWorld-trained model improves math insight retrieval too.
Together, these results support a simple design claim: math heuristic retrieval and agentic insight
retrieval share the same progress-oriented matching structure, so a model can learn it from public
reasoning data without expensive environment-specific supervision.