Reinforcement Learning for Prompt Optimization in Language Models: A Comprehensive Survey of Methods, Representations, and Evaluation Challenges
Review Article  ·  Published: 14 September 2025
Issue cover
ICCK Transactions on Emerging Topics in Artificial Intelligence
Volume 2, Issue 4, 2025: 173-181
Review Article Open Access

Reinforcement Learning for Prompt Optimization in Language Models: A Comprehensive Survey of Methods, Representations, and Evaluation Challenges

1 Brown University, Providence, RI 02912, United States
* Corresponding Author: Zhangqi Liu, [email protected]
Volume 2, Issue 4

Abstract

The growing prominence of prompt engineering as a means of controlling large language models has given rise to a diverse set of methods, ranging from handcrafted templates to embedding-level tuning. Yet, as prompts increasingly serve not merely as input scaffolds but as adaptive interfaces between users and models, the question of how to systematically optimize them remains unresolved. Reinforcement learning, with its capacity for sequential decision-making and reward-driven adaptation, has been proposed as a possible framework for discovering effective prompting strategies. This survey explores the emerging intersection of RL and prompt engineering, organizing existing research along three interdependent axes: the representation of prompts (symbolic, soft, and hybrid), the design of RL-based optimization mechanisms, and the challenges of evaluating and generalizing learned prompt policies. Rather than presenting a single unified framework, the discussion reflects the fragmented, often experimental nature of current approaches, many of which remain constrained by unstable reward signals, limited generalizability, and a lack of reproducible evaluation standards. By analyzing methodological innovations and points of friction alike, this work aims to foster a more critical and reflective understanding of what it means to "learn to prompt" in complex, real-world language modeling contexts.

Graphical Abstract

Reinforcement Learning for Prompt Optimization in Language Models: A Comprehensive Survey of Methods, Representations, and Evaluation Challenges

Keywords

prompt engineering reinforcement learning language models prompt optimization reward design prompt representation

Data Availability Statement

Not applicable.

Funding

This work was supported without any funding.

Conflicts of Interest

The author declares no conflicts of interest.

Ethical Approval and Consent to Participate

Not applicable.

References

  1. Zheng, H., Shen, L., Tang, A., Luo, Y., Hu, H., Du, B., ... & Tao, D. (2025). Learning from models beyond fine-tuning. Nature Machine Intelligence, 7(1), 6-17.
    [CrossRef] [Google Scholar]
  2. Liu, P., Yuan, W., Fu, J., Jiang, Z., Hayashi, H., & Neubig, G. (2023). Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing. ACM computing surveys, 55(9), 1-35.
    [CrossRef] [Google Scholar]
  3. Banerjee, C., & Nazir, M. S. (2025). Zero-Shot Llms in Human-in-The-Loop Rl: Replacing Human Feedback for Reward Shaping. Available at SSRN 5218722. https://dx.doi.org/10.2139/ssrn.5218722
    [Google Scholar]
  4. Zhang, T., Wang, X., Zhou, D., Schuurmans, D., & Gonzalez, J. E. (2022). Tempera: Test-time prompting via reinforcement learning. arXiv preprint arXiv:2211.11890.
    [CrossRef] [Google Scholar]
  5. Li, X. L., & Liang, P. (2021, August). Prefix-tuning: Optimizing continuous prompts for generation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers) (pp. 4582-4597).
    [CrossRef] [Google Scholar]
  6. Lester, B., Al-Rfou, R., & Constant, N. (2021, November). The Power of Scale for Parameter-Efficient Prompt Tuning. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (pp. 3045-3059).
    [CrossRef] [Google Scholar]
  7. Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., ... & Lowe, R. (2022). Training language models to follow instructions with human feedback. Advances in neural information processing systems, 35, 27730-27744.
    [Google Scholar]
  8. Liu, X., Zheng, Y., Du, Z., Ding, M., Qian, Y., Yang, Z., & Tang, J. (2024). GPT understands, too. AI open, 5, 208-215.
    [CrossRef] [Google Scholar]
  9. Wang, X., Li, C., Wang, Z., Bai, F., Luo, H., Zhang, J., ... & Hu, Z. (2024, May). Promptagent: Strategic planning with language models enables expert-level prompt optimization. In International Conference on Learning Representations (Vol. 2024, pp. 23967-24001).
    [Google Scholar]
  10. Xing, Y., & Liu, P. (2023, June). Prompt and instruction-based tuning for response generation in conversational question answering. In International conference on applications of natural language to information systems (pp. 156-169). Cham: Springer Nature Switzerland.
    [CrossRef] [Google Scholar]
  11. Meynhardt, C., Meybohm, P., Kranke, P., & Hölzing, C. R. (2025). Advanced Prompt Engineering in Emergency Medicine and Anesthesia: Enhancing Simulation-Based e-Learning. Electronics, 14(5), 1028.
    [CrossRef] [Google Scholar]
  12. Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., & Amodei, D. (2017). Deep reinforcement learning from human preferences. Advances in neural information processing systems, 30.
    [Google Scholar]
  13. Deng, M., Wang, J., Hsieh, C. P., Wang, Y., Guo, H., Shu, T., ... & Hu, Z. (2022, December). Rlprompt: Optimizing discrete text prompts with reinforcement learning. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing (pp. 3369-3391).
    [CrossRef] [Google Scholar]
  14. Xu, M., Shen, Y., Zhang, S., Lu, Y., Zhao, D., Tenenbaum, J., & Gan, C. (2022, June). Prompting decision transformer for few-shot policy generalization. In international conference on machine learning (pp. 24631-24645). PMLR.
    [Google Scholar]
  15. Shin, T., Razeghi, Y., Logan IV, R. L., Wallace, E., & Singh, S. (2020, November). AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) (pp. 4222-4235).
    [CrossRef] [Google Scholar]
  16. Do, X. L., Dinh, D., Nguyen, N. H., Kawaguchi, K., Chen, N., Joty, S., & Kan, M. Y. (2025, July). What Makes a Good Natural Language Prompt?. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 5835-5873).
    [CrossRef] [Google Scholar]
  17. Pryzant, R., Iter, D., Li, J., Lee, Y., Zhu, C., & Zeng, M. (2023, December). Automatic prompt optimization with “gradient descent” and beam search. In Proceedings of the 2023 conference on empirical methods in natural language processing (pp. 7957-7968).
    [CrossRef] [Google Scholar]
  18. Saletta, M., & Ferretti, C. (2024, July). Exploring the prompt space of large language models through evolutionary sampling. In Proceedings of the Genetic and Evolutionary Computation Conference (pp. 1345-1353).
    [CrossRef] [Google Scholar]
  19. Liu, S., Fang, Y., Cheng, H., Pan, Y., Liu, Y., & Gao, C. (2023, November). Large Language Models guided Generative Prompt for Dialogue Generation. In 2023 International Conference on Cyber-Enabled Distributed Computing and Knowledge Discovery (CyberC) (pp. 10-17). IEEE.
    [CrossRef] [Google Scholar]
  20. Schulhoff, S., Ilie, M., Balepur, N., Kahadze, K., Liu, A., Si, C., ... & Resnik, P. (2024). The prompt report: A systematic survey of prompt engineering techniques. arXiv preprint arXiv:2406.06608.
    [CrossRef] [Google Scholar]
  21. Nadizar, G., Rovito, L., De Lorenzo, A., Medvet, E., & Virgolin, M. (2024). An analysis of the ingredients for learning interpretable symbolic regression models with human-in-the-loop and genetic programming. ACM Transactions on Evolutionary Learning and Optimization, 4(1), 1-30.
    [CrossRef] [Google Scholar]
  22. Jain, A. M., & Jindal, M. (2025, March). Systematic survey of various prompt optimization methods and their classifications. In 2025 11th International Conference on Computing and Artificial Intelligence (ICCAI) (pp. 524-536). IEEE.
    [CrossRef] [Google Scholar]
  23. Wen, Y., Jain, N., Kirchenbauer, J., Goldblum, M., Geiping, J., & Goldstein, T. (2023). Hard prompts made easy: Gradient-based discrete optimization for prompt tuning and discovery. Advances in Neural Information Processing Systems, 36, 51008-51025.
    [Google Scholar]
  24. Shorinwa, O., Mei, Z., Lidard, J., Ren, A. Z., & Majumdar, A. (2024). A Survey on Uncertainty Quantification of Large Language Models: Taxonomy, Open Research Challenges, and Future Directions. arXiv preprint arXiv:2412.05563.
    [CrossRef] [Google Scholar]
  25. He, Y., Wang, J., Wang, Y., Li, K., Zhong, Y., Song, X., ... & Chen, J. (2025). Enhancing intent understanding for ambiguous prompt: A human-machine co-adaption strategy. arXiv preprint arXiv:2501.15167.
    [Google Scholar]
  26. Milani, S., Topin, N., Veloso, M., & Fang, F. (2024). Explainable reinforcement learning: A survey and comparative review. ACM Computing Surveys, 56(7), 1-36.
    [CrossRef] [Google Scholar]
  27. Zhou, Y., Muresanu, A. I., Han, Z., Paster, K., Pitis, S., Chan, H., & Ba, J. (2023). Large language models are human-level prompt engineers. In The Eleventh International Conference on Learning Representations (ICLR).
    [CrossRef] [Google Scholar]
  28. Ciniselli, M., Cooper, N., Pascarella, L., Mastropaolo, A., Aghajani, E., Poshyvanyk, D., ... & Bavota, G. (2021). An empirical study on the usage of transformer models for code completion. IEEE Transactions on Software Engineering, 48(12), 4818-4837.
    [CrossRef] [Google Scholar]
  29. Sheilsspeigh, P., Larkspur, M., Carver, S., Roskilde, F., & Longmore, S. (2024). Dynamic context shaping: A new approach to adaptive representation learning in large language models. Authorea.
    [CrossRef] [Google Scholar]
  30. Yang, C., Wang, X., Lu, Y., Liu, H., Le, Q. V., Zhou, D., & Chen, X. (2024, May). Large language models as optimizers. In International Conference on Learning Representations (Vol. 2024, pp. 12028-12068).
    [Google Scholar]
  31. Deng, Z., Ma, W., Han, Q. L., Zhou, W., Zhu, X., Wen, S., & Xiang, Y. (2025). Exploring DeepSeek: A Survey on Advances, Applications, Challenges and Future Directions. IEEE/CAA Journal of Automatica Sinica, 12(5), 872-893.
    [CrossRef] [Google Scholar]

Cited By (8)

  1. Shengjie Ye. Emotional Engagement with AI-generated vs Traditional Art. Scientific Research Bulletin, 2026 , 2 (5).
    [CrossRef]
  2. Qian Cui, Enhao Ning, Minghua Du, Baoli Lu, Shuang Li, Jiong Xiang, Shuyao He. Uncertainty-aware traffic prediction and routing for emergency vehicles: a multi-objective risk-aware framework. Connection Science, 2026 , 38 (1).
    [CrossRef]
  3. Quan Su. Designing Emotion-adaptive Human - AI Interfaces: An Empirical Study on Empathy, Trust, and Context-aware Interaction. Scientific Research Bulletin, 2026 , 2 (5).
    [CrossRef]
  4. Wenqiang Lu, Shengjie Ye. <i><b>The Role of Generative AI in Economic Research: Enhancing Productivity and Cognitive Automation</b></i>. Al lnnovations and Applications, 2026 , 2 (1).
    [CrossRef]
  5. Pelin Karaçay, Polat Goktas. The emerging future of healthcare simulation with artificial intelligence integration. Teaching and Learning in Nursing, 2026 , 21 (3).
    [CrossRef]
  6. Kai Ye. <i><b>Economic Policy Challenges in the Age of Artificial General Intelligence</b></i>. Al lnnovations and Applications, 2025 , 1 (1).
    [CrossRef]
  7. Shengjie Ye. <i><b>Cultural Representations in Artificial Intelligence and Finance from a Cross-Cultural Perspective</b></i>. Al lnnovations and Applications, 2025 , 1 (1).
    [CrossRef]
  8. Fan hao. <i><b>Maximum Macroeconomic Impacts of AI:</b><b> </b><b>Automation, Task Complementarity, and Their Effects on Productivity and Inequality</b></i>. Al lnnovations and Applications, 2025 , 1 (1).
    [CrossRef]
* Citation data provided by Crossref Cited-by.

Cite This Article

APA Style
Liu, Z. (2025). Reinforcement Learning for Prompt Optimization in Language Models: A Comprehensive Survey of Methods, Representations, and Evaluation Challenges. ICCK Transactions on Emerging Topics in Artificial Intelligence, 2(4), 173-181. https://doi.org/10.62762/TETAI.2025.790504
Export Citation
RIS Format
Compatible with EndNote, Zotero, Mendeley, and other reference managers
TY  - JOUR
AU  - Liu, Zhangqi
PY  - 2025
DA  - 2025/09/14
TI  - Reinforcement Learning for Prompt Optimization in Language Models: A Comprehensive Survey of Methods, Representations, and Evaluation Challenges
JO  - ICCK Transactions on Emerging Topics in Artificial Intelligence
T2  - ICCK Transactions on Emerging Topics in Artificial Intelligence
JF  - ICCK Transactions on Emerging Topics in Artificial Intelligence
VL  - 2
IS  - 4
SP  - 173
EP  - 181
DO  - 10.62762/TETAI.2025.790504
UR  - https://www.icck.org/article/abs/TETAI.2025.790504
KW  - prompt engineering
KW  - reinforcement learning
KW  - language models
KW  - prompt optimization
KW  - reward design
KW  - prompt representation
AB  - The growing prominence of prompt engineering as a means of controlling large language models has given rise to a diverse set of methods, ranging from handcrafted templates to embedding-level tuning. Yet, as prompts increasingly serve not merely as input scaffolds but as adaptive interfaces between users and models, the question of how to systematically optimize them remains unresolved. Reinforcement learning, with its capacity for sequential decision-making and reward-driven adaptation, has been proposed as a possible framework for discovering effective prompting strategies. This survey explores the emerging intersection of RL and prompt engineering, organizing existing research along three interdependent axes: the representation of prompts (symbolic, soft, and hybrid), the design of RL-based optimization mechanisms, and the challenges of evaluating and generalizing learned prompt policies. Rather than presenting a single unified framework, the discussion reflects the fragmented, often experimental nature of current approaches, many of which remain constrained by unstable reward signals, limited generalizability, and a lack of reproducible evaluation standards. By analyzing methodological innovations and points of friction alike, this work aims to foster a more critical and reflective understanding of what it means to "learn to prompt" in complex, real-world language modeling contexts.
SN  - 3068-6652
PB  - Institute of Central Computation and Knowledge
LA  - English
ER  - 
BibTeX Format
Compatible with LaTeX, BibTeX, and other reference managers
@article{Liu2025Reinforcem,
  author = {Zhangqi Liu},
  title = {Reinforcement Learning for Prompt Optimization in Language Models: A Comprehensive Survey of Methods, Representations, and Evaluation Challenges},
  journal = {ICCK Transactions on Emerging Topics in Artificial Intelligence},
  year = {2025},
  volume = {2},
  number = {4},
  pages = {173-181},
  doi = {10.62762/TETAI.2025.790504},
  url = {https://www.icck.org/article/abs/TETAI.2025.790504},
  abstract = {The growing prominence of prompt engineering as a means of controlling large language models has given rise to a diverse set of methods, ranging from handcrafted templates to embedding-level tuning. Yet, as prompts increasingly serve not merely as input scaffolds but as adaptive interfaces between users and models, the question of how to systematically optimize them remains unresolved. Reinforcement learning, with its capacity for sequential decision-making and reward-driven adaptation, has been proposed as a possible framework for discovering effective prompting strategies. This survey explores the emerging intersection of RL and prompt engineering, organizing existing research along three interdependent axes: the representation of prompts (symbolic, soft, and hybrid), the design of RL-based optimization mechanisms, and the challenges of evaluating and generalizing learned prompt policies. Rather than presenting a single unified framework, the discussion reflects the fragmented, often experimental nature of current approaches, many of which remain constrained by unstable reward signals, limited generalizability, and a lack of reproducible evaluation standards. By analyzing methodological innovations and points of friction alike, this work aims to foster a more critical and reflective understanding of what it means to "learn to prompt" in complex, real-world language modeling contexts.},
  keywords = {prompt engineering, reinforcement learning, language models, prompt optimization, reward design, prompt representation},
  issn = {3068-6652},
  publisher = {Institute of Central Computation and Knowledge}
}

Article Metrics

Citations
Views
11265
PDF Downloads
6580

Publisher's Note

ICCK stays neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Rights and Permissions

CC BY Copyright © 2025 by the Author(s). Published by Institute of Central Computation and Knowledge. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https://creativecommons.org/licenses/by/4.0/), which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made.
ICCK Transactions on Emerging Topics in Artificial Intelligence
ICCK Transactions on Emerging Topics in Artificial Intelligence
ISSN: 3068-6652 (Online)
Portico
Preserved at
Portico