ChatGLM-RLHF: Practices of Aligning Large Language Models with Human Feedback

AI-generated keywords: ChatGLM-RLHF

AI-generated Key Points

  • ChatGLM-RLHF is a reinforcement learning from human feedback system designed to enhance ChatGLM's alignment with human preferences
  • The system consists of three main components: collecting human preference data, training the reward model, and optimizing policies
  • Challenges addressed include mitigating reward variance for large-scale training stability, implementing model parallelism with fused gradient-descent, and designing regularization constraints to prevent catastrophic forgetting in large language models (LLMs)
  • Experimental results show that ChatGLM-RLHF outperforms the supervised fine-tuned (SFT) version of ChatGLM in alignment tasks, achieving an average of 15% more wins in Chinese alignment tasks
  • Human evaluations support the effectiveness of RLHF, with the PPO model demonstrating advantages over the SFT model in diverse tasks like creative writing, logical reasoning, semantic analysis, language comprehension, and mathematics
  • Ablation study reveals that both DPO and PPO models significantly increase response length compared to the SFT model, leading to improved automatic evaluation scores for creative writing tasks
  • Insights from evaluating the reward model suggest it can guide RLHF algorithms with approximately 65% accuracy in mirroring human judgment
Also access our AI generated: Comprehensive summary, Lay summary, Blog-like article; or ask questions about this paper to our AI assistant.

Authors: Zhenyu Hou, Yiin Niu, Zhengxiao Du, Xiaohan Zhang, Xiao Liu, Aohan Zeng, Qinkai Zheng, Minlie Huang, Hongning Wang, Jie Tang, Yuxiao Dong

License: CC BY 4.0

Abstract: ChatGLM is a free-to-use AI service powered by the ChatGLM family of large language models (LLMs). In this paper, we present the ChatGLM-RLHF pipeline -- a reinforcement learning from human feedback (RLHF) system -- designed to enhance ChatGLM's alignment with human preferences. ChatGLM-RLHF encompasses three major components: the collection of human preference data, the training of the reward model, and the optimization of policies. Throughout the process of integrating ChatGLM-RLHF into production, we encountered and addressed several unprecedented challenges. We introduce the strategies to mitigate reward variance for stabilized large-scale training, implement model parallelism with fused gradient-descent, and design regularization constraints to avoid catastrophic forgetting in LLMs. Experiments show that ChatGLM-RLHF brings significant improvements in alignment tasks compared to the supervised fine-tuned (SFT) version of ChatGLM. For instance, it achieves on average 15\% more wins against ChatGLM-SFT in Chinese alignment tasks. The work presents our practices of aligning LLMs with human preferences, offering insights into the challenges and solutions in RLHF implementations.

Submitted to arXiv on 01 Apr. 2024

Ask questions about this paper to our AI assistant

You can also chat with multiple papers at once here.

AI assistant instructions?

Results of the summarizing process for the arXiv paper: 2404.00934v1

, , , , In this paper, the authors introduce ChatGLM-RLHF, a reinforcement learning from human feedback system designed to enhance ChatGLM's alignment with human preferences. The system consists of three main components: collecting human preference data, training the reward model, and optimizing policies. Throughout the integration of ChatGLM-RLHF into production, the authors address various challenges such as mitigating reward variance for large-scale training stability, implementing model parallelism with fused gradient-descent, and designing regularization constraints to prevent catastrophic forgetting in large language models (LLMs). Experimental results show that ChatGLM-RLHF outperforms the supervised fine-tuned (SFT) version of ChatGLM in alignment tasks. For example, in Chinese alignment tasks, ChatGLM-RLHF achieves an average of 15% more wins against ChatGLM-SFT. Human evaluations further support the effectiveness of RLHF, with the PPO model demonstrating a clear advantage over the SFT model in diverse tasks like creative writing, logical reasoning, semantic analysis, language comprehension, and mathematics. Additionally, an ablation study reveals that both DPO and PPO models significantly increase response length compared to the SFT model. This increase in response length correlates with improved automatic evaluation scores for creative writing tasks. Furthermore, insights from evaluating the reward model suggest that it can guide RLHF algorithms with approximately 65% accuracy in mirroring human judgment. Overall,<keyword1>ChatGLM-RLHF</keyword1> is a successful implementation of <keyword2>reinforcement learning</keyword2> techniques to improve AI services powered by large language models. Through its three main components, the system effectively collects human preference data, trains the reward model, and optimizes policies to align with human preferences. The authors also address challenges such as mitigating reward variance and implementing model parallelism to ensure stable training for large-scale models. Experimental results demonstrate that <keyword1>ChatGLM-RLHF</keyword1> outperforms traditional supervised fine-tuning methods in <keyword3>alignment tasks</keyword3>, achieving an average of 15% more wins in Chinese alignment tasks. Human evaluations further support the effectiveness of <keyword1>ChatGLM-RLHF</keyword1>, with the PPO model showing a clear advantage over the SFT model in diverse tasks such as creative writing, logical reasoning, semantic analysis, language comprehension, and mathematics. An ablation study also reveals that both DPO and PPO models significantly increase response length compared to the SFT model. This increase in response length is correlated with improved automatic evaluation scores for creative writing tasks. Insights from evaluating the reward model suggest that it can guide RLHF algorithms with approximately 65% accuracy in mirroring human judgment. This highlights the importance of incorporating human evaluations alongside automatic evaluations to accurately assess model performance across various tasks. In conclusion,<keyword1>ChatGLM-RLHF</keyword1> provides valuable insights into aligning large language models (LLMs) with human preferences through reinforcement learning techniques. Its success in improving response length and task-specific performance underscores its potential to enhance AI services powered by LLMs. The findings emphasize the importance of incorporating human feedback into AI systems to better align them with human preferences and improve their overall performance.
Created on 29 Jul. 2026

Assess the quality of the AI-generated content by voting

Score: 0

Why do we need votes?

Votes are used to determine whether we need to re-run our summarizing tools. If the count reaches -10, our tools can be restarted.

Similar papers summarized with our AI tools

Navigate through even more similar papers through a

tree representation

Look for similar papers (in beta version)

By clicking on the button above, our algorithm will scan all papers in our database to find the closest based on the contents of the full papers and not just on metadata. Please note that it only works for papers that we have generated summaries for and you can rerun it from time to time to get a more accurate result while our database grows.

Disclaimer: The AI-based summarization tool and virtual assistant provided on this website may not always provide accurate and complete summaries or responses. We encourage you to carefully review and evaluate the generated content to ensure its quality and relevance to your needs.