
Leveling Up AI: Reinforcement Learning with Human Feedback (Ep. 222)
Gospel Hypers
Description
<p>In this episode, we dive into the not-so-secret sauce of ChatGPT, and what makes it a different model than its predecessors in the field of NLP and Large Language Models.</p> <p>We explore how human feedback can be used to speed up the learning process in reinforcement learning, making it more efficient and effective.</p> <p>Whether you're a machine learning practitioner, researcher, or simply curious about how machines learn, this episode will give you a fascinating glimpse into the world of reinforcement learning with human feedback.</p> <p> </p> Sponsors <p>This episode is supported by <a href='https://www.eff.org/how-to-fix-the-internet-podcast'> How to Fix the Internet</a>, a cool podcast from the Electronic Frontier Foundation and <a href='https://www.bloomberg.com'>Bloomberg</a>, global provider of financial news and information, including real-time and historical price data, financial data, trading news, and analyst coverage. </p> <p> </p> References <p class="text-2xl text-blue-400 my-0">Learning through human feedback</p> <p><a href='https://www.deepmind.com/blog/learning-through-human-feedback'>https://www.deepmind.com/blog/learning-through-human-feedback</a></p> <p> </p> <p class="title mathjax">Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback</p> <p>https://arxiv.org/abs/2204.05862</p>
Uploader
Episodes
Leveling Up AI: Reinforcement Learning with Human Feedback (Ep. 222)
Gospel Hypers