Selected notes of Subhash Chandran on ML, engineering, and assorted curiosities. Generated from a knowledged-managed repository.
PPO — Proximal Policy Optimization
Overview of PPO, the clipped policy-gradient RL algorithm used in RLHF for InstructGPT and original ChatGPT.