WorkshopOnline

Multi-Agent Reinforcement Learning: AlphaZero Self Play Training

17:30Online, United States

  • This workshop explores modifications to the AlphaZero self-play training algorithm for two-player zero-sum games, focusing on handling player switching and population-based training progress.
  • It targets students and researchers interested in multi-agent reinforcement learning and game AI algorithms.
  • Attendees will learn how to build a training pipeline using search simulations for policy improvement.

Details

Date
Start time
17:30
Venue
Online
Country
United States
Format
Online
Organiser
Silicon Valley Generative AI ~ The AI Collective Network
Type
Workshop

About this event

Over the last three meetings (AlphaZero Introduction, MCTS Realtime Policy Improvement, AlphaZero Training Pipeline) we learned how to apply the Monte Carlo tree search algorithm defined in the following paper:

Danihelka, I., Guez, A., Schrittwieser, J., & Silver, D. (2022). Policy Improvement by Planning with Gumbel (ICLR 2022). openreview.net/forum?id=bERaNdoegnO

As a follow-up to the famous AlphaZero algorithm, this paper focuses on using search simulations to improve an existing policy/value function. We showed the details of how this improvement occurs and developed a training pipeline to use search simulations as a data collection tool and master an environment.

So far, we have used an MDP environment with a fixed opponent based on an extended tic-tac-toe game. This time, we turn our attention to the full two-player zero sum game scenario in which the search is used in conjunction with self-play to produce increasingly strong policies. We will show the modifications needed to the original search algorithm needed to handle player switching and then build a training pipeline based on self play.

In an MDP environment, measuring performance is straightforward and we can expect monotonic improvement over time. In the self-play setting, we have two competing agents and the overall success of one player may go through cycles. Therefore, we need to develop other criteria for determining training progress that involves studying the performance across a population of policies over time. We will begin to touch on section 9.9 of Multi-Agent Reinforcement Learning: Foundations and Modern Approaches which covers populations of agents learning simultaneously.

As usual you can find below links to the textbook, previous chapter notes, slides, and recordings of some of the previous meetings.

Meetup Links:
Recordings of Previous RL Meetings
Recordings of Previous MARL Meetings
Short RL Tutorials
My exercise solutions and chapter notes for Sutton-Barto
My MARL repos

Links

Websiteopenreview.net/forum

More upcoming events in United States

FEB
20

USC Marshall Alumni SD Lunch - Vittorio's Italian Trattoria - Monthly

Vittorio's Italian Trattoria, San Diego, United States · USC Lloyd Greif Center for Entrepreneurial Studies
runs 20 February – 20 November 2026
In person
  • This monthly lunch event in San Diego brings together USC Marshall and Leventhal alumni to network and reconnect.
  • Attendees can build new connections while enjoying Italian food at Vittorio's Italian Trattoria.
  • It welcomes all Trojan alumni, friends, and colleagues and requires registration through Eventbrite.
NetworkingFree to Attend
11:30 · Free · NetworkingDetails →
JUN
09

USC Marshall Alumni SD North County Luncheon - Vista Village Pub (monthly)

Vista Village Pub, Vista, United States · USC Lloyd Greif Center for Entrepreneurial Studies
runs 9 June – 11 August 2026
In person
  • USC Marshall and Leventhal alumni gather monthly in Vista to network and reconnect over lunch.
  • The event is open to all Trojan alumni, friends, and colleagues in the San Diego North County area.
  • Attendees have the chance to expand their professional network and enjoy a social setting at Vista Village Pub.
Networking
11:30 · Free · NetworkingDetails →
AUG
01

Code & Coffee with Mentorship Saturdays @ Roseline Coffee

Roseline Coffee, Portland, United States · Portland Code & Coffee
In person
  • This event offers in-person coding sessions combined with mentorship opportunities on Saturdays at Roseline Coffee in Portland.
  • It caters to developers and founders seeking hands-on coding time and expert guidance.
  • Attendees can improve their skills and connect with mentors in a casual coffeehouse environment.
FoundersHands-OnFree to Attend
10:00 · EventDetails →
Share this event