PPO self-play agents + fine-tuned GPT-2 dialogue (4.6x BLEU over baseline). UCLA ECE 147 final project.
We model two aspects:
- Decision Making - via PPO based reinforcement learning
- Conversation - via finetuning a generative pre-trained transformer
We achieve compelling results on both tasks!
Screenshot of the game board after the impostor sucessfully kills a crewmate:
Sample conversation generation

