Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, Carol Chen, Catherine Olsson, et al.
15 December 2022
Abstract
Trains a model to critique and revise its own outputs against a written set of principles, then learns from the revised outputs. The aim is a helpful assistant that can explain why it declined, rather than one that simply refuses.
@article{bai2022constitutional,
author = {Yuntao Bai and Saurav Kadavath and Sandipan Kundu and Amanda Askell and Jackson Kernion and Andy Jones and Anna Chen and Anna Goldie and Azalia Mirhoseini and Cameron McKinnon and Carol Chen and Catherine Olsson and others},
title = {Constitutional AI: Harmlessness from AI Feedback},
year = {2022},
eprint = {2212.08073},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2212.08073}
}
Ejentic did not author this paper. Credit belongs to the authors named above; the note is ours.
Cut the RLHF pipeline down to one training run. For teams without the infrastructure to run a reward model and PPO, this is usually the practical route into preference tuning.