A crucial reason that it is possible to train ML systems to outperform human experts in narrow domains such as protein folding or chess, is because for these well-defined problems, it is easy to produce a reliable reward signal. However, current techniques for aligning frontier models with human goals, such as human feedback, are only […]
Read More