The paper, led by Anthropic Fellow Chen Yueh-Han, details how these systems improved performance on every one of 10 specific misaligned behaviors. This research provides early evidence that automated alignment post-training could become practical in the near term, representing a step toward recursive self-improvement. The automated systems cost roughly $4 per hour in API inference, significantly less than the $150 per hour paid to human researchers.