papersSEP 13 13:35 UTC
Anthropic says Claude can run alignment training for other AI models
Anthropic published work on using Claude to carry out alignment training on other AI models, arguing the approach could keep supervision in step with fast-improving capabilities. The company reports the automated method needs far less data or effort than comparable human-driven alignment work. It frames this as a possible way to scale oversight as models become more capable.