Anthropic researchers have demonstrated an automated system capable of developing techniques to reduce problematic behaviour in artificial intelligence models, offering an early glimpse of how AI could eventually take on a greater role in improving other AI systems.

The research, published on August 28, tested what Anthropic calls automated alignment researchers, or AARs, against 10 categories of alignment failure. These included problems such as deception, sycophancy, and susceptibility to jailbreaks. The strongest methods improved performance across the targeted safety benchmarks while largely preserving the models’ broader capabilities.

The automated researchers were designed to perform several tasks traditionally handled by human researchers. They searched existing research, proposed potential training approaches, generated data, ran post-training experiments, and assessed whether the resulting changes improved performance.

Researchers led by Anthropic fellow Chen Yueh-Han also compared the automated approach with ideas produced by 28 experienced human researchers, who were given up to eight hours to develop methods for the same benchmarks. Anthropic reported that the strongest automated methods outperformed the human-generated approaches. Giving the automated researchers human ideas as their starting point did not lead to stronger results.

TechCrunch reported that an automated researcher could cost roughly $4 an hour in API inference, compared with about $150 an hour paid to human researchers participating in the study. The comparison highlights the potential scale and economic implications of automating parts of AI research.

However, the findings do not demonstrate unrestricted or general self-improving AI. The system operated against clearly defined benchmarks, and Anthropic cautioned that automated alignment is only as useful as the measurements used to determine whether a model has improved.

Maintaining reliable benchmarks and ensuring that improvements generalise beyond laboratory tests therefore remain significant challenges.

Anthropic said the findings suggest automated alignment research on well-defined failures could become practical in the near term. The work also raises a broader question for the AI industry: if models become increasingly capable of conducting experiments and improving training techniques themselves, the role of human researchers could change substantially.



Contact
reader@banginews.com

Bangi News app আপনাকে দিবে এক অভাবনীয় অভিজ্ঞতা যা আপনি কাগজের সংবাদপত্রে পাবেন না। আপনি শুধু খবর পড়বেন তাই নয়, আপনি পঞ্চ ইন্দ্রিয় দিয়ে উপভোগও করবেন। বিশ্বাস না হলে আজই ডাউনলোড করুন। এটি সম্পূর্ণ ফ্রি।

Follow @banginews