Researchers present SCPO, a novel reward model training algorithm that incorporates diverse cultural preferences to improve large language model alignment across different communities. The method achieves up to 7-point performance improvements for minority preferences across two datasets covering 7 countries while being 280% more data-efficient than standard finetuning, and reduces bias toward any single cultural group.
A X discussion thread covers multiple cryptocurrency security incidents: Harmony's Layer-1 chain shutdown due to past exploits, with ONE tokens migrating to Ethereum; and a SingularityNET bridge exploit causing unauthorized WMTx minting. A third post discusses an ICLR26 paper on preference learning that uses natural-language rationales to improve reward model robustness.