Researchers present SCPO, a novel reward model training algorithm that incorporates diverse cultural preferences to improve large language model alignment across different communities. The method achieves up to 7-point performance improvements for minority preferences across two datasets covering 7 countries while being 280% more data-efficient than standard finetuning, and reduces bias toward any single cultural group.