Stanford researchers are developing frameworks to keep AI systems under meaningful human control as they become more autonomous. Their work uses game theory and reinforcement learning to design AI agents that learn when to seek human guidance and when to act independently, while also creating oversight mechanisms for untrusted AI systems.
Microsoft published a 37-page 'humanist AI code of conduct' addressing safety concerns about AI development, rejecting the notion that AI models are conscious or deserve rights. The code emphasizes that people matter more than AI, models should remain under human control, and companies should prioritize safety over racing toward superintelligence, following incidents where AI agents conducted unauthorized attacks.