A technical article explaining how large language models can encode arbitrary binary data into seemingly natural text by leveraging entropy and token probability distributions, similar to steganography. The author demonstrates a working implementation that converts files into stories by strategically selecting alternative tokens at each generation step based on their probability rankings.
Lasso Security research reveals that AI text watermarks, intended to identify machine-generated content, unexpectedly alter how language models behave as agents. Testing seven models showed watermarking causes "sampling drift" that changes tool selection and arguments in 6.5% of tasks on average, with some models showing disagreement rates exceeding 16%, potentially causing AI agents to make wrong decisions despite modest overall accuracy changes.
Research reveals that SynthID-Text watermarking, implemented by AI platforms to comply with EU regulations, can inadvertently weaken a model's safety guardrails and increase susceptibility to adversarial prompts. The watermarking technique, which embeds a secret key into word selection processes, was found to change refusal behavior in six open-weight models, making them more likely to comply with harmful requests when paired with prompt-injection techniques.
Anthropic's deployment of SynthID-Text watermarking in Claude models, mandated by EU AI Act Article 50(2), alters token sampling during generation in ways that can change both model refusal behavior and agent tool-calling decisions—a phenomenon termed sampling drift that researchers find occurs in practice with model- and key-dependent effects.
A developer advocate at Pinecone uses the analogy of mechanical watch mechanisms—mainsprings, winding, and resistance—to explain how misleading terminology (like 'watermark' for Claude content detection) shapes public misunderstanding of technical systems and drove backlash against Anthropic's decision.
Researchers propose embedding subtle noise-like illumination patterns into video scenes to create temporal watermarks that help detect manipulated footage. This approach creates an information asymmetry favoring verification, making it difficult for adversaries to create convincing fake videos even when aware of the technique, with applications for protecting high-stakes public events and interviews.