A recent Guardian report showed researchers asking frontier models to help plan violence and crime. Frontier labs (OpenAI, Google, Anthropic) responded that they’d already employed better safeguards since the research concluded. Ignoring the fact that you can still work your way around safeguards in many cases, they make no mention of local models.

I’ve experimented with highly capable local models - both from Google and from Alibaba - Gemini and Qwen - and found them to have similar safeguarding to hosted models. But then I checked the “Uncensored” versions of the same models, where researchers and enthusiasts take the base model and surgically isolate and remove the weights relating to guardrails and alignment. They report that the models answer 100% or 95+% of the unsafe questions asked of it, vs 1-5% before the further work.

I’ve been able to replicate similar findings to the researchers but with no hesitation at all on local models. How-to guides on illicit substances, planning unethical and dangerous acts, encouragement to act, or absence of discouragement.

So far the barrier to use for these models has been requiring expensive hardware, however as model architecture efficiency has improved, the hardware required for strong performance has dropped. Using a pre-pandemic era GPU (Nvidia 3060) I’m able to run these models at 20-25 tokens per second, which is plenty fast enough to get a response within a minute or two at most.

The next optimisation, hardware generation or shady company hosting one of these models for free against ads or selling around the model will enable widespread use of these models by anyone aware of them.

The models are free to manipulate, given the state of competition in the space; the Chinese labs release high quality free models hot on the heels of US frontier companies’ latest releases.

Just to reiterate: there are zero safeguards when an uncensored model is running on a private machine. The benefits of this are many, for the sake of privacy and exploring ideas and personal development. The downsides are equally as many, with risk of psychosis, dangerous suggestions, homebrew versions of dangerous materials and underage use of adult information.

So OpenAI, Anthropic, Google, can defend their own stances till they are blue in the face, but the industry is tipping towards ubiquitous free use of “good enough, but uncensored” models that will very soon be available on any hardware. Which do you think a teenager or unstable person will pick - the paid and locked down offer at $200+ per year, or the free uncensored version at no cost, other than electricity?

The safeguarding we need is likely much broader than regulation of the corporations. We potentially need to handle distribution of open models like dangerous substances.