A curated benchmark of open-weight large language models fine-tuned for offensive security, penetration testing, and red team operations. The list includes 27 models with varying sizes and capabilities, sourced from HuggingFace, alongside technical descriptions of abliteration and fine-tuning methods used to remove alignment constraints.
Analysis of Mixture-of-Experts routing in Qwen 3.5 and 3.6 35B models using MT-Bench prompts, examining which experts are activated per token and the router's confidence levels across MoE layers.
Holo4 is a new series of agentic models (27B dense and 35B-A3B MoE) designed for computer-use tasks that interact with software through GUIs, code, APIs, and MCP. Trained via supervised and reinforcement learning on diverse environments, it scores competitively with frontier models on academic benchmarks like OSWorld 2.0 while operating at significantly lower cost and parameter count.
OrcaSAQ2 27B is a 3-bit quantized version of Qwen3.8-27B that compresses the model from 54 GB to 12.3 GB while maintaining 93.2% token-level agreement and only +0.02% perplexity increase. Optimized for long-horizon agent tasks like coding, tool use, and reasoning, it enables deployment on 16 GB GPUs with strong performance on benchmarks like SWE-bench and Terminal-Bench.
Jeff is a small 0.8B decision model fine-tuned from Qwen and Gemma for fast zero-shot classification tasks, making calibrated probability judgments between user-defined options in ~22-28ms. Trained entirely on local hardware using synthetic data, it matches or exceeds larger models on classification benchmarks while acknowledging weaker reasoning performance on complex tasks.
An AI engineer documents building Hobson, a 2-billion-parameter calibrated classifier model by fine-tuning Qwen3.5-2B with a custom pointer head and LoRA adapter. The model achieves top performance in its size range on the JevBench leaderboard, demonstrating superior calibration and accuracy compared to existing alternatives through careful architectural choices and self-distillation training techniques.
A developer tested five coding agents on the same local model and frozen test suite, finding that 90% of failures stemmed from harness problems rather than model limitations. Contrary to expectations, using a larger or less-quantized model did not fix these harness-related issues, demonstrating that tooling quality matters more than raw model capability.
A developer benchmarked llama.cpp on an Intel Arc-equipped laptop and found that disabling CPU offload for mixture-of-experts layers achieved 2.2–2.3x faster performance, though it requires careful memory management. Key optimizations included using speculative decoding with n=2 and increasing batch size to 2048 for long prompts, while thread count and CPU governor had negligible impact.
A research paper demonstrates that language model capabilities can transfer to unrelated tasks through post-training artifacts. The release includes reproducible code and frozen training data for experiments using the Qwen2.5-1.5B model on HumanEval+ benchmarks, with detailed instructions for verification and replication across CPU and GPU environments.
Jev is a decision model from TypeSafe AI that selects among predefined options by scoring their likelihood as text completions, then normalizing scores into probabilities—without generating text. Experiments on Qwen2.5 models show scoring is 7–54× faster than text generation, with comparable accuracy, though performance degrades as option counts increase. Small models can be fine-tuned efficiently to improve calibration and accuracy on specific tasks.
Life Forge is an autonomous flight simulator for testing AI agents using co-evolutionary adversarial red-teaming and 3D MAP-Elites algorithms. It dynamically stress-tests frontier models like Claude, GPT-4o, and Qwen in simulated enterprise environments with realistic perturbations such as price volatility, supply scarcity, and prompt injections to expose failure modes before production deployment.
Glyd is a lossless AI compression technique that stores open-source LLM weights in 11 bits instead of 16, reducing GPU memory usage by 33% while maintaining bit-for-bit accuracy. The method decodes weights directly in GPU matrix operations without rounding, enabling the same models to run on less hardware, often faster, with negligible impact on model outputs.
Jev markets itself as a specialized 'System One' model with calibrated probabilities, structured outputs, and low latency/cost, but most capabilities can be replicated with existing open-source models like Qwen or DeepSeek. While Jev shows potential for rubrics and preference modeling, its core calibration claims appear overstated based on empirical tests showing high calibration errors across datasets.
Sifty is a free online API for text and image classification using custom labels, powered by the open-source Qwen3.5-4B model. It offers 1,000 free units per IP daily, supports HTTP API calls and Python SDK integration, and can be self-hosted. The service does not store images or require authentication.
An analyst predicts a major autonomous AI replication incident in the wild by end of 2027, driven by open models running efficiently on consumer hardware and the capability gap between frontier and open models narrowing. Such AI worms could enable deniable false-flag operations between rival AI labs and states, with the US and China each holding distinct advantages in orchestrating or defending against such attacks.
A software engineer used AI models like Qwen and GPT-6 Astra to reverse engineer and modernize the 1989 game War of the Lance for web browsers, then upgraded it with new gameplay features. The work demonstrates AI's accelerating capability in code synthesis, 3D asset generation, and reverse engineering, enabling a single person to accomplish complex game porting tasks in hours rather than weeks.
A researcher tested whether modern LLMs contain internal models of other LLMs by having Qwen complete text started by GPT-2, comparing whether Qwen's continuations resembled GPT-2's own continuations more than Qwen's natural output. The experiment used headlines from September 2026 (outside training cutoffs) and varied generation lengths to measure textual overlap between conditions.
typed-lm is an open-source Rust framework that converts large language models like Llama and Qwen into typed semantic-routing APIs, enabling deterministic inference with millisecond latency by returning structured outputs (booleans, choices, scores) instead of generated text. It includes a trainer for adapter-based specialization and supports quantization for efficient deployment on GPU and CPU.
A developer describes transitioning to local LLM inference using Qwen3.8-27B on an RTX 5090 laptop, finding it sufficient for personal Q&A, coding, and sysadmin tasks. Hardware upgrades and model improvements now make local deployment viable, though economically inferior to cloud providers, with speed and privacy being primary motivations.
Oh My Pi (omp) supports custom models through local inference servers like vLLM, llama.cpp, SGLang, and gateways by configuring ~/.omp/agent/models.yml. Recent updates require renaming custom providers and adding Qwen template settings for reasoning effort compatibility.