A framework for training large language models using LoRA with GGUF quantization on limited VRAM, enabling training of models like Qwen3.8-Flash-Next (125B parameters) on 40GB GPUs. The approach combines optimized kernels, quantization techniques, and memory-efficient training methods to make open-weight model fine-tuning accessible on consumer hardware.
VibeSys, a multi-agent system, built a serving engine for Qwen3.5-397B on AMD MI300As that achieves 2,242 tok/s of goodput, 2.33× better than tuned SGLang. Over 105 hours, agents iteratively optimized the engine for the hybrid mixture-of-experts model on a multi-turn chat workload without human code writing.
A technical article on Kolibri, Aleph Alpha's sovereign language model, examining architectural trade-offs in autoregressive models including parameter allocation, FLOPs scaling, and sequence-mixer state. The piece provides an interactive tool for configuring and comparing model architectures against recent open-weight releases.