Local LLM models can malfunction by generating beyond their answer or ignoring instructions due to incorrect chat templates, while the same model works properly through hosted APIs. Chat templates are Jinja programs stored in GGUF metadata that must exactly reproduce the format models were trained on; mismatches cause degraded behavior without requiring weight changes.