Vons is a research project providing compact local decision models for AI agents that run in the browser. It includes a TypeScript SDK, Chrome extension, and Python tools for training and evaluation, designed to keep model sizes under 64 MiB and enable CPU/WASM execution while maintaining user control over policy and execution.
Potluck is a local AI application that lets users run language models across their own computers without internet connectivity. It supports single-machine inference, distributed computing across multiple devices via WireGuard, and integrates with coding tools through MCP for persistent project context.
A tool that analyzes Hacker News comments for sentiment, sarcasm, emotion, toxicity, and usefulness using Laya, a local System One decision model. It fetches random comments via the Hacker News API and scores them without requiring an LLM, API keys, or internet after the initial model download.
Running LLM models locally for AI agents involves significant tradeoffs. Unlike simple chatbots, agents require large context windows (64k+) because they inject system prompts, tool descriptions, and persistent information into every message, causing token usage to spike 10-40k tokens for simple inputs. This overhead is necessary for agent functionality but makes local inference impractical for most users.
Hugging Face's Transformers library now supports running llama.cpp quantized models (GGUF format) locally on Apple Silicon Macs, making it easier to run AI models on personal machines. The integration reuses llama.cpp's ggml kernels for performance and supports various quantization levels like Q4_K_M to balance model size and quality.
EveryCli is a natural-language CLI assistant written in Rust that runs entirely locally using ONNX for semantic search. It combines lexical matching with AI-powered reranking to help users find shell commands by describing their intent in English or French, without network calls or requiring external dependencies.