Janus is a single Go binary that runs .gguf models on your machine (GPU or CPU) and exposes an OpenAI-compatible API plus a built-in web UI. No Python, no Docker, no Ollama required — though Ollama is supported as a backend if you prefer.
Use it your way: call it from the command line (curl, PowerShell, scripts), wire it into Cursor / Cline / any OpenAI client, or use the Web UI — same local models, whatever workflow fits you. Janus is the runner and router; you choose the front end.
Design idea: the model decides what to do; Go runs inference, routes requests, executes tools, and keeps everything local.
- Local inference — llama.cpp via Vulkan (AMD / Intel / NVIDIA) or CPU fallback
- OpenAI-compatible API — /v1/chat/completions,/v1/models, tool listing/calling
- Web UI — Assistant, Chat, Kernel (tool loop), Config, Memory, Skills
- Built-in tools — read/write files, run commands, math, docx in/out, PDF output, OCR (with Tesseract), and more
- Hot-swap models — change .ggufin the UI without restarting
- Optional auth — Basic Auth for admin endpoints when JANUS_AUTH=true
Disk: plan for the model size (often 2–8 GB per model) plus ~50 MB for Janus + llama.dll.
Optional: Tesseract OCR if you want scanned-document OCR tools.
git clone https://github.com/Vibra-Ingenn/Janus.git
cd janus
.\build.ps1build.ps1 downloads pre-built llama.cpp Vulkan DLLs and compiles dist\janus.exe.
If you already have DLLs in lib\windows\:
.\build.ps1 -SkipDownloadPut a .gguf file in the models\ folder. Easiest path — use the included downloader:
go build -o dist\modelget.exe .\cmd\modelget
.\dist\modelget.exe -repo meta-llama/Llama-3.2-3B-Instruct -file Llama-3.2-3B-Instruct-Q8_0.gguf -out .\models\Or download any compatible GGUF from Hugging Face manually.
copy .env.example .envEdit .env — you must set the model path:
INFERENCE_BACKEND=vulkan
JANUS_MODEL_PATH=./models/Llama-3.2-3B-Instruct-Q8_0.gguf
JANUS_MAX_TOKENS=4096
JANUS_AUTH=false.\dist\janus.exeOr rebuild + launch in one step:
.\run.ps1Janus opens http://127.0.0.1:8990 in your browser (disable with JANUS_NO_BROWSER=1).
First startup loads the model into VRAM — expect 10–60 seconds depending on model size and disk speed.
curl http://127.0.0.1:8990/healthYou should see {"status":"ok",...}.
git clone https://github.com/Vibra-Ingenn/Janus.git
cd janus
go mod tidy
go build -o dist/janus ./cmd/janus
cp .env.example .env
# edit .env — set JANUS_MODEL_PATH and INFERENCE_BACKEND=cpu if no Vulkan
./dist/janusOn Linux you need libllama.so next to the binary or on LD_LIBRARY_PATH. See docs/AIR_GAPPED_INSTALL.md for offline setup.
After .\dist\janus.exe starts, your browser should open http://127.0.0.1:8990 (or open that URL yourself). Wait for the status pill to show your model — first load can take 10–60 seconds.
Assistant — best for “get this done” tasks:
- Type what you want in the big text box
- Optional: drag a file onto the upload zone or click attach
- Click Get Started
- Read the result; use Copy or start a New Request
Chat — best for back-and-forth conversation:
- Open the Chat tab
- Type a message and press Enter (Shift+Enter for a new line)
- Pick a model from the dropdown if you have more than one
Click ⚙ Advanced in the top nav to reveal:
- Advanced → Config
- Choose a model from the list (or paste a path to ./models/your-model.gguf)
- Click Save & Load Model — no restart needed
Focus the terminal where Janus is running and press Ctrl+C.
More walkthroughs (uploads, troubleshooting, tools): docs/USER_MANUAL.md
Janus is meant to stay out of your way — run prompts from a terminal, a script, or a third-party app by pointing it at Janus like any other OpenAI endpoint:
curl http://127.0.0.1:8990/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "local",
"messages": [{"role": "user", "content": "Hello!"}]
}'Base URL: http://127.0.0.1:8990/v1
API key: not required when JANUS_AUTH=false
Full tool reference: docs/TOOLS_REFERENCE.md
Operator guide: docs/USER_MANUAL.md
If you already use Ollama instead of local GGUF:
INFERENCE_BACKEND=ollama
OLLAMA_BASE_URL=http://127.0.0.1:11434
OLLAMA_MODEL=mistral:7bJanus proxies chat to Ollama; tools and the web UI still work.
If you're building or hacking on Janus, these are the gotchas that burned us repeatedly. None of this is obvious the first time.
If something still feels possessed, check logs/janus.log in the project folder and the terminal output from startup.
- Check JANUS_MODEL_PATHin.envmatches a real file undermodels/
- Use Config → Save & Load Model in the web UI to pick from detected .gguffiles
- Paths are relative to where you run janus.exe(usually project root ordist/)
Normal — the model loads on first request or at startup. Smaller quantizations (Q4, Q5) load faster than Q8.
- Update GPU drivers
- Try CPU: INFERENCE_BACKEND=cpuandJANUS_GPU_LAYERS=0
- Or use Ollama backend (above)
See the pitfalls table — almost always a stale janus.exe on Windows, or the wrong port (default 8990, not 8080). If another app owns the port, set JANUS_LISTEN_ADDR=127.0.0.1:8991 in .env.
When JANUS_AUTH=true, Janus prints an admin password on first run. Use it for Config/Model admin routes, or set JANUS_ADMIN_PASSWORD in .env before first launch.
- OCR needs Tesseract installed and on PATH
- PDF text extraction in OSS is limited — use ocr_extractfor scans, or attach plain text / Word files
- Output PDFs via render_pdfwork without extra installs
cmd/janus/ Main server
cmd/modelget/ Hugging Face model downloader
internal/kernel/ ReAct loop (model + tools)
internal/tools/ Built-in and community tools
internal/engine/ llama.cpp Vulkan/CPU backend
internal/bridge/ DLL loader
models/ Put .gguf files here (not committed)
dist/ janus.exe + llama.dll after build
docs/ Manuals and references
Hacking on Janus is welcome. Read Pitfalls (learned the hard way) first — especially the Go-on-Windows rule: stop all janus.exe processes before go build, or you’ll think your fix didn’t work.
Typical loop:
Get-Process -Name "janus" -ErrorAction SilentlyContinue | Stop-Process -Force
go test ./...
go build -o dist\janus.exe .\cmd\janus
.\dist\janus.exeOr just .\run.ps1 (stop → build → run in one step). For a full Windows build including llama DLLs, use .\build.ps1.
Contributions welcome — see CONTRIBUTING.md.
Janus is complete on its own — nothing here is missing or locked behind a paywall. Use it from the CLI, scripts, Cursor, the Web UI, or any OpenAI-compatible client.
If you enjoy running local models and want to see what else is out there, the same team makes Vibe Engine PRO — a separate product focused on recipe-based automation (multi-step workflows, MCP, pipelines that run in Go without calling the model every step). It is one option, not a requirement. You can also extend Janus yourself via the API.
Vibe Engine connects to a local OpenAI-compatible endpoint — Janus works for that:
Base URL: http://127.0.0.1:8990/v1
API key: (leave blank when JANUS_AUTH=false)
Early subscribers at $8.99/month stay at that rate if the listed price goes up later.
Take a look if it sounds interesting. If Janus is all you need, that's fine too.
MIT — see LICENSE.