When exploring complex problems, traditional linear dialogue architectures introduce two fundamental bottlenecks:
- Context Contamination and Attention Dispersion (Lost in the Middle): Interleaving exploratory side quests, syntax corrections, and dead ends directly into the main thread rapidly exhausts the finite context window, diluting essential signals.
- Lack of Persistent Reasoning Depth: Conventional RAG relies strictly on semantic similarity across isolated text snippets, lacking an overarching derivation chain and structural reasoning progression.
GGUF Libra addresses these challenges through a cognitive-inspired design:
- Interaction Layer: Abstracts conversation history into a directed tree, providing Git-like branching (Fork), lossless distillation (Squash & Merge), and lifecycle archiving (Archive).
-
Retrieval Layer: Employs spectral graph theory (Laplacian matrix eigenvalues) to mathematically quantify the reasoning depth of a dialogue tree ($\text{Logic } T$ ), generating topological memory maps.
- Compute Layer: Runs entirely on local hardware with direct, slot-level control over the inference engine and physical KV cache.
Transitions conversational context from a linear log to a version-controlled tree:
- Side Chat: Branch off from any point to explore sub-problems in an isolated sandbox. Vector database querying is cleared on branching to prevent contextual crosstalk.
- Squash & Merge: Once an exploratory topic concludes, the branch is distilled into a concise summary node grafted back onto the main trunk. Verbose intermediate steps are pruned from the active context while remaining safely archived.
- Lifecycle Archiving: Converts mature conversation chapters into compressed seeds, resetting the active token budget while retaining continuity.
Replaces superficial text embeddings with structural topological metrics:
-
Graph Laplacian Construction: Deconstructs conversation trees into adjacency matrices to formulate the unnormalized Laplacian operator:
$$L = D - A$$ (where $D$ is the degree matrix and $A$ is the adjacency matrix).
-
Algebraic Connectivity and Relaxation Time ($\text{Logic } T = 1 / \lambda_2$ ): Solves for the second smallest eigenvalue (Fiedler value$\lambda_2$ ). Deep, sequential derivation chains exhibit lower algebraic connectivity and significantly higher relaxation times ($\text{Logic } T$ ), granting them highest priority during retrieval indexing (data.bin).
-
Derivation Chain Injection: Retrieves and prepends structured causal steps (Step n: Based on [A] -> Derived [B]) directly into the prompt to reinforce deductive momentum.
- Leverages Tree-sitter for native syntax tree parsing, chunking source code along natural semantic boundaries (functions, classes, methods) rather than arbitrary character or token counts.
- Currently supports fine-grained indexing for .cpp,.py,.go, and.jssource files.
- Physical KV Cache Erasure: Initiating a new session dispatches explicit purge requests to the underlying llama.cppinstance, evicting the slot's physical KV cache directly from GPU VRAM.
- Context Capacity Monitoring: Continuously polls slot utilization via /props(n_past / n_ctx), raising visual alerts when memory usage exceeds 85%.
GGUF Libra relies on llama.cpp for local inference. Two concurrent services are required:
- Chat Model: Handles conversational responses and summarization (e.g., Gemma 4 E4B on port 8021).
- Embedding Model: Handles vectorization for RAG retrieval (e.g., Embedding Gemma 300M on port 8022).
Use the following Python script to launch both services in the background:
import subprocess
import os
import sys
def main():
os.system("")
# Working directory (adjust to your local llama.cpp path)
work_dir = r"D:\llama_cpp"
if not os.path.exists(work_dir):
print(f"Error: Directory not found: {work_dir}")
sys.exit(1)
log_file1_path = os.path.join(work_dir, "llama_main_8021.log")
log_file2_path = os.path.join(work_dir, "llama_embedding_8022.log")
# Service 1: Main chat model (Port 8021)
cmd1 = [
"llama-server.exe",
"-m", r"models\gemma-4-E4B-it-Q5_K_M.gguf",
"--mmproj", r"models\gemma-4-e4b-mmproj-F16.gguf",
"-ngl", "99",
"-c", "32768",
"--port", "8021",
"--host", "0.0.0.0",
"--cache-ram", "512",
"--cache-type-k", "q8_0",
"--cache-type-v", "q8_0",
"--keep", "0",
"-fa", "on",
"-np", "1",
"--reasoning", "off",
"--repeat-penalty", "1.15",
"--mlock"
]
# Service 2: Embedding model (Port 8022)
cmd2 = [
"llama-server.exe",
"-m", r"models\embeddinggemma-300M-Q8_0.gguf",
"-ngl", "99",
"-c", "8192",
"--port", "8022",
"--host", "0.0.0.0",
"--embedding",
"--pooling", "mean"
]
print("Starting llama.cpp services...")
print(f"[Service 1] Main model: http://127.0.0.1:8021")
print(f"[Service 2] Embedding : http://127.0.0.1:8022")
print("Loading models into VRAM...")
try:
log1 = open(log_file1_path, "w", encoding="utf-8")
log2 = open(log_file2_path, "w", encoding="utf-8")
process1 = subprocess.Popen(cmd1, cwd=work_dir, stdout=log1, stderr=subprocess.STDOUT)
process2 = subprocess.Popen(cmd2, cwd=work_dir, stdout=log2, stderr=subprocess.STDOUT)
print("Both services are running in the background.")
print("Press Ctrl + C to terminate both services.\n")
process1.wait()
process2.wait()
except KeyboardInterrupt:
print("\nShutting down services...")
if 'process1' in locals(): process1.terminate()
if 'process2' in locals(): process2.terminate()
print("Services shut down successfully.")
finally:
if 'log1' in locals() and not log1.closed: log1.close()
if 'log2' in locals() and not log2.closed: log2.close()
if __name__ == "__main__":
main()Run the compiled executable gguf-libra.exe (or go run . during development) and navigate to:
http://127.0.0.1:8099
Use the action bar beneath messages to dynamically manage conversation flow:
- Side Chat: Branch into a focused sandbox.
- Summary: Generate an editable summary of the active node.
- Merge: Condense and graft branch insights back to the mainline.
- Archive: Compress the current thread into a foundational seed for the next chapter.
To index high-value reasoning trees into the persistent memory store:
- Navigate to the parser directory:
cd parser_py uv run main.py
- In the interactive terminal interface:
- Run save: Computes the Graph Laplacian eigenvalues of the active conversation, calculates the topological relaxation weight ($\text{Logic } T$ ), and persists the sorted derivation tree todata.bin.
- Run data: Inspects topological weights and metadata across all indexed trees.
- Run
From the repository root directory, execute:
# 1. Install frontend dependencies and bundle static distribution
npm install
npm run build
# 2. build parser py
cd parser_py
uv sync
cd ..
# 3. Compile the Go backend binary
go build -xThe resulting gguf-libra.exe runs self-contained alongside the generated dist/ directory.