Author: Med Karim Bchini @karimtn
Ports of jeffhub.ai use cases onto Google's EmbeddingGemma-2 embedding model, with a runnable demo and a latency + quality performance benchmark.
jeffhub ships a 0.8B "System 1" decider model plus one small adapter per task; each adapter takes an input and a list of options and returns a probability for every option. This project keeps that request shape ("pick one of your options, with a probability") but swaps the 0.8B decider for EmbeddingGemma-2 embeddings: each option is scored by cosine similarity to a few labelled exemplars (classification) or by dense passage ranking (retrieval), and a softmax turns the scores into probabilities.
The runtime here has no Python, no GPU, and a small disk budget, so the examples use:
- Node.js + @huggingface/transformersv4 (ONNX Runtime, CPU).
- onnx-community/embeddinggemma-2-ONNXat int8 (- dtype: "q8"→- model_quantized.onnx, ~314 MB).
EmbeddingGemma-2 is a 740M-parameter model (270M text backbone) that maps text into a unified 768-d space. It is task-steered: a short instruction prefix tells it what representation to produce. The prefixes used here come straight from the model card:
npm install # installs transformers.js + onnxruntime-node
npm run demo # runs every use case on sample inputs
npm run bench # latency + quality benchmark (writes bench/results.json)
npm run bench:quick # smaller timing loopsThe first run downloads the ONNX weights (~850 MB, all three encoders) into
./models/, which is git-ignored. Later runs load from disk in ~1–2 s
(warm page cache).
src/model.js EmbeddingGemma-2 loader, task prefixes, mean pooling, MRL truncation
src/knn.js nearest-exemplar classifier -> per-option probabilities (softmax)
src/similarity.js cosine / dot / softmax / ranking helpers
src/metrics.js accuracy, macro-F1, Recall@k, MRR, percentile
src/usecases.js the four use cases (ground, support-intents, spam, triage)
src/index.js public re-exports
data/*.json small labelled datasets (train exemplars + held-out test)
examples/demo.js runnable end-to-end example
bench/perf.js latency + throughput + quality harness
jeffhub use case: ground (Retrieval — pick the passage that answers)
Q: Which planet is known as the Red Planet?
1. [0.8068] d1 Mars, known for its reddish appearance, is often referred to as ...
2. [0.7056] d3 Jupiter is the largest planet in the solar system, with a ...
3. [0.6650] d2 Venus is often called Earth's twin because of its similar size ...
jeffhub use case: spam
Input: Your account has been suspended, verify your password immediately at the link below.
-> phishing (confidence 99.7%)
options: phishing=99.7% spam=0.2% legitimate=0.1%
Measured on this machine (CPU-only, int8, single process). Numbers vary a little run-to-run with machine load.
Embedding latency (batch = 1)
p50 168 ms | p95 177 ms | mean 168 ms (n=30)
Embedding throughput (batched)
24 docs/s (128 docs in 5.36 s, 41.8 ms/doc)
Per-use-case quality + end-to-end latency
use case | n | quality | p50 ms | p95 ms
------------------+----+-------------------------------------------+--------+-------
ground | 17 | Recall@1 100.0% | Recall@5 100.0% | MRR 1.000 | 166.9 | 171.5
support-intents | 13 | accuracy 100.0% | macro-F1 100.0% | 184.0 | 233.2
spam | 13 | accuracy 100.0% | macro-F1 100.0% | 170.7 | 185.2
triage | 12 | accuracy 100.0% | macro-F1 100.0% | 163.6 | 169.2
How to read this honestly:
- Latency/throughput are the meaningful performance numbers here: ~170 ms per single-text embedding on CPU, ~24 docs/s batched. That is the price of a 740M model on CPU with no GPU.
- Quality is at ceiling (100%) because the datasets are small and curated, and the held-out rows are paraphrases of the exemplars. This shows the ports work end-to-end; it is not a competitive accuracy claim. A rigorous claim needs a real labelled benchmark (e.g. an MTEB retrieval subset or genuine support tickets).
- Pooling / normalization: mean pooling, L2-normalized, so dot product == cosine.
- Classification: each option is scored by its nearest labelled exemplar (cosine), then a softmax (temperature 0.05) yields the per-option probabilities jeffhub returns.
- Retrieval: documents are embedded with the document prefix, queries with the retrieval-query prefix, then ranked by cosine.
- Matryoshka (MRL): embed(texts, { dim: 256 })truncates to 128/256/512-d and re-normalizes, trading a little quality for a 3–6× smaller vector store.
- No GPU here, so throughput is CPU-bound; the same code runs on WebGPU by
passing device: "webgpu"toloadModel.
- The ONNX build loads all three encoders (text + vision + audio). A text-only
deployment could ship just onnx/model_quantized.*to save space.
- Datasets are tiny by design (they live in data/), meant to demonstrate the pipeline rather than to benchmark the model.