Author: Med Karim Bchini @karimtn

Ports of jeffhub.ai use cases onto Google's EmbeddingGemma-2 embedding model, with a runnable demo and a latency + quality performance benchmark.

jeffhub ships a 0.8B "System 1" decider model plus one small adapter per task; each adapter takes an input and a list of options and returns a probability for every option. This project keeps that request shape ("pick one of your options, with a probability") but swaps the 0.8B decider for EmbeddingGemma-2 embeddings: each option is scored by cosine similarity to a few labelled exemplars (classification) or by dense passage ranking (retrieval), and a softmax turns the scores into probabilities.

The runtime here has no Python, no GPU, and a small disk budget, so the examples use:

- Node.js + @huggingface/transformersv4 (ONNX Runtime, CPU).

- onnx-community/embeddinggemma-2-ONNXat int8 (- dtype: "q8"→- model_quantized.onnx, ~314 MB).

EmbeddingGemma-2 is a 740M-parameter model (270M text backbone) that maps text into a unified 768-d space. It is task-steered: a short instruction prefix tells it what representation to produce. The prefixes used here come straight from the model card:

npm install # installs transformers.js + onnxruntime-node

npm run demo # runs every use case on sample inputs

npm run bench # latency + quality benchmark (writes bench/results.json)

npm run bench:quick # smaller timing loopsThe first run downloads the ONNX weights (~850 MB, all three encoders) into

./models/, which is git-ignored. Later runs load from disk in ~1–2 s

(warm page cache).

src/model.js EmbeddingGemma-2 loader, task prefixes, mean pooling, MRL truncation

src/knn.js nearest-exemplar classifier -> per-option probabilities (softmax)

src/similarity.js cosine / dot / softmax / ranking helpers

src/metrics.js accuracy, macro-F1, Recall@k, MRR, percentile

src/usecases.js the four use cases (ground, support-intents, spam, triage)

src/index.js public re-exports

data/*.json small labelled datasets (train exemplars + held-out test)

examples/demo.js runnable end-to-end example

bench/perf.js latency + throughput + quality harness

jeffhub use case: ground (Retrieval — pick the passage that answers)

Q: Which planet is known as the Red Planet?

1. [0.8068] d1 Mars, known for its reddish appearance, is often referred to as ...

2. [0.7056] d3 Jupiter is the largest planet in the solar system, with a ...

3. [0.6650] d2 Venus is often called Earth's twin because of its similar size ...

jeffhub use case: spam

Input: Your account has been suspended, verify your password immediately at the link below.

-> phishing (confidence 99.7%)

options: phishing=99.7% spam=0.2% legitimate=0.1%

Measured on this machine (CPU-only, int8, single process). Numbers vary a little run-to-run with machine load.

Embedding latency (batch = 1)

p50 168 ms | p95 177 ms | mean 168 ms (n=30)

Embedding throughput (batched)

24 docs/s (128 docs in 5.36 s, 41.8 ms/doc)

Per-use-case quality + end-to-end latency

use case | n | quality | p50 ms | p95 ms

------------------+----+-------------------------------------------+--------+-------

ground | 17 | Recall@1 100.0% | Recall@5 100.0% | MRR 1.000 | 166.9 | 171.5

support-intents | 13 | accuracy 100.0% | macro-F1 100.0% | 184.0 | 233.2

spam | 13 | accuracy 100.0% | macro-F1 100.0% | 170.7 | 185.2

triage | 12 | accuracy 100.0% | macro-F1 100.0% | 163.6 | 169.2

How to read this honestly:

- Latency/throughput are the meaningful performance numbers here: ~170 ms per single-text embedding on CPU, ~24 docs/s batched. That is the price of a 740M model on CPU with no GPU.

- Quality is at ceiling (100%) because the datasets are small and curated, and the held-out rows are paraphrases of the exemplars. This shows the ports work end-to-end; it is not a competitive accuracy claim. A rigorous claim needs a real labelled benchmark (e.g. an MTEB retrieval subset or genuine support tickets).

- Pooling / normalization: mean pooling, L2-normalized, so dot product == cosine.

- Classification: each option is scored by its nearest labelled exemplar (cosine), then a softmax (temperature 0.05) yields the per-option probabilities jeffhub returns.

- Retrieval: documents are embedded with the document prefix, queries with the retrieval-query prefix, then ranked by cosine.

- Matryoshka (MRL): embed(texts, { dim: 256 })truncates to 128/256/512-d and re-normalizes, trading a little quality for a 3–6× smaller vector store.

- No GPU here, so throughput is CPU-bound; the same code runs on WebGPU by

passing device: "webgpu"toloadModel.

- The ONNX build loads all three encoders (text + vision + audio). A text-only

deployment could ship just onnx/model_quantized.*to save space.

- Datasets are tiny by design (they live in data/), meant to demonstrate the pipeline rather than to benchmark the model.