Skip to content

Blog

EmbeddingGemma 2 local search: building Finer on a local embeddings API

By Saif Qureshi21 min read
  • Embeddings
  • Semantic search
  • Local AI
  • MCP

I write my search queries in English. Many of the documents I need are in German, and some are scans with no text layer. Hosted search means uploading every file, so I built Finer, a local search engine for my Mac. Its core is EmbeddingGemma 2, released by Google on 6 October 2026 and used here as a local embeddings API: text in, 768 numbers out, with a child process on my machine as the endpoint.

On 28 questions about my own files, hybrid search reached 0.90 MRR and a gated reranker 0.95. Those are README figures from my M4 Pro, on a small set the weights were tuned on, so read them as in-sample, and the repo records no embedder name next to them (the table notes explain). Finer was built under the working name MySearch, and the code, the mse command and the repo still use that name.

Key takeaways

  • One model and one 768-dimensional space cover text, images, PDF pages, audio and video, in 100+ languages. That replaces a text embedding API plus image and audio models.
  • README figures from my M4 Pro on 28 questions, with the embedder inferred to be EmbeddingGemma 2 8-bit: keywords alone 0.63 MRR, hybrid search 0.90, hybrid plus a reranker 0.95. The weights were tuned on those questions, so the numbers are in-sample, and 0.82 to 0.89 hit@1 is two questions.
  • The reranker runs only on near-ties. My first version took about 3 seconds and scored 0.88 MRR. The current one fuses ranks, adds 0.25 to 0.45 seconds and scores 0.95 on the same 28 questions.
  • Finer is not fully local. Search, embedding, reranking, OCR and transcription run on the Mac. The optional Ask feature is off by default; when it is on, a cloud model writes the answer unless you choose the local model, and the masking of identifiers before that is regex-only.
  • Hosted embeddings are cheap per token. The case for local is privacy, offline use, no per-query bill, no lock-in and files that stay on disk.

Why I built Finer

My documents are a mess in the usual way: PDFs, scans, Word files, e-mails, screenshots, voice notes, in English and German. I search in English and often need the German file.

Finer watches the folders I choose and indexes what is in them. I use it from a Mac app, a command line, a web page and, through MCP, from coding agents. SQLite FTS5 does keywords, EmbeddingGemma 2 does meaning, a small bridge links English and German vocabulary, and weighted reciprocal rank fusion merges the lists. A Qwen3 reranker checks near-ties, and a nightly job re-tunes a few settings behind guardrails.

Two limits first. Finer is Mac-only: the model runtime is Apple's MLX and OCR uses Apple Vision. And it is not fully local: the optional Ask feature can send passages to a cloud model, described below. Where a part is Mac-only, I name what a Linux or Windows reader could swap in. Those are suggestions; Finer does none of it.

Why a local embeddings API works at this size

The size depends on which encoders you load: 270 million parameters for text only (the EmbeddingGemma 270m configuration), 440 million with images, 570 million with audio, 740 million for everything. I call it a local API because the contract matches a hosted endpoint. In Finer the endpoint is a worker process that reads JSON lines on stdin and replies on stdout, with the vectors as base64 float32. The calling code does not know there is no network.

Hosted embeddings are cheap per token. These are the list prices shown on the vendors' pages on 10 October 2026, per million text tokens:

  • OpenAI: text-embedding-3-small $0.02, text-embedding-3-large $0.13.
  • Google Gemini API: gemini-embedding-2 text $0.20, or $0.10 batch; at standard rates also $0.00012 per image, $0.00016 per second of audio and $0.00079 per video frame.
  • Voyage AI: voyage-4-large $0.12, voyage-4 $0.06, voyage-4-lite $0.02.

Why run it locally anyway

At $0.02 to $0.20 per million tokens, price is not the reason. Mine are these. Files stay on disk and the index stays on the machine. Ask is the exception: if you switch it on, passages go to a cloud model by default, as described below. Search works offline. There is no per-query bill, and no vendor can retire a model version under my vectors. Scans, photos and recordings become searchable without being uploaded. The cost is hardware, setup time and disk: about 1.3 GB of models for a minimal install, 4.4 GB for all of them. The repo publishes no query-embedding latency, so I give none.

Run EmbeddingGemma 2 with MLX on a Mac

I run the mlx-community 8-bit build, about 1.2 GB, through mlx-vlm. Its text encoder and audio projection are 8-bit, and the vision and audio towers stay in BF16. The card says to run inference in bfloat16 or float32, never float16, because the activations overflow float16's range. My README says 8-bit against BF16 gives cosine 0.9996 and a 99.6% identical top-10; no script in the repo reproduces that, and the README names no corpus. Finer also keeps a 320 MB text-only copy for search. That size is not a Google figure. The setup notes say the copy was cut from the full folder, and the method is not recorded.

A launchd service serves a local HTTP API on 127.0.0.1 behind a random token and runs search, the job queue and the file watcher. It never imports MLX. Each model lives in its own worker process: text embedding, full multimodal embedding, speech, reranking and Ask. A worker starts on first use, stops after 10 idle minutes (5 for Ask), and is killed and restarted if it hangs.

Stopping a process frees its memory, the service stays small (50 to 170 MB by my README), and a stuck model cannot take search down. The repo records no measurement of the alternatives, loading models in the service or an HTTP model server. One recorded comparison is in a worker code comment: the transformers tokenizer wrapper costs about 575 MB for Gemma's 262,144-token vocabulary, the raw tokenizers library about 260 MB. A Linux or Windows reader could try a torch or llama.cpp runtime. That is a suggestion; Finer does none of it.

The shape of the system

Read the diagram top to bottom. The first half runs when files change. The second half runs when I search.

Finer's pipeline in plain text.
text
files on disk
  |  file-system events (day)      nightly job 02:30, on AC power
  v
extract   text layer, Office, e-mail  |  OCR, page images, Whisper (night)
  v
chunk     ~1100 characters, 150 overlap, plus one "title" item per file
  |---> FTS5 table (name x4, heading x2, text x1)
  v
embed     EmbeddingGemma 2 worker, 768-d, cached by prompt hash
  v
SQLite    float16 vectors + 1-bit copy, one database per folder

query --> parse --> embed query --+--> keyword list   (FTS5 bm25)
                                  +--> bridge list    (EN/DE terms, FTS5)
                                  +--> meaning lists  (1-bit shortlist x16, exact re-score)
          weighted RRF (k = 60) --> files --> fold copies --> favourites
          near-tie?  yes --> Qwen3-Reranker, top 10 x 128 tokens --> rank fusion
          results: app, CLI, web, MCP        Ask: optional, off by default

Ingestion, OCR and chunking: what gets embedded

Extraction is split by cost. Light work runs by day: text, Office files, e-mail, the PDF text layer. Heavy work runs at night or on a manual sync, on AC power: OCR, page images, audio, video, transcription. Up to five new scans or photos get their OCR and picture embedding at once. Unchanged files are never opened.

A PDF page with fewer than 40 non-space characters counts as a scan and goes through Apple Vision OCR. Page 1 and every scanned page are also rendered and embedded as pictures, up to 30 per file, so a scan is findable by its OCR text and by how the page looks. Audio is cut into 60-second segments and video into 30-second clips, and Whisper large-v3-turbo transcribes speech. Files that look like secrets (keys, .env files, credential folders) are never read. Mac-only here: Apple Vision, PDFKit, textutil, file-system events and launchd. A Linux or Windows reader could try Tesseract or EasyOCR with pdfium, pandoc, watchdog, faster-whisper and systemd. Those are suggestions; the repo has no benchmark of Tesseract against Vision.

I chunk by characters: about 1,100 per chunk, texts up to 1,600 kept whole, 150 of overlap. Markdown splits by heading, PDFs by page, slides by slide, sheets by sheet. Counting characters needs no tokenizer, and 1,600 characters will almost always fit under the 2,048-token cut Finer applies. The repo records no comparison with token-based or semantic chunking.

Documents are embedded in the format the card prescribes, "title: {title} | text: {content}", and queries as "task: search result | query: {query}". The card says leaving the task prefix out may lower quality. Each file also gets a title item made of its path words, so a scan with poor OCR can still be found by name. A cache keyed by a hash of the prompt reuses vectors for identical prompts; my README says re-embedding an unchanged folder takes 6 seconds instead of 320, again with no script behind it. Copies of a file are all indexed and fold into one "N versions" result at query time.

Hybrid search: keywords plus embeddings, with numbers

The keyword side is one FTS5 table with name, heading and text columns, diacritics removed, a prefix index and bm25 weights of 4, 2 and 1, so a hit in the file name counts four times a hit in the body. A query becomes an OR of its words, and words of four letters or more also match as prefixes.

The table after these notes is from my README: 28 hand-written questions about one private folder of a few hundred business documents in English and German, run with mse eval on my M4 Pro. The weights were tuned on these same questions, so every row is in-sample. The repo records no embedder name or revision with the table. I attribute it to EmbeddingGemma 2 8-bit by timeline: the project's files date from 7 October 2026, the day after the model came out, and name no other embedder. Read it with four caveats:

  • Keywords alone scored 0.63 MRR. That run had no bridge, and the gold set's own description says English questions mostly target German documents, so it says little about BM25 on a single-language folder.
  • "Meaning only" includes the bridge. Hybrid beats it by 0.03 MRR with the same hit@1, a small edge on 28 questions.
  • The README describes the current default as tuned weights, a penalty on overview files and version grouping together, so I cannot say which did the work.
  • One question is 3.6 points of hit@k. Going from 0.82 to 0.89 hit@1 is two questions.
Finer README quality table. 28 questions, one private folder, in-sample.
text
Setting                                  Hit@1   Top 3   Top 10   MRR
First version, equal weights              0.68    0.89     0.96   0.81
Current default                           0.82    0.96     1.00   0.90
Meaning only (vectors + bridge)           0.82    0.93     0.96   0.87
Keywords only (FTS5 bm25)                 0.57    0.64     0.79   0.63
Current default + near-tie reranker       0.89    1.00     1.00   0.95
First reranker attempt (replaced)         0.79    0.96     1.00   0.88

Merging the lists with rank fusion

Keyword scores and cosines live on different scales, so I fuse ranks. Cormack, Clarke and Büttcher (SIGIR 2009) score a document as the sum of 1 / (k + rank) over the lists, with k = 60 fixed in a pilot. They call 60 near-optimal and say the choice was not critical. I kept it and added a weight per list, so each hit contributes weight / (60 + rank). That is weighted reciprocal rank fusion, and the build guide shows it in Python. The default weights are meaning 1.0, keyword and bridge 0.6, titles 0.55, pictures 0.75 and audio 0.6, plus a small 0.004 bonus for the top keyword and text-vector hit.

Chunk scores roll up to files as best + 0.3 × second + 0.1 × third. Files named like a summary, readme, index or to-do list are multiplied by 0.6, because they tend to mention everything. The repo records no comparison of RRF with a linear blend for the retrieval lists, so I report none. The nightly job tunes the keyword, bridge and title weights, that penalty and two reranker settings. The rest, including k, stays fixed.

Searching German files with English questions

FTS5 cannot cross languages: an English query has no German words to match. The bridge uses each folder's own vocabulary. At night Finer embeds up to 25,000 of the folder's most frequent terms, with the query prompt, into a cosine table. At query time it takes the 12 terms nearest to the query's embedding, drops any below cosine 0.62 and the query's own words, keeps at most five, and runs them as an extra FTS5 query at weight 0.6. The interface shows them as "also searched". For an English query like "rental contract" it might add "mietvertrag" or "miete"; that example is illustrative and does not come from a real index.

The vocabulary carries the folder's jargon, product names and German compounds, with no translation model, dictionary or network. Translating the query with a model, using a dictionary or letting an LLM rewrite it are the alternatives. Finer uses none of them, and the repo has no comparison.

The evidence is thin. A no-bridge test config exists, but I have no published number for what the bridge adds. A code comment dated 7 October records query-to-query cosines of 0.75 to 0.86 for English and German versions of one question, against up to 0.889 for different questions; no script reproduces them. That measures query similarity and says nothing about retrieval. The mechanism should work for any language pair the model covers; the stop-word lists cover English and German only.

1-bit vectors with exact re-scoring: the decision

Vectors live in the same SQLite file as the text. Each chunk stores 768 values as float16 for exact scoring, plus a copy keeping only the sign of each dimension, packed into 96 bytes in a sqlite-vec table. A search asks that table for the k × 16 nearest chunks by Hamming distance, 960 candidates at the default k of 60, then re-scores them with exact cosine on the float16 vectors. Storing finished unit vectors as float16 is fine: the cast happens after inference.

This saves no disk, since both copies are stored, about 1.6 KB per chunk. It buys speed. My README says a 99% identical top-10 against exact search at about 10× the speed, with no script behind it.

Qdrant published numbers for EmbeddingGemma 2 (Qdrant says it had early access from Google DeepMind). On five text retrieval datasets, its 1-bit TurboQuant on all 768 dimensions kept 99% of float32 nDCG@10 at 30× less vector RAM without rescoring. In the paragraph on its separate Quora run (523k documents) it adds that, without rescoring, about 74% of the top 10 results matched exact search. Shorter 256-dimension vectors with rescoring (4× oversampling) kept 94.5% at 77× less. That is Qdrant's quantiser on text benchmarks, and Finer's sign bits inside a hybrid system are a different test. It supports the idea; it does not prove the README number.

I do not use Matryoshka truncation. The card supports 768, 512, 256 and 128 dimensions and calls quality close to lossless down to 256. Finer stores all 768, which is unused headroom, not a recorded decision. Testing 256 on your own files, and re-normalising after cutting, would be the first step.

Is a reranker worth it? Only for near-ties

Rank fusion merges retrieval lists. A reranker is a different job: it re-reads the top documents. Qwen3-Reranker-0.6B is Apache 2.0. It reads the query and one document in a single prompt and answers yes or no, and the score is the probability of yes. I run an 8-bit MLX conversion with a custom instruction.

My first attempt blended that probability into the retrieval score over 20 documents of 640 tokens. It took about 3 seconds a query and scored 0.79 hit@1 and 0.88 MRR, below plain hybrid at 0.90. The probabilities saturate: most plausible files land between 0.95 and 0.999, so the raw value adds noise. The order still carried signal: alone, the reranker put the right file first about as often as hybrid search, on different queries.

So the current version fuses ranks. A document at retrieval rank i and reranker rank j scores 0.75 / (10 + i) + 1 / (10 + j), so the reranker wins ties. It reads only the top 10, at 128 tokens each, and only when the runner-up scores at least 0.8 of the leader. A comment in the code says that on the 28-question gold set this skips about half of all queries, and hybrid search was right on every skipped one.

Result: 0.89 hit@1, 1.00 top-3, 0.95 MRR, for 0.25 to 0.45 seconds more when it runs, cached for repeats. That is still two questions, on a set I tuned on. The app shows the hybrid list at once and re-orders a moment later if the order changed.

Favourites, and a LoRA fine-tune of the reranker

When I confirm the right file for a query, I star it, and for that exact query it comes first from then on. A similar query gets a boost only if the stored query embedding has cosine 0.90 or more with the new one and the file was retrieved anyway. The 0.90 comes from query-to-query cosines recorded in a code comment, with no script behind them: rewordings of one question sit at 0.94 to 0.99, different questions reach 0.889. Evaluation runs with favourites off, so they cannot grade themselves.

Favourites also train the reranker, with a LoRA adapter (rank 8, last 16 layers) that learns from each starred query against hard negatives. It runs on AC power, only after 300 favourites over 100 distinct queries, and never trains on gold-set questions. An adapter is promoted only if held-out favourites gain 0.02 MRR, no gold set loses more than 0.01, and median rerank time grows by at most 10%. Finer tunes the reranker and leaves the embedder alone. A likely reason is that a new embedder means re-embedding every file; the repo states no reason and has no comparison.

The only recorded run is a test fixture: held-out favourites MRR rose 0.113, one gold set rose 0.015, the other fell 0.049, and the guardrail refused it. Treat the fine-tune as a mechanism with one test fixture. The repo does not show whether an adapter is active.

A nightly loop that tunes itself, behind guardrails

After the night sync, a learn job builds known-item questions from the index (the cleaned title or file name of up to 150 recent files per folder), scores the current settings on four sets (two hand-written gold sets, the favourites, the known items at half weight), and tries up to 39 settings on a small grid over six knobs.

A candidate is applied only if weighted MRR rises by at least 0.01, no hand-written gold set loses more than 0.01 MRR, no gold set loses the rank-1 hit of any question, and the favourites do not get worse. Each run writes one JSON line, and one command rolls back the last change.

A second gold set was written to check generalisation, but the nightly job scores it too, so it is not a clean hold-out, and no number for it is published. The gold sets stay private, so nobody else can reproduce the table.

Giving AI agents access through MCP

Finer speaks MCP with eight tools: search, get, multi_get, similar, collections, mark_best, sync and ask. By default an agent searches the folder holding its working directory, and one command writes a server entry locked to a folder into the configs of Claude Code, Cursor, Codex, Grok and OpenCode. With MSE_LOCK=1 the session may use only that folder. Anything else is refused, and if the folder is in no collection, every call fails. Results are checked again on the way out, and identifiers are masked by default.

The limits: the HTTP /mcp endpoint inside the service is token-gated but not locked. mark_best, which changes ranking, and sync, which queues re-indexing, are writes an agent inside the lock can call.

Ask: where the cloud comes in

Ask is off by default, and while it is off no model is called or loaded. When it is on, Finer searches with the reranker, takes the best passage from each top document (up to 8 passages, about 3,000 tokens) and numbers them. The prompt says to use only those passages, cite [n] after each claim, and give one fixed sentence when they do not answer.

By default a cloud model writes the answer, Claude unless you pick another. The local Qwen3-4B-Instruct-2507 at 4-bit answers when you choose Local, when a folder in scope has cloud turned off, or, if you switch on local_fallback (off by default), when no cloud model is reachable. With the defaults, an offline Mac or a cloud call that fails before the first token gets an error and no local answer.

Before anything leaves, regexes mask IBANs, passport, ID, permit, tax and social-security numbers that follow their label words, dates of birth and machine-readable passport lines. Names and addresses are not masked, and the question is sent as typed. Metadata-only files are never read. On a 24-question test the local model cited the expected document in 15 of 17 answers where it reached the passages, against 14 of 17 for Qwen3.5-4B, and gave the fixed refusal sentence 4 of 4 times, against 0 of 4 (it refused in its own words). That is a tiny sample, and the cloud path is unmeasured against a live API.

Will it slow my Mac down? What it costs to run

These are README figures from my M4 Pro, except Ask's local model, which comes from the Ask model notes. No script in the repo produces the README figures. Memory is measured as physical footprint, because ps misses GPU buffers. The repo has no comparison of Finer's load with Spotlight's.

  • The service: 50 to 170 MB, about 0% CPU when idle.
  • The text embedding worker takes about 0.8 GB while loaded (with the text-only copy), the reranker about 0.9 GB, heavy work 1.5 to 2 GB. Embedding uses 30 to 50% of one CPU core. Ask's local model, measured: 2.46 GB loaded, 3.76 GB at peak while answering.
  • A changed file shows up in about 3.6 seconds, 3.0 of them settle delay. Models unload after 10 idle minutes, and night work runs on AC power at low priority.
  • Disk: about 1.3 GB of models for a minimal install, 4.4 GB for all of them.

What is rough or missing

  • Mac-only: MLX, Apple Vision, file-system events, textutil and launchd.
  • A bug in keyword search. The query code folds ß to ss, but the FTS5 index keeps ß, so by that code a query for "Straße" only finds documents that spell it "Strasse". The index side can be reproduced with the schema in the build guide; the full path has not been checked against a real Finer index. "Mueller" does not match "Müller" either.
  • No stemming or decompounding, and stop-words and language detection cover English and German only.
  • Thin evidence: 28 questions, one private folder, weights tuned on the same set, no held-out number, gold sets and result logs private, no recorded embedder name or revision next to the table, and no scripts behind the 8-bit, 1-bit, cache, live-indexing and memory figures.
  • Two pieces cannot be reproduced from the repo: it records neither how the 320 MB text copy was cut nor the command that converted the 8-bit reranker by hand, so a fresh setup cannot download the reranker. The minimal install path has not been run on a fresh Mac.
  • Ask's masking is regex only, and the per-folder cloud permission defaults to on once Ask is on. The app is ad-hoc signed, not notarised.

What to measure on your own files first

Before you copy any of my settings, write about 30 real questions from your own files, each with the file that should answer it, and score keyword-only, meaning-only and hybrid on them. Hold some questions back, because tuning on all of them gives in-sample numbers like mine. The build guide has the code, including MRR and hit@k.

The source code is linked under Sources. Known gaps, as the repo records them: no recipe for the 320 MB text copy or the 8-bit reranker, an install path that has not been run on a fresh Mac, and no scripts behind the figures above. I scope AI work for founders as milestones with evals; see how I work with founders. At Parker AI I designed the hybrid retrieval layer, with query-aware re-ranking.

Frequently asked questions

Is EmbeddingGemma open source, and can I use it commercially?

Only facts, no legal advice. Google's model card and launch post say Apache 2.0 for version 2, and the card says deployments must adhere to the Gemma Prohibited Use Policy. Version 1 used the Gemma Terms of Use. I found no Google page saying how the two fit together, so ask Google or a lawyer if it affects your product.

How to use EmbeddingGemma?

Load it with sentence-transformers or transformers, put "task: search result | query: " before queries and "title: {title} | text: {content}" before documents, and run in bfloat16 or float32, never float16. Google's launch post names transformers, sentence-transformers, vLLM and llama.cpp among the runtimes, so it is not Mac-only; ggml-org publishes a GGUF build for llama.cpp. The build guide has code.

Does reranking even make sense?

On my 28 questions, a reranker that blended its probability into every query scored 0.88 MRR against 0.90 for plain hybrid. A version that runs only on near-ties and fuses ranks reached 0.95. Several things changed between the two (the gate, the fusion, 10 documents instead of 20, 128 tokens instead of 640), so I cannot say which one did the work. The gain is two questions on a set the weights were tuned on, so test it on yours.

Is EmbeddingGemma 2 the same as Gemma 2B or Gemini Embedding 2?

No. EmbeddingGemma 2 (also written embedding gemma 2 or gemma embedding 2) is a Google DeepMind embedding model with open weights, 740 million parameters, built on Gemma 4. Gemini Embedding 2 is a hosted model on Google's Gemini API pricing page.

Is Finer fully local?

No. Search, embedding, reranking, OCR and transcription run on the Mac. The optional Ask feature is off by default. When it is on, a cloud model (Claude by default) writes the answer from the top results, with personal identifiers masked first. A local Qwen3-4B writes it only if you choose Local, a folder in scope has cloud turned off, or you enable local_fallback (off by default); with the defaults an offline Mac gets an error. The masking is regex only: names and addresses are not masked, and the question is sent as typed.

Sources

  1. EmbeddingGemma 2 launch post (Google) blog.google
  2. EmbeddingGemma 2 model card (Hugging Face) huggingface.co
  3. EmbeddingGemma 2 model card (ai.google.dev) ai.google.dev
  4. EmbeddingGemma 1 model card (Hugging Face) huggingface.co
  5. Gemma Prohibited Use Policy ai.google.dev
  6. mlx-community 8-bit conversion of EmbeddingGemma 2 huggingface.co
  7. llama.cpp GGUF build of EmbeddingGemma 2 (ggml-org) huggingface.co
  8. Qdrant: How small can Google's new EmbeddingGemma 2 get? qdrant.tech
  9. Qwen3-Reranker-0.6B model card huggingface.co
  10. Qwen3-Embedding-0.6B model card huggingface.co
  11. Cormack, Clarke, Büttcher: Reciprocal Rank Fusion outperforms Condorcet and individual Rank Learning Methods (SIGIR 2009) research.google
  12. SQLite FTS5 documentation sqlite.org
  13. MLX, Apple's array framework for Apple silicon github.com
  14. OpenAI API pricing (embeddings) developers.openai.com
  15. Gemini API pricing (gemini-embedding-2) ai.google.dev
  16. Voyage AI pricing docs.voyageai.com
  17. Finer source code github.com

Talk to me

Building an AI product and need someone senior to own the technical side? Book a free 30-minute call.

Have something to build?

Tell me what you're building — start with a free call

Send a message

Founder, SolutionPlus · AI Product Engineer

SQ
Saif Qureshi
  • Berlin, Germany · Production AI agents and systems for companies and enterprises
  • Outcomes-focused delivery: measurable impact, not demos.

Contact

Available for new projects
© 2026 Made withby Saif Qureshi
ImpressumReact · TypeScript · Tailwind CSS