GENOME

by northtekdevs

Not rated
GitHub

About

Fully local memory for AI agents: zero LLM calls in the write path, runs air-gapped, accuracy parity with Mem0 on published benchmarks.

Details

Author
northtekdevs
Categories
AI

Setup

Install GENOME in your MCP client (Claude Desktop, Cursor, Windsurf, and others).

Repository: https://github.com/northtekdevs/genome

Follow the installation instructions in the repository README, then restart your MCP client.

Open memory for AI agents. Same answer accuracy as Mem0 - but ~1,000× cheaper to store, runs fully offline, and keeps an auditable record.

Papers:Do Agents Need an LLM to Remember?(the core evaluation, 2026) andWhat Does Each Memory Feature Buy?(a measured audit of all five optional features, wins and failures alike, 2026). PDFs inpapers/; result tables inbenchmarks/AUDIT-RESULTS.md.

Most agent-memory tools (like Mem0) call an LLM onevery messageto decide what to remember. That's the slow, expensive part - and GENOME's bet is that you don't need it. GENOME just embeds each message locally: no LLM, no API, no network in the write path.

Benchmarked honestly on public datasets (LoCoMo, LongMemEval), GENOMEanswers just as accurately as Mem0- while storing memories for a tiny fraction of the cost and running completely offline.

Honest up front:on answer accuracy, GENOMEtiesMem0 - we donotclaim to beat it there (six independent benchmark configurations confirm parity, none significant in either direction). The advantage is cost, speed, offline operation, and a temporal/auditable record Mem0 can't produce.

Every frame is real output fromexamples/demo_timeline.py, captured bytools/render_demo_gif.py. Run it yourself, no API key required:

The interesting part is step 3. The same question gets three different correct answers depending onwhenyou ask about, because the store keeps when each fact became true rather than overwriting it:

The "thinking about maybe moving to Denver, nothing decided" turn is stored but never becomes an answer: it is a plan, not a durable fact.

The write path is deliberately dumb and cheap. All the intelligence happens at read time, when there is a query to focus it.

flowchart LR M["incoming message"] --> E["local embedder<br/>all-MiniLM-L6-v2"] E --> S[("local store<br/>SQLite or Postgres")] M -. "optional, opt-in" .-> B["belief extraction<br/>(the only LLM call)"] B --> K[("bi-temporal<br/>fact log")] Q["query"] --> R["exact cosine search<br/>over this tenant's rows"] S --> R R --> RR["optional cross-encoder<br/>rerank"] RR --> A["context for the agent"] Q --> PIT["as-of resolution<br/>facts_valid_at(entity, T)"] K --> PIT PIT --> A style E fill:#0A84FF,color:#fff style S fill:#1c2530,color:#fff style K fill:#1c2530,color:#fff style B fill:#3a3a3a,color:#fff

Write: embed locally, store. About 10 ms, zero LLM calls, zero network calls. The embedding is deterministic -- the same text always yields the same vector, with no sampled extraction step deciding what matters -- so what gets stored is a function of the input, and replaying a journal reproduces that store exactly. (Ids and timestamps are stamped per write, so two independent ingests of the same conversation agree on content and vectors, not on record ids.)

Read: exact cosine search within the tenant's scope (no ANN index to build or update), with an optional local cross-encoder reranker.

Bi-temporal layer (opt-in): records each fact at itsdomain time, the moment it became true in the world, not the moment it was ingested. That is what makes point-in-time questions answerable even when facts arrive out of order.

flowchart TB subgraph LLM["LLM-extraction memory"] A1["message"] --> A2["LLM decides what matters<br/>(sampled, non-deterministic)"] A2 --> A3[("store")] A3 --> A4["replaying the same input<br/>can produce a different store"] end subgraph GEN["GENOME"] B1["message"] --> B2["local embedding<br/>(deterministic)"] B2 --> B3[("store")] B3 --> B4["replaying the same input<br/>reproduces the same store"] end style A4 fill:#5c1f1f,color:#fff style B4 fill:#1f4d33,color:#fff

A record that cannot be re-derived is difficult to audit. That property, not accuracy, is the actual argument for this design.

Don't believe it? Prove it yourself

Thecost, speed, and offlineclaims need no API key - measure them onyourmachine in 60 seconds:

git clone https://github.com/NORTHTEKDevs/genome && cd genome pip install -e . && python -m genome.verify

Thefirstrun downloads the local embedding model (~90 MB, one time) before printing anything, so expect 30-120 seconds of apparent silence on a cold machine. Every run after that is instant.

It writes memories with youroutbound network physically blockedand prints a live pass/fail receipt - 0 network calls, 0 LLM calls, single-digit-ms writes, retrieval that works:

[PASS] Air-gapped write path: wrote 200 memories with every outbound socket blocked -> 0 network attempts, 0 LLM calls [PASS] Write latency: 7.1 ms/message (Mem0's measured write path: ~2,055 ms + 1 LLM call/message) [PASS] Retrieval works: top hit score 0.598

That receipt covers the cost/speed/offline story only. Theaccuracy-parity with Mem0claim is a separate, larger check that needs an LLM key - reproduce it head-to-head on the same questions with your own key viapython benchmarks/head_to_head.py(one OpenRouter key works; seebenchmarks/RESULTS.mdfor the n=90 / n=205 runs, the paired significance tests, and the published nulls). The full test suite runs in public CI (badge above). The pitch isn't "trust me" - it's "run it."

Add persistent memory to your agent in one line (MCP)

GENOME ships afully-local MCP server- cross-session memory for Claude Desktop, Claude Code, or Cursor withno API key and no data leaving your machine:

pip install "genome-memory[mcp]"
{ "mcpServers": { "genome": { "command": "genome-mcp" } } }

Or zero-install via uv:{ "command": "uvx", "args": ["--from", "genome-memory[mcp]", "genome-mcp"] }

Tools the agent gets:remember,recall,forget,reset_memories. Memories persist locally in~/.genome/memories.db.Full MCP details ↓

Every number is measured within one harness - same responder, judge, embedder, and top-k; only the memory layer changes - with paired significance tests. Full detail and per-number provenance:benchmarks/RESULTS.md. Formatted report:benchmarks/GENOME-LoCoMo-Report.pdf.

Why it's ~1,000× cheaper: it never calls an LLM to remember

Storing one message costsone LLM call in Mem0, zero in GENOME(just a local embedding). That's not a benchmark you can argue with - it's arithmetic, and it holds no matter which LLM you price it against. At 10,000 users × 50 messages/day (15M messages/month):

The gap survives the cheapest model andgrowsin production (Mem0 re-sends stored memories to the LLM as the store fills). Reproduce:python benchmarks/tco_project.py(no API key).

GENOME's default embedder is local. We proved the write path is genuinely offline byblocking all network during writes- they still succeed:

- ~10 ms/message, 0 network calls, 0 LLM calls(python benchmarks/local_writepath.py)
- Mem0 can't do this - it needs an LLM API call to ingest.

That makes GENOME usable on-prem, in regulated environments, or fully offline. It's a yes/no capability, not a price point.

- Write:embed the message locally and store it. No LLM, no network. (~10 ms)
- Read:vector search over your memories, with an optional local cross-encoder reranker for harder queries.
- Optional bi-temporal layer:track how facts change over time and answer "what was true at time T" - see below.

Because nothing on the write path interprets your content, GENOME can do things an LLM-ingest memory system cannot do in principle:

-

Memory firewall(genome.firewall): tag every write with where it came from (user,agent,tool,web), quarantine low-trust origins from recall, and enforce origin-bound authority - web content can never UPDATE or DELETE what your user said, even when a prompt-injected conflict resolver asks for it. There is also no extraction step for injected content to attack: the write path has no LLM.

from genome import Memory from genome.firewall import TrustPolicy m = Memory(trust_policy=TrustPolicy(recall_min_trust=1)) m.add("I live in Anchorage", user_id="u1", provenance="user") m.add(scraped_page_text, user_id="u1", provenance="web") # quarantined

Explainable recall(genome.explain):explain_search()reports every candidate's dense score, BM25 rank, fused score, and - when it was not returned - the exact reason (parent-filtered, quarantined, beyond the limit). Two runs agree, so a recall bug can be committed as a regression test instead of a shrug.

Journal + replay(genome.journal): record every mutation and provably reproduce the store -verify_journal()replays the history and compares canonical hashes. Replay a prefix to roll back; replay into different storage to branch a memory for a what-if run. The journal sits after extraction, so replay is deterministic even if you configured an LLM extractor. Each line chains to its predecessor, so a removed or edited line is detected even when the change cancels out in the final state.

# Tamper-EVIDENT by default. Pass a key (kept outside the journal's directory) # to make it tamper-PROOF: an unkeyed chain can be recomputed by anyone with # write access, an HMAC chain cannot. m = Memory(journal="mem.journal", journal_key=os.environb[b"GENOME_JOURNAL_KEY"])

Multi-agent belief attribution(record_fact(..., believed_by="agent-a")): agents sharing a store keep their own belief timelines - agent B disagreeing does not clobber agent A's fact - andbelief_conflicts()surfaces disagreements for deliberate resolution instead of silently picking a winner.

A neutral benchmark harness(benchmarks/neutral/): run GENOME, Mem0, and a full-context baseline through the same responder, judge, and embedder, with a pairwise McNemar matrix and a full-disclosure block. GENOME is one row in the table, not the house.

The default embedder is local (sentence-transformers/all-MiniLM-L6-v2) - no API key, works offline; the first run downloads the ~90 MB model once. OpenAI embeddings are optional for higher-dimensional retrieval.

Dependency footprint, honestly:the core install isnumpy,sentence-transformers,scikit-learn, andrank-bm25. Local embeddings run on PyTorch (pulled in by sentence-transformers), so it isn't a tiny install - that's the deliberate tradeoff for offline, zero-cost embedding. Plotting/benchmark-chart deps live in an optional[viz]extra, not the core. Migrating from Mem0? Seedocs/migrating_from_mem0.md.

from genome import Memory mem = Memory(storage="genome.db") # local embedder by default; ":memory:" for ephemeral # Store a message -- embedded locally, no LLM call, no network mem.add("Ada met Lin at the robotics summit in Berlin.", user_id="u1") mem.add("They are collaborating on an open-source planning library.", user_id="u1") # Retrieve the most relevant memories for hit in mem.search("Where did Ada meet Lin?", user_id="u1", limit=5): print(f"{hit.score:.3f} {hit.content}")

Memorymirrors Mem0's API (add/search/get/delete/reset) - a near drop-in swap. To use OpenAI embeddings instead (setOPENAI_API_KEY):

from genome import Memory, EmbeddingProvider mem = Memory(storage="genome.db", embedding_provider=EmbeddingProvider(model_name="openai:text-embedding-3-small"))

Use it as an MCP server (fully-local memory for any agent)

GENOME ships an MCP server, so any MCP client (Claude Desktop, Claude Code, Cursor, ...) gets persistent cross-session memory that runsentirely on the local machine- no LLM calls, no API keys, no data leaves the box. Most memory MCPs can't say that.

Install with themcpextra, then add it to your client's config:

pip install "genome-memory[mcp]"
{ "mcpServers": { "genome": { "command": "genome-mcp" } } }

Tools the agent gets:remember(store a fact/preference, local + 0 LLM),recall(semantic search),forget(delete the memory matching a query),reset_memories(clear a user's memories). Memories persist in~/.genome/memories.db(override with theGENOME_MCP_DBenv var). Run standalone withgenome-mcporpython -m genome.mcp.server.

Prefer HTTP? GENOME ships a FastAPI server that mirrors the library 1:1 (add/search/get/update/delete/reset/synthesize), with an auto-generated OpenAPI spec at/docs.

pip install "genome-memory[fastapi]"

Try it locally(keyless, loopback only - one flag makes the "no auth" intent explicit):

GENOME_ALLOW_NO_AUTH=1 python -m genome.server # serves on 127.0.0.1:8080
curl -X POST localhost:8080/v1/memories \ -H 'Content-Type: application/json' \ -d '{"text": "Ada met Lin at the robotics summit in Berlin.", "user_id": "u1"}' curl -X POST localhost:8080/v1/search \ -H 'Content-Type: application/json' \ -d '{"query": "Where did Ada meet Lin?", "user_id": "u1", "limit": 5}'

Safe by default.The server refuses to serve unauthenticated unless you opt in as above, and it will not bind a non-loopback interface without a key. To expose it, set an API key (sent asX-API-Key) - required to bind beyond localhost:

GENOME_API_KEY=$(openssl rand -hex 32) GENOME_HOST=0.0.0.0 python -m genome.server # then add: -H "X-API-Key: $GENOME_API_KEY" to every request

For multi-tenant deployments, setGENOME_REQUIRE_SCOPE=1to requireuser_id/agent_idon every call and disable the global reset. Docker:docker-compose up(needsGENOME_API_KEYandPOSTGRES_PASSWORD; Postgres is published on loopback only). Full guide, including the Postgres backend and every env var:docs/tutorial_quickstart.md.

@northtek/genome-memorymirrors the PythonMemoryAPI shape against this server (ESM, Node 20+ or browser):

import { Memory } from "@northtek/genome-memory"; const mem = new Memory({ baseUrl: "http://localhost:8080" }); await mem.add({ text: "Ada met Lin in Berlin.", userId: "u1" }); const hits = await mem.search({ query: "Where did Ada meet Lin?", userId: "u1" });

Full client docs:sdks/typescript/README.md.

Same responder + judge + embedder for every system; only the memory layer changes.

What we tested thatdidn'thelp (so you don't have to)

We publish our nulls - it's how you know the wins are real:

- Synthesis / consolidation:accuracy-neutral at equal token budget (p = 0.86).
- Hybrid (BM25 + dense) and graph retrieval:hybrid underperformed plain dense on LoCoMo; graph was not validated here.
- Reranking's accuracy gain is embedder-dependent:it reliably improvesretrieval hit-rate, but its effect on finalanswer accuracydepends on the embedder - treat it as a retrieval-quality tool, not a guaranteed accuracy win.

Bi-temporal memory: "what was true at time T"

GENOME can track how facts change over time and answer point-in-time questions - something overwrite-based memory structurally can't do (it only keeps the latest value):

from genome.memory.belief import ingest_belief_turn, answer_belief_context mem = Memory(storage="genome.db", llm_call=my_llm_fn) # facts land at their DOMAIN time (parsed from the text), not wall-clock ingest time ingest_belief_turn(mem, "In March 2024, Jordan moved to Seattle.", session_time=t0, user_id="u") ingest_belief_turn(mem, "Jordan just moved to Austin.", session_time=t2, user_id="u") answer_belief_context(mem, "Where does Jordan live now?", user_id="u") # -> Austin answer_belief_context(mem, "Where did Jordan live in early 2024?", user_id="u") # -> Seattle answer_belief_context(mem, "List every city Jordan has lived in.", user_id="u") # -> Seattle; Austin

On the TempBelief benchmark it answers as-of queries at0.870vs Mem0's 0.676, with the knowledge graph audited at 0.97 precision / 0.96 recall.Caveat:TempBelief is synthetic text with explicit dates; the edge shrinks on natural speech. Real capability, bounded proof.

Opt-in; the default path stays LLM-free and local at ingest.

mem = Memory( storage="genome.db", llm_call=my_llm_fn, # LLM-based fact extraction on add() resolve_conflicts=True, # ADD/UPDATE/DELETE vs existing memories auto_extract_entities=True, # entity graph for graph retrieval auto_consolidate_threshold=200, # summarize-or-prune when a scope grows past N ) mem.search("...", user_id="u1", mode="hybrid") # modes: "dense" (default), "hybrid", "graph"
from genome.memory.rerank import CrossEncoderReranker mem = Memory(storage="genome.db", reranker=CrossEncoderReranker()) # lazy-loaded mem.search("Where did the user go on vacation?", user_id="u1", limit=5) # reranked

The LoCoMo and LongMemEval datasets arenot bundled(they carry their own licenses - LoCoMo is CC BY-NC 4.0). Seebenchmarks/data/README.mdto download them. The first two lines need no dataset and no API keys:

python benchmarks/local_writepath.py # local write path: ~10ms/msg, 0 network python benchmarks/tco_project.py # deployment cost projection python benchmarks/verdict.py # in-window accuracy + McNemar python benchmarks/haystack_report.py # overflow / context-window crossover python benchmarks/ingest_cost.py --n 80 # measured ingestion cost vs Mem0 python benchmarks/lme_qa.py --n 90 # LongMemEval head-to-head vs Mem0 python benchmarks/tempbelief_run.py --convs 6 # bi-temporal point-in-time vs baselines
No reviews yet — be the first

Sign in to leave a review

Use Google, GitHub, or an email account so ratings stay tied to real people.

Email sign in

No reviews posted yet.