Arcaeon Ledger
About
A tamper-evident, hash-chained action log for AI agents, plus a stdio proxy that records tool calls at the seam so the record does not depend on the agent's cooperation.
Details
- Author
- dan8433-user
- Categories
- Developer Tools
Jump to
Setup
Install Arcaeon Ledger in your MCP client (Claude Desktop, Cursor, Windsurf, and others).
Repository: https://github.com/dan8433-user/ledger
Follow the installation instructions in the repository README, then restart your MCP client.
Observability tools show you what your agent did.arcaeon-ledgerlets youproveit.
Every record is hash-chained to the one before it. Edit a row, delete one, or reorder history, and every later link breaks —verifynames the exact line. You own the record, and you can prove it wasn't altered. Zero dependencies, one JSONL file, two verbs.
pip install arcaeon-ledger # then: from arcaeon_ledger import Ledger
from arcaeon_ledger import Ledger log = Ledger("agent.log.jsonl") log.append({"tool": "web.search", "query": "weather in LA", "result_ok": True}) log.append({"tool": "payment", "amount": "49.00", "currency": "USD"}) log.verify() # VerifyResult(ok=True, rows=2, chained=2, ...)
# someone edits row 2's amount in the file by hand... log.verify() # VerifyResult(ok=False, first_break="line 2: chain mismatch")
CLI (wire it into CI or a pre-ship gate — a tampered log exits nonzero, and a log that could only bepartiallyvouched for no longer exits like a fully verified one):
python -m arcaeon_ledger.cli append agent.log.jsonl '{"tool":"search","ok":true}' python -m arcaeon_ledger.cli verify agent.log.jsonl python -m arcaeon_ledger.cli verify --strict agent.log.jsonl
python -m arcaeon_ledger.cli verify agent.log.jsonl case $? in 0) echo "fully verified" ;; 3) echo "chain intact but prechain rows skipped unverified — inspect, or use --strict" ; exit 1 ;; ) echo "ledger broken" ; exit 1 ;; esac
A hash chain proves sequence integrity — it can't prove who wrote each entry or whether they were allowed to. Attach anauthorityblock to bind the actor and their permission surface into the chained (tamper-evident) row:
from arcaeon_ledger import Ledger, authority log = Ledger("agent.log.jsonl") log.append( {"tool": "payment", "amount": "49.00"}, authority=authority( "agent://billing-7", capability_version="v3", # what they were allowed to do tool_schema={"name": "payment", "args": ["amount"]}, # hashed, not just named time_source="ntp", # trust surface of the clock ), )
Now the audit question sharpens from"was this edited?"to"was this editedandwas the writer authorized?"— editing the principal, capability, or schema hash breaks the chain like any other tamper. This composes tamper-evidence with permission-replay. (Shipped in response to community feedback on launch.)
The loudest unmet pain for agent builders in 2026 is the reliability/audit gap: an agent "completes" a task and the result is quietly wrong, and you can't reconstruct — or prove — what actually happened. Observability platforms trace runs; none give you atamper-evident, portable, ownablerecord. Regulation is arriving too: the EU AI Act requires high-risk systems to technically allow automatic recording of events over their lifetime (Art. 12(1)) and requires providers and deployers to keep those logs, to the extent under their control, for at least six months (Art. 19(1), Art. 26(6)). The Act mandates recording and retention — tamper-evidence is not its word, it is ours: when someone asks whether a retained log is still the log, that question needs an answer stronger than trust.arcaeon-ledgeris the smallest honest version: a cryptographically chained action log you drop in, own, and verify.
chain = sha256(prev_chain + canonical_json(row_without_chain))[:32]
The chain value istruncated_sha256_128— the first 32 hex chars (128 bits) of SHA-256, not the full digest. Named so nobody cites it as full SHA-256: 128 bits is plenty for edit/accident detection, thinner if you want the chain itself to be expensive to grind after a rewrite (credit: atomic-raven's review).
Each row commits to the entire history before it. The first row chains from a fixed"genesis"seed. Rows without achainfield are tolerated only before the first chained row (so you can adopt it on an existing log); an unchained row appearingafterthe chain begins is itself flagged. On a mismatch, verify keeps going from the claimed value so it counts later damage honestly instead of cascading one break into noise.
What it proves — and the five things it doesn't
Being precise here is the product, not a disclaimer. A hash chain proves the recordedcontentof each row was not alteredin placeafter writing: mid-file edit, delete, and reorder all break it andverifynames the row.
One word in that sentence changed in 0.5.8, and the reason is the kind of thing this section exists for. It used to say "the recordedbytes", which claims more than the chain does. The chain is computed over each row parsed back from the file, and the reader normalises byte sequences it cannot decode — so two different byte strings inside such a region read identically and produce the same verdict. What is protected is the meaning of every row, not the exact bytes of the file. If you need byte-level custody, hash the file itself alongside this.
It doesnotby itself prove five other things:
1. Truncation.Lop off the most recent rows and what remains verifies clean — no append-only chain catches this alone. Close it by publishing the head somewhere outside your own control, on a cadence:
pin = log.head().as_pin() # -> "arcaeon-ledger head chain=9f3c… rows=204 as_of=2026-08-13T17:40:00Z" # post pin to a git commit / public comment / notarization anchor. # a reader compares a fresh head() against the last pin; a truncated or # re-minted history disagrees. the MAX gap between pins is your security # parameter, not the average — an attacker picks the gap.
2. Truth.The chain notarizes whatever was written — a tamper-evident record of a hallucination is still a hallucination with a checksum. To make a row speak about the world, hash a re-fetchable artefact (URL+bytes, a snapshot, tool stdout) and store that digest in the row, so a third party can re-get it and compare.
3. Authorship.authority()(above) records who-claimed-what, but it is data in the row, not a signature — a rewriter who re-mints from genesis re-mints it too. External head-anchoring (#1) is the thing a re-minter cannot advance.
4. Fabricated-legacy-prepend.Rows with nochainfield are toleratedbeforethe first chained row — that is deliberate, so you can adopt the chain on top of an existing log without rewriting its history. But skipped rows areunverifiedrows, and the verifier cannot tell real legacy history from a fabricated prepend. So (0.5.7) a non-strict verify that skipped any rows never mints a green:okisNone— "no break found, verified within scope" — falsy, with the scope in-band (verified_scope: "bounded_prechain_skipped") and the count inprechain; the CLI exits3, not0. Only a scan that checked every row returnsok=True. If your log is chained from genesis and must have no legitimate legacy rows, passverify(strict=True)/--strict— it treats any unchained row as a break, hard red. (An unchained row insertedafterthe chain begins is already flagged in every mode.)
5. Completeness.This is the big one, and it is structural: the agent decides what to callappendon. A tamper-evident log of the calls an agentchose to reportis still self-report. Nothing inside this library can close that, because anything the agent invokes, the agent can decline to invoke.
Close it by moving the pen out of the agent's reach — record at the seam instead, in a separate OS process the agent does not own, cannot skip, and cannot see:
pip install arcaeon-adapter python -m arcaeon_adapter --ledger seam.log.jsonl -- <your mcp server command...>
Scoped honestly, the primitive is"this file was not rewritten in place"— small, true, and testable. The layers above (external anchoring viahead(), artefact binding, signed authorship, seam recording) are how you extend it toward a full evidence claim.
The two look like the same thing — "no data" — andverify()treats them as opposites, on purpose:
Ledger("never/written.jsonl").verify() # VerifyResult(ok=False, rows=0, first_break="unreadable: [Errno 2] No such file...") open("touched/empty.jsonl", "w").close() Ledger("touched/empty.jsonl").verify() # VerifyResult(ok=True, rows=0, chained=0, first_break=None)
A path that was never created can't be vouched for —ok=False, "unreadable," same as any other read failure. A path that exists and is genuinely empty has zero rows to tamper with, so there's nothing for the chain to disagree about —ok=True, rows=0. Automation that branches onverify().okto decide "is this log intact" needs to checkfirst_break(or catch the missing-file case upstream) if it also needs to distinguish "never existed" from "exists, legitimately empty" —okalone collapses that distinction into two different answers, not one.
Bind what the agent actually read (artefact-binding)
The chain proves a row wasn't edited. It doesnotprove the row was evertrue— it will notarize a hallucination as faithfully as a fact.bind_artefactcloses that gap for the cases where you can point at a re-fetchable source: hash the actual bytes the agent read and store that digestinthe row, so a third party can re-get the source and compare.
from arcaeon_ledger import Ledger, bind_artefact log = Ledger("agent.log.jsonl") art = bind_artefact("https://example.com/pricing") # or bytes, a file path, or a dict log.append({"tool": "web.read", "url": "https://example.com/pricing", "artefact": art}) # art -> {"subject": {"name": "...", "digest": {"sha256": "..."}}, # "recipe": "sha256:raw-bytes:v1", # "digest": "sha256:raw-bytes:v1:<hex>", "bound_at": "...", "source_meta": {...}}
Digests areself-describing— never a bare hex hash. Each one issha256:<recipe>:<version>:<hex>, carrying its own recipe so a stranger reproduces it from the string alone:raw-bytes:v1(opaque bytes as-read) orjson-c14n:v1(a pinned, documented JSON canonicalization — sorted keys, compact, UTF-8). Recipes are frozen and versioned append-only, so old rows keep their recipe forever and a changed rule never makes history look tampered.
from arcaeon_ledger import verify_artefact verify_artefact(art) # recipe reproducible + string self-consistent verify_artefact(art, refetch=True) # for a URL: re-fetch and compare # -> {"verdict": "live_match", # <- THE answer; read this field # "digest_ok": True, "reason": None, # "refetch": "match" | "mismatch" | "unavailable" | "skipped", "notes": [...]}
Readverdict, not justdigest_ok(0.5.7).digest_oknames only theofflineleg — recipe reproducible, string self-consistent — and it staysTrueeven when a live re-fetch disagrees. The top-levelverdicttag mints the whole answer in one field:"digest_consistent"(offline leg passed, no live comparison made),"live_match","live_mismatch"(live content no longer matches — changedortampered, indeterminate),"live_unavailable"(the requested live check could not run), or the typed failure reason itself when the offline leg fails.if out["digest_ok"]afterrefetch=Trueused to read green through a live mismatch;out["verdict"] == "live_match"cannot.
A label this build cannot reproduce is a typed failure, never a pass.If the digest names an algorithm, recipe, or recipeversionoutside the supported registry,verify_artefactreturnsdigest_ok=Falsewith a machine-readablereason— one ofunknown_algorithm,unknown_recipe,unknown_recipe_version,malformed_digest,subject_digest_mismatch— and never reaches the re-fetch stage, so an unverifiable recipe can't come back as"match". A digest we cannot recompute is a digest we did not check, and "did not check" must not be reported as "verified." Old versions stay verifiable by staying listed inSUPPORTED_RECIPE_VERSIONSwhen a new one is minted, so the append-only recipe promise holds without the verifier waving through labels it has never shipped.
The honest boundary, stated loudly because it is the point:a re-fetchmismatchmeans the contentchanged orwas tampered —indeterminate. It is never reported as proof of tampering. The web mutates, 404s, paywalls, and personalizes; binding proves"this is the digest of the bytes the agent said it read at time T,"nothing stronger. For a neutral capture rather than your own fetch, route the source through a notarizing snapshot; forexisted-before-T, anchor the digest externally. Each is a layer you add — stated, not implied.
The chain can't catch truncation alone — lop off the most recent rows and what remains verifies clean (stated in "what it doesn't prove", above). The fix is awitness: a record-keeper outside your own control that holds your head(rows, chain)on a cadence. Once a witness has a pin from time T, a truncated log hasfewer rowsthan the witness saw, and a rewritten one has adifferent chainat the witnessed row. Neither can hide.
from arcaeon_ledger import Ledger, WitnessStore, publish_head, verify_against_witness log = Ledger("agent.log.jsonl") witness = WitnessStore("witness_pins.jsonl") # ideally on a host you don't control publish_head(witness, "billing-agent", log) # record the current head — do this on a cadence # later — did the log survive intact? v = verify_against_witness(witness, "billing-agent", log) v.verdict # "consistent" | "truncated" | "rewritten" | "no_record" bool(v) # truthy ONLY on "consistent" — a missing pin is no_record, never a false ok
WitnessStoreis the reference witness: one append-only JSONL file of pins. A hosted witness is a thin HTTP wrapper over exactly this object; run it locally and you have a complete, offline, zero-cost witness you fully control (with the obvious caveat that a witness you control is only as independent as its host).
What this proves, exactly.A witness proves your log wasn't truncated or rewrittenonly relative to what the witness saw, and only as recently as the last pin. Rows appended after the last pin are unprotected until the next one — sothe MAX gap between pins is your real security parameter, not the average, because an attacker picks the gap.And it says nothing about whether the logged content wastrue*— that's artefact-binding's job (above); the witness only guards the history's shape.
What the witness holds.Only fingerprints —(namespace, rows, chain, time)— never your log content. Password-nowhere by design: if the witness is breached, there is nothing sensitive to steal, only hashes useless without the original log.
arcaeon-ledgerships a zero-dependency MCP server, so any MCP client (Claude Code, etc.) can give its agent tamper-evident logging with no code. Wire it in:
{ "mcpServers": { "ledger": { "command": "python", "args": ["-m", "arcaeon_ledger.mcp_server", "--log", "agent.log.jsonl"] } } }
The agent then has two tools:ledger_append(record)to log an action (returns its chain hash) andledger_verify(strict?)to prove the log is intact (or get the exact tampered line back). The verify verdict is three-valued, same as the library:ok: true= every row verified,ok: null= chain intact but unchainedprechainrows were skipped unverified (verified_scope: "bounded_prechain_skipped"— not a green),ok: false= broken. Passstrict: trueto make any unchained row a hard failure. MCP is JSON-RPC over stdio and this server speaks it directly — no SDK, no extra install.
Core library, CLI, and a drop-inMCP server, all tested: the library against edit / delete / reorder tampering (test_ledger.py), the MCP server through a full initialize → tools/list → append → verify handshake including tamper detection over the wire. Extracted from a hash-chained action ledger running in production. External anchoring ships viahead()(publish the pin yourself) and the reference witness (WitnessStore, above); a hosted witness tier (retention, automatic pin cadence, compliance export) is the next layer.
This is a web browser that enables your coding agent, such as Claude Code, to visit websites on your behalf and assist you in identifying bugs or creating UI test cases.
Create crafted UI components inspired by the best 21st.dev design engineers.
Bring agent evaluations, observability, and synthetic test set generation directly into your IDE for free with Galileo's new MCP server
An MCP server to help AI assistants to answer questions and generate AccelByte Extend SDK code more effectively .
MCP server for AI Diagram Maker — generate beautiful software engineering diagrams directly inside Cursor, Claude Desktop, Claude Code, or any MCP-compatible AI agent
ALAPI MCP Tools,Call hundreds of API interfaces via MCP
AI-powered SVG animation generator that transforms static files into animated SVG components using the Allyson platform
MCP server that gives AI assistants on-demand access to 1,500+ amCharts docs, ~300 code examples, and 1000+ class API references.
APIMatic MCP Server is used to validate OpenAPI specifications using APIMatic. The server processes OpenAPI files and returns validation summaries by leveraging APIMatic’s API.
One shared context layer for AI agents and humans — live API specs, DB schemas, and versioned contracts across repos so every agent and teammate works from the same source of truth.
Build and deploy full-stack Next.js apps with 98 tools for React, AWS, and MongoDB
Sign in to leave a review
Use Google, GitHub, or an email account so ratings stay tied to real people.
No reviews posted yet.





