Entroly

by juyterman1000

Not rated
GitHub

Description

Local-first context-control plane and MCP server for AI coding agents. Selects evidence under a token budget, keeps omitted content byte-exactly recoverable via CCR handles, and emits auditable Context Receipts. Verify locally (no API key): entroly verify-claims (12/12 pass)…

About

Local-first context-control plane and MCP server for AI coding agents. Selects evidence under a token budget, keeps omitted content byte-exactly recoverable via CCR handles, and emits auditable Context Receipts. Verify locally (no API key): entroly verify-claims (12/12 pass). Savings are workload-dependent. Apache-2.0.

Details

Author
juyterman1000
Categories
Developer Tools

Setup

Install Entroly in your MCP client (Claude Desktop, Cursor, Windsurf, and others).

Repository: https://github.com/juyterman1000/entroly

Follow the installation instructions in the repository README, then restart your MCP client.

Entroly — Drop-In Context Assurance to Lower AI Operational Cost

Reduce unnecessary context without losing control of critical evidence.
Select the highest-value evidence first, compress it, keep originals recoverable, and emit a receipt — without rewriting your codebase or agent architecture.

Entroly is a local-first Context OS: content-addressed evidence, recoverable compression, and auditable receipts. Works through proxy, MCP, plugin, wrapper, and SDK paths with Claude Code, Codex, OpenClaw, GitHub Copilot, Cursor, Aider, and OpenAI/Anthropic-compatible apps.

100,438 downloads · growing day by day
Measured across different distribution sources.

⭐ If Entroly is useful to you, please star the repository on GitHub.
⭐ Star Entroly on GitHub— it helps the project grow and reach more developers.

English ·简体中文·繁體中文·日本語·한국어·Español·हिन्दी·Français·Deutsch·Português·Italiano·Türkçe·Tiếng Việt·Bahasa Indonesia·Polski·Nederlands·ไทย·Svenska·Čeština·Tagalog·Română

Tokens saved·Estimated cost avoided·Compression savings·Tool-schema deferral savings

Live means measured by Entroly, not a fabricated global number.Exact totals stay in each installation's local Value Receipt. Separately opted-in proxy installations may contribute a conservative community lower bound: every provider-bound delta is rounded down to whole 1,000-token units and whole cents before upload, with no prompt, content, model, price, or exact per-request value. It is not an exact worldwide total or provider invoice. Runentroly value,entroly value --json, or openentroly dashboardfor your exact local cumulative totals. For the public-counter contract and proxy metrics, seeLive tokenomicsandMetrics & Monitoring.

Tool schemas are never hidden by a relevance guess. To opt in for a request, send a comma-separated active set such asX-Entroly-Active-Tools: search_files,read_file. Forced tool choices and unnamed provider tools remain available; an invalid or non-matching set leaves the request unchanged.

Measurement contract·AI efficiency hub·Cost methodology·Metrics & monitoring·Privacy-safe telemetry

Token savings·Integrations·What is it?·Install·Quickstart·See it work·Benchmarks·Questions

Use Entroly at the SDK, framework, proxy, MCP, plugin or agent boundary. A listed name is not automatically a claim that hosted subscription inference is intercepted; provider-bound savings exist only when the request traverses an Entroly-controlled route.

Open the complete verified integration and operations hub →

AI coding assistants have a memory limit. Hand one your whole codebase and it gets slow, expensive, and distracted — like giving someone a 500-page manual when they only needed page 47.

It sits between your code and the AI, reads everything, and passes along only the parts that matter for the question actually being asked. Three things make that safe to do:

Do I need to pay for anything to try it?No. The two commands in theInstallsection below run entirely on your own machine, with no API key, and show you real numbers on your own project before you connect anything paid.

Not sure which one?PickPython. It's the complete version and what most people use. The others are alternate ways to run the same engine. | Platform | Install | What you get | |---|---|---| | 🐍Python(pip) —recommended|pip install -U entroly| Everything: the command-line tool, the server your AI editor talks to, and the code library | | 📦Node / npm|npm install -g entroly| The same engine, nothing Python required | | 🦀Rust(source build) |cd entroly-core && cargo build --release --bin entroly-rs --features proxy| One self-contained program, no Python or Node needed | | 🍺Homebrew|brew install juyterman1000/entroly/entroly| The command-line tool on macOS/Linux | | 🐳Docker|docker pull ghcr.io/juyterman1000/entroly:latest| Runs in a container, nothing installed on your machine |Now check that it worked — free, offline, no API key:

cd /your/repo entroly verify-claims entroly simulate

Both run locally. Neither one calls an AI or costs anything.

Extras (entroly[proxy],entroly[native],entroly[full]), the standalone Rust binary, and uninstall steps:Engine & install options.

Just want it working?pip install -U entroly && entroly go— that's the whole thing. It finds your editor, sets itself up, and shows you a before/after dashboard. The rest of this table is for specific setups. | Your situation | Do this | What it gets you | |---|---|---| | 🟢"I just want it on."(pip / Python user)|pip install -U entroly && entroly go| Auto-detects your editor, wraps your agent, opens a dashboard showing tokens before and after | |"I use Node, not Python."(npm user)|npm install -g entroly && entroly init| Same engine, nothing Python required | |"I want one binary, no runtime."(Rust user)|cargo build --release --bin entroly-rs --features proxy(fromentroly-core/) | A single native program with no dependencies | |"I use Claude Code / Cursor / Windsurf / VS Code."(MCP user)|entroly attach create --client claude --project . --ttl 4h --install(orentroly initfor Cursor/VS Code) | Your editor gets compression, receipts, and recovery as built-in tools — access expires on its own, and you change zero code | |"I'm building my own app in Python."(SDK user)|from entroly import compress, compress_messages, optimize| Call it straight from your code, anywhere you assemble a prompt | |"I have an API key and my own app."(proxy user)|entroly proxy→ pointANTHROPIC_BASE_URL/OPENAI_BASE_URL/GOOGLE_GEMINI_BASE_URLatlocalhost:9377| Every request gets optimized on the way past — no code changes on your side |Why bother:less unnecessary context reaches the model (lower bill, less distraction for the model), nothing is silently lost (every drop is recoverable and receipted), and you can prove it —entroly verify-claimsandentroly simulateshow real numbers on your own repo before you connect a paid key.

from entroly import compress, compress_messages, optimize compressed = compress(api_response, budget=2000) messages = compress_messages(messages, budget=30000) context = optimize(fragments, budget=8000, query="fix the login bug")
entroly compress response.json --out small.json entroly recover sha256:0b957c79... --out restored.json

Full setup paths for every agent, IDE, and CI use case:Get started in depth·Command reference.

Not mocked recordings — each video is rendered from a checked-in command that verifies its source artifact before printing a number.

entroly verify-claims— import, compression, receipts, WITNESS checks, recovery, proxy routing, replay. No API key.

On a frozen 24-case holdout, Entroly answered24/24; a published baseline answered18/24at roughly 1.5x the effective context.python scripts/readme_proof.py model-recovery

Omitted evidence recoveredbyte-exactafter a process restart, 66/66 payloads.python scripts/readme_proof.py restart-recovery

Full protocols, sample sizes, and every caveat:docs/BENCHMARKS.md.

The question that matters:if you send less, does the AI start getting things wrong?These are standard public tests, run with and without Entroly.

How to read this:Retentionis how well the AI still answered — 100% means it did just as well on far less text.Token savingsis how much less was sent (and therefore paid for). Measured withgpt-4o-mini; intervals are Wilson 95% CIs.

Being straight with you:look at the SQuAD 2.0 row — accuracy wentdown(80% → 72%). Compression is a trade, not magic, and it doesn't win everywhere. That's whyentroly simulateexists: run it on your own project and see your own numbers before you commit to anything.

Hallucination detection (WITNESS, local, no API):84.92%accuracy /0.7976 AUROCon 20,000HaluEval-QAdecisions — within the reported uncertainty ofgpt-4o-minias an API judge on the same shared sample.

Frozen evidence-selection benchmark (opt-in PRISM-R research prototype, not the default compressor): a disagreement guard kept the answer-bearing passage in 298 of 300 cases while selecting an average of 1.02 of 16 passages (paired exact McNemar p=0.21875 vs. BM25 alone) — this experiment measures retrieval of the known-answer passage, not generated-answer quality. Full protocol:PRISM-R neural evidence frontier.

Recovery, latency, and head-to-head frontier results are indocs/BENCHMARKS.mdwith raw artifacts linked. None of these numbers are a universal or production-savings guarantee for your workload — reproduce them on your own repo withentroly simulateandentroly value.

- Picks first, shrinks second— it works out which files actually answer your question,thencompresses them.
- Gives you the original back, exactly— anything left out can be restored character-for-character and checked against a fingerprint.
- Shows its work— a receipt for every decision: what was kept, what was left out and why, and what risk remains.
- Fact-checks answers— compares what the AI said against the evidence it was given, on your machine, without paying for a second AI call.
- Doesn't wreck your caching— keeps the unchanging parts of your prompt stable so your provider's discount for repeated text still applies.
- Rescues sessions before they crash— when a conversation grows too big, it trims recoverable output instead of letting the provider reject the request mid-task.
- Can route cheap work to cheap models— optional and fail-closed when uncertain.

Runs as aCLI,Python/TypeScript SDK,MCP server,HTTP proxy, orlibrary import. Full surface map:docs/product-surface.md. Architecture and Rust internals:docs/DETAILS.md.

Status describes integration depth, not a savings guarantee — provider-observed savings require requests to actually traverse an Entroly proxy route. Entroly does not claim interception of GitHub-hosted subscription inference on Copilot's native path. Full compatibility matrix:docs/agent-compatibility.md.

Entroly carries verified public metadata for GPT-5.6 Sol, Terra, and Luna; Gemini 3.6 Flash; and Gemini 3.5 Flash-Lite, and it can discover installed NVIDIA Nemotron 3.5 Lightning Ollama tags. Gated or private-preview announcements are not promoted into the verified matrix without a usable public model ID and limits. For example, Gemini 3.5 Flash Cyber remains outside the generally available matrix because its documented CodeMender access is restricted to selected governments and trusted partners. SeeVerified model supportfor model IDs, transport paths, limits, and availability boundaries.

NVIDIA Nemotron 3.5 Lightning with Ollama

Entroly supportsnemotron-3.5-lightningthrough its existing local Ollama discovery and OpenAI-compatible proxy path. This is a model-neutral integration: Entroly manages evidence selection, budgets, recovery handles, Context Receipts, and optional verification around the request; Ollama runs the model.

ollama pull nemotron-3.5-lightning python -m entroly.models discover ollama --inspect-ollama-context # Set ENTROLY_OPENAI_BASE=http://127.0.0.1:11434 in your shell, then: entroly proxy

Ollama lists the standardnemotron-3.5-lightningtag as a 30B mixture-of-experts model with 3B active parameters and a 1M context window. Its Apple-silicon30b-mlxtag is listed separately with a 256K window, so Entroly discovers the installed tag's metadata instead of assuming that every build has the same limit. Local Ollama inference can keep model prompts on the device; agent tools, configured remote providers, and other applications retain their own network and privacy boundaries.Compatibility, setup, and official sources.

Great fit:large repos where the agent only sees a few files at a time · chatty multi-turn agents · anywhere you want answers checked against evidence · cutting a real, growing AI bill.

Skip it:tiny repos or short prompts that already fit the budget · judgment-heavy tasks where you always want the full flagship model.

Also available:entroly wrap,entroly unwrap,entroly serve,entroly daemon,entroly dashboard,entroly demo,entroly capabilities,entroly ingest,entroly select,entroly receipt,entroly explain,entroly context-commit,entroly proof,entroly benchmark,entroly cache,entroly ravs,entroly perf,entroly batch. Full description:command reference.

- AI efficiency hub— token economics, AI cost optimization, memory, hallucination reduction, model routing, adaptive context, and verified code intelligence.
-
AI cost optimization— provider-bound input savings, billing boundaries, and workload-specific measurement.
-
Token economics— token saving, context compression, cache-aware context control, and more room in the context window.
-
Best token compression tools— comparison across 6 token reduction surfaces, ratios, byte-exact recoverability, and benchmark results.
-
Memory OS— budget-aware working, episodic, and semantic memory.
-
Hallucination reduction— WITNESS evidence-support verification.
-
Guarded model routing— RAVS routing, uncertainty control, and fail-closed escalation.
-
Adaptive context— bounded self-improving context.
-
Verified Code Context— parser-backed repository intelligence, typed graphs, architecture, value flow, LSP enrichment, source verification, and refactoring contracts.
-
Full benchmark evidence— every number, protocol, artifact, and caveat.
-
Model-triggered recovery holdout— frozen recovery protocol, evidence boundary, and reproduction details.
-
Context Commit conformance artifact— checked-in conformance evidence for Context Commit contracts.
-
Product surface map— CLI, SDK, MCP, proxy, verification, memory, security.
-
Architecture & full spec— Rust modules, compression, provenance, command reference.
-
Agent compatibility— every supported client and its exact authentication boundary.
-
First-run trust guide— exactly what to run before wiring a paid model key.
-
For teams— ROI, security, deployment one-pager.
-
Limitations— where Entroly helps, where it passes through, what it doesn't guarantee.
-
Public evidence policy— claim tiers and package links.
-
Context Commits·Context Receipts·Proof-guided recovery
-
Cookbook— copy-paste recipes.
-
Discord·Discussions·Issues

Compressing abadselection is still a bad selection. Entroly ranks first, then compresses — so the model gets structure, not just fewer tokens.

Apache-2.0 · local-first · no outbound analytics by default

pip install entroly && entroly go

This is a web browser that enables your coding agent, such as Claude Code, to visit websites on your behalf and assist you in identifying bugs or creating UI test cases.

Create crafted UI components inspired by the best 21st.dev design engineers.

Bring agent evaluations, observability, and synthetic test set generation directly into your IDE for free with Galileo's new MCP server

An MCP server to help AI assistants to answer questions and generate AccelByte Extend SDK code more effectively .

MCP server for AI Diagram Maker — generate beautiful software engineering diagrams directly inside Cursor, Claude Desktop, Claude Code, or any MCP-compatible AI agent

ALAPI MCP Tools,Call hundreds of API interfaces via MCP

AI-powered SVG animation generator that transforms static files into animated SVG components using the Allyson platform

MCP server that gives AI assistants on-demand access to 1,500+ amCharts docs, ~300 code examples, and 1000+ class API references.

APIMatic MCP Server is used to validate OpenAPI specifications using APIMatic. The server processes OpenAPI files and returns validation summaries by leveraging APIMatic’s API.

One shared context layer for AI agents and humans — live API specs, DB schemas, and versioned contracts across repos so every agent and teammate works from the same source of truth.

Build and deploy full-stack Next.js apps with 98 tools for React, AWS, and MongoDB

No reviews yet — be the first

Sign in to leave a review

Use Google, GitHub, or an email account so ratings stay tied to real people.

Email sign in

No reviews posted yet.