MCP CCPP Project Indexer

by walti1972

144 downloads
Not rated
GitHub

About

mcp-cpp-project-indexer is a deterministic C++ source-range indexer for large, module-heavy projects and MCP-based AI code navigation.

Details

Author
walti1972
Downloads
144
Categories
Developer Tools, Database, Other

- Deterministic C++ source-range indexing
- Maps symbols, files, and C++20 modules to exact source ranges
- Built for MCP-based AI code navigation
- Focused on large, module-heavy projects
- Does not guess code—only finds and reads it
- Not a compiler, LSP replacement, refactoring engine, semantic analyzer, or call-graph builder

mcp-cpp-project-indexeris a deterministic C++ source-range indexer for large, module-heavy projects and MCP-based AI code navigation.

It is not a compiler, LSP replacement, refactoring engine, semantic analyzer, or call-graph builder.

Find code. Read code. Do not guess code.

The indexer maps C++ symbols, files, and C++20 modules to exact source ranges so an AI can read only the code it needs.

Project and related AI orchestration work is documented on theMEF Programming homepage, including the ongoing relay/governance layer we are building around MCP tool use.

mcp-cpp-project-indexerbuilds a lightweight routing index over a C++ source tree. MCP clients can then ask deterministic questions such as:

- where is this function/class/data member?
- which exact source range should be read?
- which module imports or exports this partition?
- which changed hunk intersects which indexed symbol or data range?

The indexer returns metadata and original source ranges. It does not claim to understand the program. The AI still has to read the returned source and reason from that evidence.

User asks about Widget::OnScroll -> find_symbol("Widget::OnScroll") -> read_symbol(symbolId) -> AI explains only what was visible in that source range

This keeps large C++ projects out of the prompt until exact source evidence is needed.

git clone https://github.com/walti1972/mcp-cpp-project-indexer.git cd mcp-cpp-project-indexer
python <indexer-root>\build_project_index.py \ --root <project-root> \ --output-root <project-root>\.mcp-cpp-project-indexer
<project-root>\.mcp-cpp-project-indexer
python <indexer-root>\code_index_mcp_server.py \ --project-root <project-root> \ --index-root <project-root>\.mcp-cpp-project-indexer

For multiple MCP clients or a long-running shared process, use HTTP transport:

python <indexer-root>\code_index_mcp_server.py \ --project-root <project-root> \ --index-root <project-root>\.mcp-cpp-project-indexer \ --transport http \ --http-host 127.0.0.1 \ --http-port 8765
{ "mcpServers": { "mcp-cpp-project-indexer": { "command": "python", "args": [ "<indexer-root>\\code_index_mcp_server.py", "--project-root", "<project-root>", "--index-root", "<project-root>\\.mcp-cpp-project-indexer" ] } } }

5. Ask for exact source, not whole files

Find the symbol Widget::OnScroll, read its implementation, and explain only what is visible in the source range.
find_symbol -> read_symbol -> source-grounded answer

For best results, give your AI the rules fromprompt_template.md. The short version is:

Use metadata to locate code. Read exact source ranges before explaining behavior. Do not infer implementation behavior from metadata alone.

The public rootreadme.mdis human-facing project documentation. The actual Python implementation lives undersrc/:

src/ README.md indexer/ build_project_index.py update_project_index.py cpp_project_index.py server/ code_index_mcp_server.py server_ui/ ui/ indexer_tui.py indexer_control.py

Root-level scripts such asbuild_project_index.py,code_index_mcp_server.py,indexer_tui.py, andupdate_project_index.pyare compatibility wrappers. They keep existing command lines and MCP client configs working while routing execution to the implementation package.

Folder-local READMEs undersrc/use the project-indexer orientation format so agents can discover where to start without turning the public root README into machine-only documentation.

This project is used on real C++ codebases, not only toy examples. Two recent scale runs show the intended range:

These numbers are machine-dependent. The Chromium run used--jobs 60on a high-core workstation with an Intel Xeon Silver 4316 system, 128 GB RAM, and enterprise NVMe SSD storage. It is a useful public stress test because it exercises a very large classic include-based C++ codebase, while the anonymized commercial project exercises dense C++20 module and partition metadata. The Chromium run also validated the data/member indexer at scale: after fixing nested-template>>depth handling, the public stress test surfaced46,529additional data declarations and66,866additional data-name aliases.

The SQLite-backed lookup index keeps server startup practical even at Chromium scale: the MCP server can start immediately and stay around200 MB RAMafter startup instead of loading millions of symbol/data/name entries into Python objects.

It is designed for workflows that combine the Codex desktop app or other MCP clients with Visual Studio navigation and, when needed, binary/decompiler evidence from tools such as IDA Pro.

In one measured workflow, exact source-range routing reduced source text read from roughly 2,000 lines to 283 lines, an86% reduction.

For daily use, the indexer includes an optional mouse-capable TUI. It turns the project index into a small local control center:

- start the HTTP MCP server and watcher from one place
- run full builds, incremental updates, fast updates, and module-map rebuilds
- watch live server, watcher, lock, process, token, and index stats
- inspect build/update logs without switching tools
- toggle diagnostic file sections for deeper parser evidence when needed

Install the optional UI dependency and start the control center with explicit project/index paths:

pip install -r <indexer-root>\requirements-ui.txt python <indexer-root>\indexer_tui.py \ --root <project-root> \ --index-root <project-root>\.mcp-cpp-project-indexer \ --jobs 20 \ --http-url http://127.0.0.1:8765

The UI is optional; the core indexer remains dependency-light and can still be driven entirely from scripts or MCP clients. For setup and keyboard shortcuts, seeControl Center.

Large C++ projects are expensive to feed into an AI model when entire files are loaded just to find one function, class, import, or declaration. C++20 modules make this harder: many IDE/LSP-style tools still struggle with large module graphs, partitions, generated SDK headers, and build-specific configuration.

This indexer solves a narrower but very practical problem: it gives the AI a small, deterministic routing map. The AI can locate the relevant symbol, module, file, or changed hunk first, then read only the exact original source lines needed for the task.

Example from a real bug-finding workflow:

Whole file context: ~2000 source lines On-demand source reads: ~283 source lines -------------------------------------------- Reduction: ~86% less source text

The result is lower token usage, lower latency, less context drift, and more source-grounded analysis.

The indexer deliberately avoids pretending to be a compiler.
- Fast token/structure scanSource files are scanned with a lightweight Python lexer and structural parser. The output is a deterministic table of contents: files, symbols, data declarations, lexical
#includedirectives, source ranges, diagnostics, and module facts.
- C++20 module mapModule interfaces, partitions, imports, export-imports, and consumers are indexed so the AI can route through module-heavy code without asking an LSP to solve the whole build.
- Incremental update and watcherThe updater tracks content hashes and rewrites only changed index data where possible. The optional watcher can keep the MCP server cache fresh while you work in Visual Studio.
- MCP tools with compact output controlsTools expose exact routing metadata first. The AI escalates only when needed: compact symbol lookup, file/module/change overview, exact
read_symbolorread_range, then deeper recursive source reads.

The indexer is only the table of contents. The AI performs recursive exploration and code review from the original source lines it explicitly reads.

Read Renderer.cpp completely: ~2000 lines
find_symbol("Renderer::Paint") read_symbol(symbolId) inspect visible calls read only relevant project callees
list_changed_files get_file_change_hunks(includeIndexedRangeSummary:true, includeSource:false) get_file_change_hunks(symbolId/dataId, includeSource:true) read_symbol/read_range only when current source behavior is needed

The scanner extracts routing facts from C++ source files:

- files and stable file IDs
- C++20 modules and partitions
- lexical
#includedirectives
- imports and exports
- namespaces
- classes / structs / enums
- functions / methods
- constructors / destructors / operators
- declarations and inline definitions
- exact
startLine/endLine
- diagnostics for structurally suspicious files

It is stream/token based, not regex based.

- no compiler-accurate whole-program call graph
- no
find_references
- no type resolution
- no template-instantiation resolution
- no compiler-accurate overload resolution
- no macro expansion
- no semantic summaries
- no bug analysis
- no
analyze_symbol(symbolId)

The AI should read source ranges and reason from the original code.

<project-root>/.mcp-cpp-project-indexer/
.mcp-cpp-project-indexer/ manifest.json files/ f_<pathHash>.json index.sqlite modules.json diagnostics.json update_state.json # written by build/update; used for fast incremental updates module_map.json # generated by build_module_map.py .watch_update_summary.json # temporary watcher/update summary .update.lock # process lock for index writers .watcher.lock # process lock for one active watcher

Global symbol and data routing indexes are stored inindex.sqlite. The per-file JSON indexes remain the source of truth for exact source ranges and incremental rebuilds.

python <indexer-root>\export_index_jsonl.py --index-root <project-root>\.mcp-cpp-project-indexer --kind symbols --output symbols.jsonl python <indexer-root>\export_index_jsonl.py --index-root <project-root>\.mcp-cpp-project-indexer --kind data --output data.jsonl

Scanner diagnostic file-index fields are emitted only with--emit-diagnosticsor--emit-diagnostic-file-indexes:

scopeIntervals structuralEvents functionBodyRanges
python <indexer-root>\build_file_index.py \ --file <project-root>\path\to\file.ixx \ --project-root <project-root> \ --output <project-root>\.mcp-cpp-project-indexer\diagnostic_file.json
python <indexer-root>\build_file_index.py \ --file <project-root>\path\to\file.ixx \ --project-root <project-root> \ --output <project-root>\.mcp-cpp-project-indexer\diagnostic_file.json \ --emit-diagnostics

If--project-rootis omitted, the file's parent directory is used.

Recommended usage from the C++ project root:

cd <project-root> python <indexer-root>\build_project_index.py
<project-root>/.mcp-cpp-project-indexer/
python <indexer-root>\build_project_index.py \ --root <project-root> \ --output-root <project-root>\.mcp-cpp-project-indexer
Built cpp.project_index.v1 Root: <project-root> Output: <project-root>/.mcp-cpp-project-indexer Files: 7076 Symbols: 97583 Names: 95674 Modules: 3774 Diagnostics: 7 Total code lines: 1750000 Total tokens: 14200000 SQLite index: <project-root>/.mcp-cpp-project-indexer/index.sqlite

Total tokensis the indexer's lexer token count over the indexed source after comment blanking. It is a project-size metric, not an LLM billing-token count.

When the project root is inside a Git worktree andgitis available, file discovery respects Git ignore rules by filtering candidates throughgit check-ignore --stdin. This excludes paths matched by.gitignore,.git/info/exclude, or the user's global Git ignore file. Non-Git projects, or systems without Git, fall back to the built-in excluded directory list. Dot-directories such as.git,.vs,.cache,.idea, or.folderare excluded by default.

For large projects with mixed source layouts, placeindexer_config.jsonin the project root or in any subdirectory. Config files are applied while walking the tree: the root config becomes the base, and subdirectory configs can override or extend it for that subtree.

{ "addExtensions": [".mm"], "addExcludeDirs": ["generated", "third_party"], "includeExtensionlessHeaders": true, "useGitIgnore": false }
{ "extensions": [".cpp", ".cc", ".h"], "addExtensions": [".mm"], "removeExtensions": [".c"], "excludeDirs": ["out", "build"], "addExcludeDirs": ["generated"], "removeExcludeDirs": ["third_party"], "includeExtensionlessHeaders": true, "useGitIgnore": false }

extensionsandexcludeDirsreplace the inherited values for that subtree.addandremovefields modify the inherited values. Extensionless header discovery is conservative and opt-in; it only accepts extensionless files whose first lines look like C/C++ headers.useGitIgnore:falsedisables the finalgit check-ignore --stdinpass for very large repositories where Git ignore filtering is more expensive than an explicit indexer config.

After a full index build, changed files can be detected and re-indexed incrementally.

python <indexer-root>\update_project_index.py \ --root <project-root> \ --index-root <project-root>\.mcp-cpp-project-indexer \ --dry-run
python <indexer-root>\update_project_index.py \ --root <project-root> \ --index-root <project-root>\.mcp-cpp-project-indexer \ --known-files-only

Watcher-style update for a known changed file:

python <indexer-root>\update_project_index.py \ --root <project-root> \ --index-root <project-root>\.mcp-cpp-project-indexer \ --known-files-only \ --changed-file path\to\changed.cpp

--known-files-onlyskips full discovery of new files. This is ideal for save/watch loops.--changed-filecan be repeated and lets a watcher avoid hashing unchanged files.

Index writes are protected by an exclusive.update.lockfile in the index root. This prevents full builds, incremental updates, and module-map rebuilds from writing the same index files at the same time.

No reviews yet — be the first

Sign in to leave a review

Use Google, GitHub, or an email account so ratings stay tied to real people.

Email sign in

No reviews posted yet.