Sourcerer

by st3v3nmw

303 downloads
Not rated
GitHub

About

MCP for semantic code search & navigation that reduces token waste

Details

Author
st3v3nmw
Downloads
303
Categories
Other, AI

- Semantic search by concept and functionality
- Retrieve specific code chunks by stable ID
- Uses Tree-sitter for AST-based code parsing
- Automatically re-indexes changed files via file watching
- Respects .gitignore rules
- Stores embeddings persistently in .sourcerer/db/

Setting up with Highlight

This MCP is not yet compatible with Highlight’s one-click setup. However, you can still use it with Highlight by following these steps:

  1. Download and install Highlight from highlightai.com/download
  2. Navigate to the plugins tab and select "Add Custom Plugin"
  3. Configure the plugin with the settings below
    Plugin Name Sourcerer
    Command (node, npx, python, etc.)

    Please refer to the README for specific instructions on how to obtain API keys or other required environment variables.

  4. Enable "Start Automatically" if you want the plugin to start when Highlight launches

From the repository

Install via go install or Homebrew. Configure with an OpenAI API key and workspace root path, then add the server to your MCP client (e.g., Claude Code or mcp.json). Once configured, agents can call tools like semantic_search and get_source_code to query and retrieve code.

Claude Desktop / Cursor

Paste into your MCP client config file to install this server.

{
    "mcpServers": {
        "sourcerer": {
            "sourcerer": {
                "command": "sourcerer",
                "env": {
                    "OPENAI_API_KEY": "your-openai-api-key",
                    "SOURCERER_WORKSPACE_ROOT": "/path/to/your/project"
                }
            }
        }
    }
}

McpServers

{
    "sourcerer": {
        "command": "sourcerer",
        "env": {
            "OPENAI_API_KEY": "your-openai-api-key",
            "SOURCERER_WORKSPACE_ROOT": "/path/to/your/project"
        }
    }
}

Sourcerer MCP 🧙

An MCP server for semantic code search & navigation that helps AI agents work efficiently without burning through costly tokens. Instead of reading entire files, agents can search conceptually and jump directly to the specific functions, classes, and code chunks they need.

Demo

asciicast

Requirements

- OpenAI API Key: Required for generating embeddings (local embedding support planned) - Git: Must be a git repository (respects .gitignore files) - Add .sourcerer/ to .gitignore: This directory stores the embedded vector database

Installation

Go

``shell go install github.com/st3v3nmw/sourcerer-mcp/cmd/sourcerer@latest `

Homebrew

`shell brew tap st3v3nmw/tap brew install st3v3nmw/tap/sourcerer `

Configuration

Claude Code

`shell claude mcp add sourcerer -e OPENAI_API_KEY=your-openai-api-key -e SOURCERER_WORKSPACE_ROOT=$(pwd) -- sourcerer `

mcp.json

`json { "mcpServers": { "sourcerer": { "command": "sourcerer", "env": { "OPENAI_API_KEY": "your-openai-api-key", "SOURCERER_WORKSPACE_ROOT": "/path/to/your/project" } } } } `

How it Works

Sourcerer builds a semantic search index of your codebase:

1. Code Parsing & Chunking

- Uses Tree-sitter to parse source files into ASTs - Extracts meaningful chunks (functions, classes, methods, types) with stable IDs - Each chunk includes source code, location info, and contextual summaries - Chunk IDs follow the pattern:
file.ext::TypeName::methodName

2. File System Integration

- Watches for file changes using
fsnotify - Respects .gitignore files via git check-ignore - Automatically re-indexes changed files - Stores metadata to track modification times

3. Vector Database

- Uses chromem-go for persistent vector storage in
.sourcerer/db/ - Generates embeddings via OpenAI's API for semantic similarity - Enables conceptual search rather than just text matching - Maintains chunks, their embeddings, and metadata

4. MCP Tools

-
semantic_search: Find code by concept/functionality - get_source_code: Retrieve specific chunks by ID - index_workspace: Manually trigger re-indexing - get_index_status`: Check indexing progress This approach allows AI agents to find relevant code without reading entire files, dramatically reducing token usage and cognitive load.

Supported Languages

Language support requires writing Tree-sitter queries to identify functions, classes, interfaces, and other code structures for each language. Supported: Go Planned: Python, TypeScript, JavaScript

Contributing

All contributions welcome!
No reviews yet — be the first

Sign in to leave a review

Use Google, GitHub, or an email account so ratings stay tied to real people.

Email sign in

No reviews posted yet.