Vision Memory MCP
About
Persistent visual cache for LLM-driven software development. Caches screenshots using perceptual hashing, vector search, and AX trees to prevent token overhead and visual hallucination loops.
Details
- Author
- putervision
- Categories
- AI, Other, Knowledge Base
Jump to
Alternative Options & CLI Usage Examples
# Run stdio MCP server directly via binary (after global install) vision-memory-mcp run # Start server skipping heavy CLIP model downloads (air-gapped / offline mode) vision-memory-mcp run --skip-model-load # Re-initialize across all registered workspace projects vision-memory-mcp init-global # Health check dependencies, sharp bindings, and git safety vision-memory-mcp doctor # Run health diagnostics & aggregate metrics across all registered projects vision-memory-mcp doctor-global # Inspect stored visual states and metadata in terminal ASCII table vision-memory-mcp inspect # Register baseline design mockup contract (Visual SDD) vision-memory-mcp spec set --name "Dashboard" --file ./dashboard-spec.png # Save visual memory checkpoint snapshot vision-memory-mcp snapshot save --name "v1.0-milestone" # Ingest WebM / MP4 video recording into visual state memory timeline vision-memory-mcp video ingest ./playwright-test.webm --category playwright_test # Open interactive force-directed visual graph viewer in browser vision-memory-mcp view
- 👁️ Perceptual Visual Caching: Sub-5ms L1/L2 dHash zero-token fast-path layout recognition.
- 🎬 WebM & MP4 Video Ingestion: Digest E2E test recordings & screen captures into searchable keyframe visual states & state transition graphs.
- ⚡ 15 Core MCP Tools: High-coherence consolidated toolset covering perception, video memory, evidence packs, trajectory comparison, semantic retrieval, element grounding, visual SDD, snapshots, and unified context & metrics.
- 🔗 Dual-MCP Synergy & Immutable Evidence Packs: Deeply bridges@putervision/state-memory-mcptask DAGs with visual state memory, generating cryptographically hashable evidence packs for compliance and audit trails.
- 📉 Reduced Token Overhead: Caches UI states locally using dHash, local CLIP vector search, and accessibility trees to maximize vision token savings.
- 🚀 Sub-5ms Fast-Path Latency: Eliminates repetitive vision LLM API calls and avoids visual hallucination loops.
- 🎯 Element Grounding & Action Target Prediction: Maps screen elements to CSS selectors and coordinates for deterministic UI interaction.
- 🎨 Visual Spec-Driven Development (Visual SDD): Register design mockups or screenshots as perceptual baseline contracts to verify visual regression.
- 🛡️ 100% Local-First Privacy: Local LanceDB vector store, local CLIP model, zero cloud telemetry, and PII redaction guarantees.
@putervision/vision-memory-mcpprovides15 production-grade consolidated MCP toolsstructured across 4 core visual perception & automation domains:
- Perception & Semantic Search:analyze_screenshot(L1/L2 perceptual dHash & AX tree parsing, single/batch),recall_memory(text & image semantic vector search),get_session_context(aggregated cache hit metrics, recent states).
- Element Grounding & Navigation:predict_next_action(deterministic CSS selectors & bounding coordinates),record_outcome(UI action transitions & visual blockers),get_navigation_paths(BFS shortest-path planner),wait_for_visual_state(polling for target UI state).
- Video Trajectories & Evidence Packs:manage_video(WebM/MP4 keyframe ingestion, timeline search),compare_states(visual layout diffs & video trajectory comparison),create_evidence_pack(cryptographic audit proof linking video keyframes to state-memory DAGs),export_trajectories(multimodal fine-tuning datasets).
- Snapshots & Visual SDD:manage_visual_spec(mockup baseline contracts & regression checks),manage_snapshot(checkpoints, export, restore),undo_visual_mutation(revert state ingestion),forget_state(privacy & PII purging).
👉 For complete parameter specifications, return schemas, and example payloads, see theFormal API ReferenceandFeatures & Architecture Guide.
Incoming Screen │ ▼ ┌──────────────────────────────┐ │ L1: In-Memory Cache Lookup │ ──(Hit)──▶ Return Cached Description & Grounded Elements └──────────────┬───────────────┘ │ (Miss) ▼ ┌──────────────────────────────┐ │ L2: Perceptual Hash Scan │ ──(Hit)──▶ Return Cached Description & Grounded Elements └──────────────┬───────────────┘ │ (Miss) ▼ ┌──────────────────────────────┐ │ L3: Local CLIP Vector Search │ ──(Hit)──▶ Return Semantically Close └──────────────┬───────────────┘ │ (Miss) ▼ ┌──────────────────────────────┐ │ L4: Vision LLM Fallback │ ──(Ingest)──▶ Save Redacted State to DB └──────────────┬───────────────┘
Explore dedicated guides and deep dives in thedocs/directory:
Whilevision-memory-mcpis designed for visual frontend state caching, UI testing, and multimodal workflows, it may not be appropriate for:
- Headless / Pure Backend Development: Non-visual CLI tools, database scripts, or pure backend microservices with no UI rendering. (Usestate-memory-mcpstandalone instead).
- High-Framerate Live Video Streaming: Continuous 60 fps live video ingest without discrete keyframe or test action boundaries.
- Ultra Low-Memory Embedded Environments (<512 MB RAM): Running full local CLIP neural embeddings requires ~300 MB RAM (use--skip-model-loadfor lightweight dHash-only perception if memory is constrained).
# Run full unit and integration test suite across all 69 test files (300 tests) npm run test
Developed and maintained byPuterVision. Released under theMIT License.
- Local Storage Guarantee: Provided "as is" without warranty. Screenshots, perceptual hashes, vector embeddings, and transition graphs are stored locally unencrypted at the application level in.vision-memory-mcp/. Zero telemetry or analytics data is ever transmitted.
- Trademarks & Non-Affiliation: Product names (Cursor, Claude Code, Gemini, Windsurf, VS Code, Sharp, LanceDB, ONNX, HuggingFace) are property of their respective owners and used solely for compatibility identification.
Private persistent memory for Claude, ChatGPT & Gemini via MCP — semantic search, zero-code setup.
An MCP server powered by txtai for semantic search, knowledge graphs, and AI-driven text processing.
A powerful Model Context Protocol (MCP) server using gemini embedding 3 that transforms any local directory into an ultrafast, visually-aware spatial search engine for AI agents.
A fully-local MCP server for question-answering over your PDFs. Ask in plain language; Claude retrieves only the relevant passages with page citations. On-device embeddings (sentence-transformers) + ChromaDB — no API keys, nothing leaves your machine.
An MCP server for web and similarity search, designed for Claude Desktop. It integrates with various external embedding and API services.
MCP server for Prompt Builder — search, retrieve, and compile prompt components from a community vault using semantic search (pgvector) and slug-based lookup. Works with Claude Desktop and Cursor.
A server providing web and similarity search functionalities, designed for Claude Desktop. It requires external embedding and API services.
Offline MCP server that ranks & summarizes code using BM25, TF-IDF, embeddings & git signals; integrates with Cursor, Claude Desktop and Windsurf; privacy preserving.
Agent-agnostic persistent memory backend. 13 MCP tools, Supabase + Jina embeddings, multi-profile isolation, semantic recall across sessions.
Local-first long-term memory for coding agents — in-process embeddings (MLX/CPU), hybrid vector+BM25 search, markdown as source of truth. No cloud, no keys.
@putervision/vision-memory-mcpis a zero-infrastructure, local-first Model Context Protocol (MCP) server and CLI tool that provides AI coding assistants (such as Cursor, Claude Code, Gemini, or Copilot) with visual state caching using perceptual hashing, local CLIP embeddings, and transition graphs to eliminate repetitive vision LLM calls.
🌐Official Documentation & Website:visionmemorymcp.com
# Global installation via npm npm install -g @putervision/vision-memory-mcp
Runinitin your project root to scaffold database directories,.gitignore,.env, and IDE rules:
Add to your MCP client config (e.g..cursor/mcp.jsonor.vscode/mcp.json):
{ "mcpServers": { "vision-memory-mcp": { "command": "vision-memory-mcp", "args": ["run"] } } }
Alternative Options & CLI Usage Examples
# Run stdio MCP server directly via binary (after global install) vision-memory-mcp run # Start server skipping heavy CLIP model downloads (air-gapped / offline mode) vision-memory-mcp run --skip-model-load # Re-initialize across all registered workspace projects vision-memory-mcp init-global # Health check dependencies, sharp bindings, and git safety vision-memory-mcp doctor # Run health diagnostics & aggregate metrics across all registered projects vision-memory-mcp doctor-global # Inspect stored visual states and metadata in terminal ASCII table vision-memory-mcp inspect # Register baseline design mockup contract (Visual SDD) vision-memory-mcp spec set --name "Dashboard" --file ./dashboard-spec.png # Save visual memory checkpoint snapshot vision-memory-mcp snapshot save --name "v1.0-milestone" # Ingest WebM / MP4 video recording into visual state memory timeline vision-memory-mcp video ingest ./playwright-test.webm --category playwright_test # Open interactive force-directed visual graph viewer in browser vision-memory-mcp view
- 👁️ Perceptual Visual Caching: Sub-5ms L1/L2 dHash zero-token fast-path layout recognition.
- 🎬 WebM & MP4 Video Ingestion: Digest E2E test recordings & screen captures into searchable keyframe visual states & state transition graphs.
- ⚡ 15 Core MCP Tools: High-coherence consolidated toolset covering perception, video memory, evidence packs, trajectory comparison, semantic retrieval, element grounding, visual SDD, snapshots, and unified context & metrics.
- 🔗 Dual-MCP Synergy & Immutable Evidence Packs: Deeply bridges@putervision/state-memory-mcptask DAGs with visual state memory, generating cryptographically hashable evidence packs for compliance and audit trails.
- 📉 Reduced Token Overhead: Caches UI states locally using dHash, local CLIP vector search, and accessibility trees to maximize vision token savings.
- 🚀 Sub-5ms Fast-Path Latency: Eliminates repetitive vision LLM API calls and avoids visual hallucination loops.
- 🎯 Element Grounding & Action Target Prediction: Maps screen elements to CSS selectors and coordinates for deterministic UI interaction.
- 🎨 Visual Spec-Driven Development (Visual SDD): Register design mockups or screenshots as perceptual baseline contracts to verify visual regression.
- 🛡️ 100% Local-First Privacy: Local LanceDB vector store, local CLIP model, zero cloud telemetry, and PII redaction guarantees.
@putervision/vision-memory-mcpprovides15 production-grade consolidated MCP toolsstructured across 4 core visual perception & automation domains:
- Perception & Semantic Search:analyze_screenshot(L1/L2 perceptual dHash & AX tree parsing, single/batch),recall_memory(text & image semantic vector search),get_session_context(aggregated cache hit metrics, recent states).
- Element Grounding & Navigation:predict_next_action(deterministic CSS selectors & bounding coordinates),record_outcome(UI action transitions & visual blockers),get_navigation_paths(BFS shortest-path planner),wait_for_visual_state(polling for target UI state).
- Video Trajectories & Evidence Packs:manage_video(WebM/MP4 keyframe ingestion, timeline search),compare_states(visual layout diffs & video trajectory comparison),create_evidence_pack(cryptographic audit proof linking video keyframes to state-memory DAGs),export_trajectories(multimodal fine-tuning datasets).
- Snapshots & Visual SDD:manage_visual_spec(mockup baseline contracts & regression checks),manage_snapshot(checkpoints, export, restore),undo_visual_mutation(revert state ingestion),forget_state(privacy & PII purging).
👉 For complete parameter specifications, return schemas, and example payloads, see theFormal API ReferenceandFeatures & Architecture Guide.
Incoming Screen │ ▼ ┌──────────────────────────────┐ │ L1: In-Memory Cache Lookup │ ──(Hit)──▶ Return Cached Description & Grounded Elements └──────────────┬───────────────┘ │ (Miss) ▼ ┌──────────────────────────────┐ │ L2: Perceptual Hash Scan │ ──(Hit)──▶ Return Cached Description & Grounded Elements └──────────────┬───────────────┘ │ (Miss) ▼ ┌──────────────────────────────┐ │ L3: Local CLIP Vector Search │ ──(Hit)──▶ Return Semantically Close └──────────────┬───────────────┘ │ (Miss) ▼ ┌──────────────────────────────┐ │ L4: Vision LLM Fallback │ ──(Ingest)──▶ Save Redacted State to DB └──────────────┬───────────────┘
Explore dedicated guides and deep dives in thedocs/directory:
Whilevision-memory-mcpis designed for visual frontend state caching, UI testing, and multimodal workflows, it may not be appropriate for:
- Headless / Pure Backend Development: Non-visual CLI tools, database scripts, or pure backend microservices with no UI rendering. (Usestate-memory-mcpstandalone instead).
- High-Framerate Live Video Streaming: Continuous 60 fps live video ingest without discrete keyframe or test action boundaries.
- Ultra Low-Memory Embedded Environments (<512 MB RAM): Running full local CLIP neural embeddings requires ~300 MB RAM (use--skip-model-loadfor lightweight dHash-only perception if memory is constrained).
# Run full unit and integration test suite across all 69 test files (300 tests) npm run test
Developed and maintained byPuterVision. Released under theMIT License.
- Local Storage Guarantee: Provided "as is" without warranty. Screenshots, perceptual hashes, vector embeddings, and transition graphs are stored locally unencrypted at the application level in.vision-memory-mcp/. Zero telemetry or analytics data is ever transmitted.
- Trademarks & Non-Affiliation: Product names (Cursor, Claude Code, Gemini, Windsurf, VS Code, Sharp, LanceDB, ONNX, HuggingFace) are property of their respective owners and used solely for compatibility identification.
Private persistent memory for Claude, ChatGPT & Gemini via MCP — semantic search, zero-code setup.
An MCP server powered by txtai for semantic search, knowledge graphs, and AI-driven text processing.
A powerful Model Context Protocol (MCP) server using gemini embedding 3 that transforms any local directory into an ultrafast, visually-aware spatial search engine for AI agents.
A fully-local MCP server for question-answering over your PDFs. Ask in plain language; Claude retrieves only the relevant passages with page citations. On-device embeddings (sentence-transformers) + ChromaDB — no API keys, nothing leaves your machine.
An MCP server for web and similarity search, designed for Claude Desktop. It integrates with various external embedding and API services.
MCP server for Prompt Builder — search, retrieve, and compile prompt components from a community vault using semantic search (pgvector) and slug-based lookup. Works with Claude Desktop and Cursor.
A server providing web and similarity search functionalities, designed for Claude Desktop. It requires external embedding and API services.
Offline MCP server that ranks & summarizes code using BM25, TF-IDF, embeddings & git signals; integrates with Cursor, Claude Desktop and Windsurf; privacy preserving.
Agent-agnostic persistent memory backend. 13 MCP tools, Supabase + Jina embeddings, multi-profile isolation, semantic recall across sessions.
Local-first long-term memory for coding agents — in-process embeddings (MLX/CPU), hybrid vector+BM25 search, markdown as source of truth. No cloud, no keys.
Sign in to leave a review
Use Google, GitHub, or an email account so ratings stay tied to real people.
No reviews posted yet.





