llm-advisor-mcp
About
Real-time LLM/VLM model comparison with benchmarks, pricing, and personalized recommendations from 5 data sources. No API key required.
Details
- Author
- daichi-kudo
- Categories
- Developer Tools, AI
Jump to
Setup
Install llm-advisor-mcp in your MCP client (Claude Desktop, Cursor, Windsurf, and others).
Repository: https://github.com/daichi-kudo/llm-advisor-mcp
Follow the installation instructions in the repository README, then restart your MCP client.
Give your AI assistant real-time LLM/VLM knowledge.Pricing, benchmarks, and recommendations — updated every hour, not every training cycle.
LLMs have knowledge cutoffs. Ask Claude "what's the best coding model right now?" and it cannot answer with current data. This MCP server fixes that by feeding live model intelligence directly into your AI assistant's context window.
- Zero config— No API keys, no registration. One command to install.
- Low token— Compact Markdown tables (~300 tokens), not raw JSON (~3,000 tokens). Your context window matters.
- 5 benchmark sources— SWE-bench, LM Arena Elo, OpenCompass VLM, Aider Polyglot, and OpenRouter pricing merged into one unified view.
- "What's the best coding model right now?"—list_top_modelswith categorycoding
- "Compare Claude vs GPT vs Gemini"—compare_modelswith side-by-side table
- "Find a cheap model with 1M context"—recommend_modelwith budget constraints
- "What benchmarks does model X have?"—get_model_infowith percentile ranks
claude mcp add llm-advisor -- npx -y llm-advisor-mcp
claude mcp add llm-advisor -- cmd /c npx -y llm-advisor-mcp
{ "mcpServers": { "llm-advisor": { "command": "npx", "args": ["-y", "llm-advisor-mcp"] } } }
Detailed specs for a specific model: pricing, benchmarks, percentile ranks, capabilities, and a ready-to-use API code example.
## anthropic/claude-sonnet-4 Provider: anthropic | Modality: text+image→text | Released: 2025-06-25 ### Pricing | Metric | Value | |--------|-------| | Input | $3.00 /1M tok | | Output | $15.00 /1M tok | | Cache Read | $0.30 /1M tok | | Context | 200K | | Max Output | 64K | ### Benchmarks | Benchmark | Score | |-----------|-------| | SWE-bench Verified | 76.8% | | Aider Polyglot | 72.1% | | Arena Elo | 1467 | | MMMU | 76.0% | ### Percentile Ranks | Category | Percentile | |----------|------------| | Coding | P96 | | General | P95 | | Vision | P90 | Capabilities: Tools, Reasoning, Vision ### API Example (openai_sdk) python from openai import OpenAI client = OpenAI( base_url="https://openrouter.ai/api/v1", api_key="<OPENROUTER_API_KEY>", ) response = client.chat.completions.create( model="anthropic/claude-sonnet-4", messages=[{"role": "user", "content": "Hello"}], )
--- ### list_top_models Top-ranked models for a category. Includes release dates for freshness awareness. Parameters | Name | Type | Required | Default | Description | |------|------|----------|---------|-------------| | category | enum | Yes | — | coding, math, vision, general, cost-effective, open-source, speed, context-window, reasoning | | limit | number | No | 10 | Number of results (1-20) | | min_context | number | No | — | Minimum context window in tokens | | min_release_date | string | No | — | YYYY-MM-DD. Excludes models released before this date | Example output
--- ### compare_models Side-by-side comparison for 2-5 models. Best values are bolded automatically. Includes a Released row so you can spot outdated models at a glance. Parameters | Name | Type | Required | Default | Description | |------|------|----------|---------|-------------| | models | string[] | Yes | — | 2-5 model IDs or partial names | Example output
--- ### recommend_model Personalized top-3 recommendations. Scores combine weighted benchmarks, pricing, capability bonuses, and a freshness bonus (+3 points for models released within 3 months, +1 within 6 months). Parameters | Name | Type | Required | Default | Description | |------|------|----------|---------|-------------| | use_case | enum | Yes | — | coding, math, general, vision, creative, reasoning, cost-effective | | max_input_price | number | No | — | Max input price (USD/1M tokens) | | max_output_price | number | No | — | Max output price (USD/1M tokens) | | min_context | number | No | — | Minimum context window in tokens | | require_vision | boolean | No | — | Require image input support | | require_tools | boolean | No | — | Require tool/function calling support | | require_open_source | boolean | No | — | Require open-source license | | min_release_date | string | No | — | YYYY-MM-DD. Excludes older models | Example output
1. anthropic/claude-sonnet-4 (score: 78)
Input: $3.00/1M | Output: $15.00/1M | Context: 200K | Released: 2025-06-25 Benchmarks: SWE-bench: 76.8%, Aider: 72.1%, Arena: 1467 Strengths: reasoning, tools, vision
Input: $0.15/1M | Output: $0.60/1M | Context: 1M | Released: 2025-05-20 Benchmarks: SWE-bench: 62.9%, Arena: 1445 Strengths: tools, vision, 1M+ context
Input: $1.10/1M | Output: $4.40/1M | Context: 200K | Released: 2025-04-16 Benchmarks: SWE-bench: 73.6%, Arena: 1430 Strengths: reasoning, tools
--- ## Data Sources All data is fetched in real time from free, public APIs. No authentication required. | Source | Data | Models | Cache TTL | |--------|------|--------|-----------| | OpenRouter | Pricing, context lengths, modalities, release dates | 300+ | 1 hour | | SWE-bench | Coding benchmark (Verified leaderboard) | 30+ | 6 hours | | LM Arena | Human preference Elo ratings | 314+ | 6 hours | | OpenCompass VLM | Vision benchmarks: MMMU, MMBench, OCRBench, AI2D, MathVista | 284+ | 6 hours | | Aider Polyglot | Multi-language coding pass rate | 63+ | 6 hours | --- ## Context Cost MCP tool definitions and responses consume your LLM's context window. This server is designed to be lean: | Component | Tokens | |-----------|--------| | All 4 tool definitions | ~1,000 | | Typical tool response | ~250-400 | For comparison, most MCP servers that return raw JSON consume 3,000-10,000 tokens per response. Every response from llm-advisor-mcp is pre-formatted Markdown, keeping context costs roughly 10x lower. --- ## Architecture
┌──────────────────────────────────────────────┐ │ MCP Client (Claude, etc.) │ └──────────┬───────────────────────────────────┘ │ stdio (JSON-RPC) ┌──────────▼───────────────────────────────────┐ │ llm-advisor-mcp server │ │ │ │ ┌─────────┐ ┌───────────┐ ┌────────────┐ │ │ │ Tools │ │ Registry │ │ Cache │ │ │ │ (4 tools)│──│ (unified) │──│ (in-memory)│ │ │ └─────────┘ └───────────┘ └────────────┘ │ │ │ │ │ ┌────────────┼────────────┐ │ │ ▼ ▼ ▼ │ │ ┌──────────┐ ┌──────────┐ ┌──────────┐ │ │ │Normalizer│ │Percentile│ │ Fetchers │ │ │ │(slug map)│ │ (5 cats) │ │(5 sources│ │ │ └──────────┘ └──────────┘ └──────────┘ │ └──────────────────────────────────────────────┘ │ │ │ OpenRouter SWE-bench Arena / VLM / Aider
- TypeScript + ESM — Single entry point, tsup build - In-memory cache — TTL-based (1h pricing, 6h benchmarks), stale-while-revalidate - Cross-source normalization — Maps inconsistent model names (e.g. Claude 3.5 Sonnet vs anthropic/claude-3.5-sonnet) to canonical IDs - Percentile computation — Ranks across 5 categories (coding, math, general, vision, cost efficiency) - Freshness scoring — Recommendation algorithm gives a bonus to recently released models (+3 for <=3mo, +1 for <=6mo) - Zero runtime deps beyond @modelcontextprotocol/sdk and zod --- ## Roadmap | Version | Status | Highlights | |---------|--------|------------| | v0.1 | Done | get_model_info + list_top_models via OpenRouter | | v0.2 | Done | compare_models + recommend_model + SWE-bench + Arena Elo | | v0.3 | Done | VLM benchmarks (MMMU, MMBench, OCRBench, AI2D, MathVista) + Aider Polyglot + percentile ranks + 43 tests | | v0.4 | Current | Release date display, date-based filtering, freshness scoring in recommendations + 51 tests | | v1.0 | Planned | Community contributions, weekly static data snapshots via GitHub Actions | --- ## Development `bash git clone https://github.com/Daichi-Kudo/llm-advisor-mcp.git cd llm-advisor-mcp npm install npm run build # Build with tsup npm run dev # Run with tsx (hot reload) npm test # Run 51 unit tests (vitest) npm run test:watch # Watch mode
src/ index.ts # Server entry point types.ts # Shared type definitions tools/ model-info.ts # get_model_info tool list-top.ts # list_top_models tool compare.ts # compare_models tool recommend.ts # recommend_model tool formatters.ts # Markdown output formatters data/ registry.ts # Unified model registry cache.ts # In-memory TTL cache normalizer.ts # Cross-source name normalization percentiles.ts # Percentile rank computation fetchers/ openrouter.ts # OpenRouter API swe-bench.ts # SWE-bench leaderboard arena.ts # LM Arena Elo ratings vlm-leaderboard.ts # OpenCompass VLM benchmarks aider.ts # Aider Polyglot scores static/ api-examples.ts # API code snippet templates``
- Fork the repository
- Create a feature branch
- Add tests for new functionality
- Run
npm test`to verify all 51 tests pass- Submit a pull request
This is a web browser that enables your coding agent, such as Claude Code, to visit websites on your behalf and assist you in identifying bugs or creating UI test cases.
next-devtools-mcp is a MCP server that provides Next.js development tools and utilities for AI coding assistants like Claude and Cursor.
Word search, crossword, and sudoku generator MCP server with printable PDF worksheets, themed word banks, and verifiable LLM evals. Local-first, from the makers of puzzletide.com.
A demonstration server for ActionKit, providing access to Slack actions via Claude Desktop.
MCP server that lets Claude Code agents delegate tasks to agents in other project directories, with parallel dispatch, sessions, and async jobs.
Statistical regression testing for LLM agents: p-value, effect size, and CI on behavior change.
A Python MCP package that gives your LLM agents complete file system and shell capabilities — production-ready, sandboxed, and wired to any LLM in minutes.
Integrates with Google AI Studio/Gemini API for PDF to Markdown conversion and content generation.
Anchor Browser (https://anchorbrowser.io) is secure infrastructure for computer-use agents — stealth cloud browsers, authentication, captcha bypass, and a hosted MCP server for Cursor, Claude, and Windsurf.
Sign in to leave a review
Use Google, GitHub, or an email account so ratings stay tied to real people.
No reviews posted yet.




