AiCore Project
About
A unified framework for integrating various language models and embedding providers to generate text completions and embeddings.
Details
- Author
- brunov21
- Categories
- Developer Tools, AI, Other
Jump to
Setup
Install AiCore Project in your MCP client (Claude Desktop, Cursor, Windsurf, and others).
Repository: https://github.com/brunov21/AiCore
Follow the installation instructions in the repository README, then restart your MCP client.
✨AiCoreis a comprehensive framework for integrating various language models and embedding providers with a unified interface. It supports both synchronous and asynchronous operations for generating text completions and embeddings, featuring:
🔌Multi-provider support: OpenAI, Mistral, Groq, Gemini, NVIDIA, and more 🤖Reasoning augmentation: Enhance traditional LLMs with reasoning capabilities 📊Observability: Built-in monitoring and analytics 💰Token tracking: Detailed usage metrics and cost tracking ⚡Flexible deployment: Chainlit, FastAPI, and standalone script support 🛠️MCP Integration: Connect to Model Control Protocol servers via tool calling 🖥️Claude Code provider: Use your Claude subscription locally or remotely via the Claude Agents Python SDK — no API key required
pip install git+https://github.com/BrunoV21/AiCore
pip install git+https://github.com/BrunoV21/AiCore.git#egg=core-for-ai[all]
from aicore.llm import Llm from aicore.llm.config import LlmConfig import os llm_config = LlmConfig( provider="openai", model="gpt-4o", api_key="super_secret_openai_key" ) llm = Llm.from_config(llm_config) # Generate completion response = llm.complete("Hello, how are you?") print(response)
from aicore.llm import Llm from aicore.llm.config import LlmConfig import os async def main(): llm_config = LlmConfig( provider="openai", model="gpt-4o", api_key="super_secret_openai_key" ) llm = Llm.from_config(llm_config) # Generate completion response = await llm.acomplete("Hello, how are you?") print(response) if __name__ == "__main__": asyncio.run(main())
more examples available atexamples/anddocs/exampes/
- Anthropic
- OpenAI
- Mistral
- Groq
- Gemini
- NVIDIA
- OpenRouter
- DeepSeek
- Claude Code(local — via Claude Agents Python SDK, no API key required)
- Remote Claude Code(remote — connects to aaicore-proxy-serverover HTTP)
- Operation tracking and metrics collection
- Interactive dashboard for visualization
- Token usage and latency monitoring
- Cost tracking
- Connect to multiple MCP servers simultaneously
- Automatic tool discovery and calling
- Support for WebSocket, SSE, and stdio transports
To configure the application for testing, you need to set up aconfig.ymlfile with the necessary API keys and model names for each provider you intend to use. TheCONFIG_PATHenvironment variable should point to the location of this file. Here's an example of how to set up theconfig.ymlfile:
# config.yml embeddings: provider: "openai" # or "mistral", "groq", "gemini", "nvidia" api_key: "your_openai_api_key" model: "text-embedding-3-small" # Optional llm: provider: "openai" # or "mistral", "groq", "gemini", "nvidia" api_key: "your_openai_api_key" model: "gpt-o4" # Optional temperature: 0.1 max_tokens: 1028 reasonning_effort: "high" mcp_config: "./mcp_config.json" # Path to MCP configuration max_tool_calls_per_response: 3 # Optional limit on tool calls
config examples for the multiple providers are included in theconfig dir
from aicore.llm import Llm from aicore.config import Config import asyncio async def main(): # Load configuration with MCP settings config = Config.from_yaml("./config/config_example_mcp.yml") # Initialize LLM with MCP capabilities llm = Llm.from_config(config.llm) # Make async request that can use MCP-connected tools response = await llm.acomplete( "Search for latest news about AI advancements", system_prompt="Use available tools to gather information" ) print(response) asyncio.run(main())
Example MCP configuration (mcp_config.json):
{ "mcpServers": { "search-server": { "transport_type": "ws", "url": "ws://localhost:8080", "description": "WebSocket server for search functionality" }, "data-server": { "transport_type": "stdio", "command": "python", "args": ["data_server.py"], "description": "Local data processing server" }, "brave-search": { "command": "npx", "args": [ "-y", "@modelcontextprotocol/server-brave-search" ], "env": { "BRAVE_API_KEY": "SUPER-SECRET-BRAVE-SEARCH-API-KEY" } } } }
AiCore supports routing completions through yourClaude subscriptionvia theClaude Agents Python SDK. No Anthropic API key is required — auth is handled entirely by the Claude Code CLI. This is exposed through two providers and an optional proxy server:
Both providers share the sameacomplete()/complete()interface and emit identical tool-call streaming events — you can switch between them with a single config change.
Runsclaude-agent-sdkdirectly on the machine where AiCore is executing.
# 1. Install the Claude Code CLI (requires Node.js 18+) npm install -g @anthropic-ai/claude-code # 2. Authenticate once claude login # 3. Install AiCore (the Python SDK is included automatically) pip install core-for-ai
from aicore.llm import Llm from aicore.llm.config import LlmConfig config = LlmConfig( provider="claude_code", model="claude-sonnet-4-5-20250929", # No api_key needed — auth is handled by the CLI ) llm = Llm.from_config(config) response = await llm.acomplete("List all Python files in this project") print(response)
# config/config_example_claude_code.yml llm: provider: "claude_code" model: "claude-sonnet-4-5-20250929" # Optional permission_mode: "bypassPermissions" # default — all tools allowed cwd: "/path/to/your/project" # working directory for the CLI max_turns: 10 # limit agentic turns mcp_config: "./mcp_config.json" # pass through an MCP config file cli_path: "/usr/local/bin/claude" # override if CLI is not on PATH allowed_tools: - "Read" - "Write" - "Bash"
The proxy server wrapsclaude-agent-sdkin a FastAPI SSE service so Claude Code can be accessed remotely over HTTP. Useful when:
- The Claude Code CLI is authenticated on adifferent machine(e.g. a dev box, a server, or WSL)
- You want toshare a single Claude subscriptionacross multiple AiCore clients
- Your AiCore workload runs in acontainer or cloud environmentthat cannot run the CLI directly
# Install AiCore with the claude-server extras pip install core-for-ai[claude-server] # Also install the Claude Code CLI and authenticate npm install -g @anthropic-ai/claude-code claude login
The[claude-server]extra installsfastapi,uvicorn[standard], andpython-dotenv.pyngrokis optional and only needed for the ngrok tunnel mode.
# Minimal — binds to 127.0.0.1:8080, prompts for tunnel choice interactively aicore-proxy-server # Fully configured aicore-proxy-server \ --host 0.0.0.0 \ --port 8080 \ --token my-secret-token \ --tunnel none \ --cwd /path/to/project \ --log-level INFO # Or via Python module python -m aicore.scripts.claude_code_proxy_server --port 8080 --tunnel none
On first run the bearer token is auto-generated and printed. SetCLAUDE_PROXY_TOKENin your environment or.envfile to reuse it across restarts, or pass--tokenexplicitly.
When--tunnelis omitted the server prompts interactively at startup.
Connects AiCore to a runningaicore-proxy-serverover HTTP SSE. The remote provider reconstructs the SDK message stream locally, giving the sameacomplete()/complete()interface as the local provider — no Claude Code CLI needed on the client side.
pip install core-for-ai # no CLI or claude-server extras required
The proxy server must be running and reachable before instantiating the provider (aGET /healthcheck is performed automatically at startup, controllable viaskip_health_check).
from aicore.llm import Llm from aicore.llm.config import LlmConfig config = LlmConfig( provider="remote_claude_code", model="claude-sonnet-4-5-20250929", base_url="http://your-proxy-host:8080", # or a tunnel URL api_key="your_proxy_token", # CLAUDE_PROXY_TOKEN from server startup ) llm = Llm.from_config(config) response = await llm.acomplete("Summarise this codebase") print(response)
# config/config_example_remote_claude_code.yml llm: provider: "remote_claude_code" model: "claude-sonnet-4-5-20250929" base_url: "http://your-proxy-host:8080" # or the ngrok / cloudflare tunnel URL api_key: "your_proxy_token" # CLAUDE_PROXY_TOKEN printed at server startup # Optional — forwarded to the proxy server permission_mode: "bypassPermissions" cwd: "/path/to/project" # must be in server's --allowed-cwd-paths max_turns: 10 allowed_tools: - "Bash" - "Read" - "Write" # Skip the GET /health connectivity check at startup skip_health_check: false
Bothclaude_codeandremote_claude_codeemit identical tool-call events:
def on_tool_event(event: dict): if event["stage"] == "started": print(f"→ Calling tool: {event['tool_name']}") elif event["stage"] == "concluded": status = "✗" if event["is_error"] else "✓" print(f"{status} Tool finished: {event['tool_name']}") llm.tool_callback = on_tool_event response = await llm.acomplete("Find all TODO comments in the codebase")
TOOL_CALL_START_TOKEN/TOOL_CALL_END_TOKENare also emitted viastream_handler, so any existing stream consumer works without changes.
Note:temperature,max_tokens, andapi_keyare ignored by both providers — the Claude Code CLI controls model parameters internally. Cost is reported fromResultMessage.total_cost_usdrather than computed from a pricing table.
You can use the language models to generate text completions. Below is an example of how to use theMistralLlmprovider:
from aicore.llm.config import LlmConfig from aicore.llm.providers import MistralLlm config = LlmConfig( api_key="your_api_key", model="your_model_name", temperature=0.7, max_tokens=100 ) mistral_llm = MistralLlm.from_config(config) response = mistral_llm.complete(prompt="Hello, how are you?") print(response)
To load configurations from a YAML file, set theCONFIG_PATHenvironment variable and use theConfigclass to load the configurations. Here is an example:
from aicore.config import Config from aicore.llm import Llm import os if __name__ == "__main__": os.environ["CONFIG_PATH"] = "./config/config.yml" config = Config.from_yaml() llm = Llm.from_config(config.llm) llm.complete("Once upon a time, there was a")
Make sure yourconfig.ymlfile is properly set up with the necessary configurations.
AiCore includes a comprehensive observability module that tracks:
- Request/response metadata
- Token usage(prompt, completion, total)
- Latency metrics(response time, time-to-first-token)
- Cost estimates(based on provider pricing)
- Tool call statistics(for MCP integrations)
- Requests per minute
- Average response time
- Token usage trends
- Error rates
- Cost projections
from aicore.observability import ObservabilityDashboard dashboard = ObservabilityDashboard(storage="observability_data.json") dashboard.run_server(port=8050)
AiCore also contains native support to augmenttraditionalLlms withreasoningcapabilities by providing them with the thinking steps generated by an open-source reasoning capable model, allowing it to generate its answers in a Reasoning Augmented way.
This can be usefull in multiple scenarios, such as:
- ensure your agentic systems still work with the propmts you have crafted for your favourite llms while augmenting them with reasoning steps
- direct control for how long you want your reasoner to reason (via max_tokens param) and how creative it can be (reasoning temperature decoupled from generation temperature) without compromising generation settings
To leverage the reasoning augmentation just introduce one of the supported llm configs into the reasoner field and AiCore handles the rest
# config.yml embeddings: provider: "openai" # or "mistral", "groq", "gemini", "nvidia" api_key: "your_openai_api_key" model: "your_openai_embedding_model" # Optional llm: provider: "mistral" # or "openai", "groq", "gemini", "nvidia" api_key: "your_mistral_api_key" model: "mistral-small-latest" # Optional temperature: 0.6 max_tokens: 2048 reasoner: provider: "groq" # or openrouter or nvidia api_key: "your_groq_api_key" model: "deepseek-r1-distill-llama-70b" # or "deepseek/deepseek-r1:free" or "deepseek/deepseek-r1" temperature: 0.5 max_tokens: 1024
A Hugging Face Space showcasing reasoning-augmented models
Instant summaries of Git activity
🌐Live App
📦GitHub Repository
CodeTideis a fully local, privacy-first tool for parsing and understanding Python codebases using symbolic, structural analysis—no LLMs, no embeddings, just fast and deterministic code intelligence. It enables developers and AI agents to retrieve precise code context, visualize project structure, and generate atomic code changes with confidence.
AgentTideis a next-generation, precision-driven software engineering agent built on top of CodeTide. AgentTide leverages CodeTide’s symbolic code understanding to plan, generate, and apply high-quality code patches—always with full context and requirements fidelity. You can interact with AgentTide via a conversational CLI or a beautiful web UI.
Live Demo:Try AgentTide on Hugging Face Spaces:https://mclovinittt-agenttidedemo.hf.space/
AiCorewas used to make LLM calls within AgentTide, enabling seamless integration between local code analysis and advanced language models. This combination empowers AgentTide to deliver context-aware, production-ready code changes—always under your control.
- Extended Provider Support: Additional LLM and embedding providers
- Add support for Speech: Integrate text2speech and speech to text objects with usage and observability4
For complete documentation, including API references, advanced usage examples, and configuration guides, visit:
This project is licensed under the Apache 2.0 License.
This is a web browser that enables your coding agent, such as Claude Code, to visit websites on your behalf and assist you in identifying bugs or creating UI test cases.
MCP server for xAI Grok API — chat, vision, search, and embeddings (Rust, stdio)
Offline MCP server that ranks & summarizes code using BM25, TF-IDF, embeddings & git signals; integrates with Cursor, Claude Desktop and Windsurf; privacy preserving.
A fully-local MCP server for question-answering over your PDFs. Ask in plain language; Claude retrieves only the relevant passages with page citations. On-device embeddings (sentence-transformers) + ChromaDB — no API keys, nothing leaves your machine.
Private persistent memory for Claude, ChatGPT & Gemini via MCP — semantic search, zero-code setup.
An MCP server for web and similarity search, designed for Claude Desktop. It integrates with various external embedding and API services.
MCP server for Prompt Builder — search, retrieve, and compile prompt components from a community vault using semantic search (pgvector) and slug-based lookup. Works with Claude Desktop and Cursor.
A server providing web and similarity search functionalities, designed for Claude Desktop. It requires external embedding and API services.
Persistent visual cache for LLM-driven software development. Caches screenshots using perceptual hashing, vector search, and AX trees to prevent token overhead and visual hallucination loops.
Query and analyze your Opik logs, traces, prompts and all other telemtry data from your LLMs in natural language.
Sign in to leave a review
Use Google, GitHub, or an email account so ratings stay tied to real people.
No reviews posted yet.





