soundside.ai
About
MCP-native AI media generation with x402 pay-per-call. Image, video, audio, and music from 6 providers — composable via resource IDs. USDC on Base.
Details
- Author
- soundside-design
- Categories
- Cloud Service, Media, Other, AI
Jump to
Setup
Install soundside.ai in your MCP client (Claude Desktop, Cursor, Windsurf, and others).
Repository: https://github.com/soundside-design/soundside-docs
Follow the installation instructions in the repository README, then restart your MCP client.
- Generate images from text prompts or character referencesusingcreate_imageacross multiple providers like Alibaba, Grok, Luma, MiniMax, Runway, or Vertex AI.
- Create videos from text, images, or extend existing clipswithcreate_video, choosing from providers such as Alibaba, Grok, Luma, MiniMax, Runway, or Vertex AI (Veo 3.1).
- Produce audio including TTS, sound effects, and voice cloningviacreate_audiobacked by MiniMax, Runway, or Vertex AI.
- Compose music from lyrics and style promptsusingcreate_music.
- Run a server-side video composition pipelinewithcompose_videothat enriches a plan, generates assets in parallel, and assembles with transitions and audio ducking.
- Train custom LoRA adapters from your library mediausingtrain_adapter, then inspect or manage them withlist_adaptersandmanage_adapter.
Soundside exposes 19 MCP tools for generating, editing, composing, extracting, and analyzing media — images, video, audio, music, text, and business artifacts — plus LoRA adapter fine-tuning and server-side video composition. Connect any MCP client. Pay with an API key (credits) or crypto (x402 USDC on Base, no account needed).
# MCP endpoint https://mcp.soundside.ai/mcp # Auth: API key or x402 crypto payment Authorization: Bearer <your-api-key>
POST https://mcp.soundside.ai/mcp {"jsonrpc":"2.0","id":"1","method":"tools/list","params":{}}
Soundside aims to break even on provider pass-through costs with a small margin (~10%). The editing engine and library are priced at $0.01/call; vision QA is $0.03.
GET https://mcp.soundside.ai/api/x402/status
This returns machine-readable per-tool, per-provider USDC prices. Prices are DB-driven and may change —always check the endpoint rather than hardcoding.
No API key needed. Pay with USDC on Base (L2) per tool call via EIP-3009transferWithAuthorization(off-chain signing, facilitator pays gas).
Network: eip155:8453 (Base mainnet) Token: USDC Facilitator: Coinbase CDP
- Getting Started— First MCP connection in 5 minutes
- x402 Pay-Per-Call— Crypto payments, no account needed
- Tool Reference— Detailed docs for all 19 tools
- Python — API Key— Connect and generate with httpx
- Python — x402— Pay-per-call with USDC
- TypeScript — API Key— Node.js MCP client
- OpenClaw Skill— One-line config for OpenClaw agents
- Website:soundside.ai
- MCP Endpoint:https://mcp.soundside.ai/mcp
- Live Pricing:https://mcp.soundside.ai/api/x402/status
- GitHub:github.com/soundside-design/soundside-docs
Turn any language model into a multimodal powerhouse that can generate images, music, videos and more on the fly. Rostro's tools are designed to be used by language models from the ground up, expanding capabilities with minimal context bloat.
AI audio tools for music producers — stem splitting, vocal removal, BPM/key detection, audio-to-MIDI, format conversion and AI song generation
All-in-one AI creative studio — generate videos, images, audio in 11 Indian languages, and 3D models via MCP. Hosted at mcp.arcframe.ai.
Official Cannon Studio MCP for AI video, image, 3D, audio, workflow, pricing, model, and developer API guidance.
A server for creating fast and free lipsync videos for digital avatars, supporting both realistic and cartoon styles.
MCP server for AI image, video, voice and music generation, routing prompts to Veo 3.1, Seedance 2.0, Nano Banana 2, GPT-Image-2 and more, from Claude, Cursor, ChatGPT or Codex.
35+ local AI tools - TCG card grading, Monte Carlo simulation, voice synthesis, 3D mesh, image gen, and autonomous M2M NFT purchase bridge.
: MCP server for AI media generation (imagesflux, videosveo3.1, music suno v5, with deterministic cost control using reserve-burn-refund billing
Analyzes image and video content from URLs or local files using the Gemini 2.0 Flash model.
An MCP server that integrates with the Jimeng AI image generation service.
Sign in to leave a review
Use Google, GitHub, or an email account so ratings stay tied to real people.
No reviews posted yet.





