Gemini Image Generator MCP Server

by qhdrl12

33 stars
377 downloads
Not rated
GitHub

About

MCP server for AI image generation and editing using Google's Gemini Flash models. Create images from text prompts with intelligent filename generation and strict text exclusion. Supports text-to-image generation with future expansion to image editing capabilities.

Details

Author
qhdrl12
GitHub stars
33
Downloads
377
Categories
Developer Tools, Other, AI

- Text-to-image generation using Gemini 2.0 Flash
- Image-to-image transformation based on text prompts
- Support for both file-based and base64-encoded images
- Automatic intelligent filename generation from prompts
- Automatic translation of non-English prompts
- Local image storage with configurable output path

Setting up with Highlight

This MCP is not yet compatible with Highlight’s one-click setup. However, you can still use it with Highlight by following these steps:

  1. Download and install Highlight from highlightai.com/download
  2. Navigate to the plugins tab and select "Add Custom Plugin"
  3. Configure the plugin with the settings below
    Plugin Name Gemini Image Generator MCP Server
    Command (node, npx, python, etc.)

    Please refer to the README for specific instructions on how to obtain API keys or other required environment variables.

  4. Enable "Start Automatically" if you want the plugin to start when Highlight launches

From the repository

Install via Smithery (npx -y @smithery/cli install @qhdrl12/mcp-server-gemini-image-gen --client claude), or manually clone the repo, create a Python 3.11+ virtual environment, and install dependencies. Configure a Gemini API key (via Google AI Studio) and an optional image output path. Add the server to your MCP client’s configuration (e.g., claude_desktop_config.json) with the command, args, and env keys. Invoke the three available MCP tools: generate_image_from_text, transform_image_from_encoded, or transform_image_from_file.

Claude Desktop / Cursor

Paste into your MCP client config file to install this server.

{
    "mcpServers": {
        "gemini image generator mcp server": {
            "mcp-server-gemini-image-generator": {
                "command": "npx",
                "args": [
                    "-y",
                    "@smithery/cli",
                    "install",
                    "@qhdrl12/mcp-server-gemini-image-gen",
                    "--client",
                    "claude"
                ]
            }
        }
    }
}

McpServers

{
    "mcp-server-gemini-image-generator": {
        "command": "npx",
        "args": [
            "-y",
            "@smithery/cli",
            "install",
            "@qhdrl12/mcp-server-gemini-image-gen",
            "--client",
            "claude"
        ]
    }
}

Gemini Image Generator MCP Server

Generate high-quality images from text prompts using Google's Gemini model through the MCP protocol.

Overview

This MCP server allows any AI assistant to generate images using Google's Gemini AI model. The server handles prompt engineering, text-to-image conversion, filename generation, and local image storage, making it easy to create and manage AI-generated images through any MCP client.

Features

- Text-to-image generation using Gemini 2.0 Flash
- Image-to-image transformation based on text prompts
- Support for both file-based and base64-encoded images
- Automatic intelligent filename generation based on prompts
- Automatic translation of non-English prompts
- Local image storage with configurable output path
- Strict text exclusion from generated images
- High-resolution image output
- Direct access to both image data and file path

Available MCP Tools

The server provides the following MCP tools for AI assistants:

1. generate_image_from_text

Creates a new image from a text prompt description.

generate_image_from_text(prompt: str) -> Tuple[bytes, str]

Parameters:
- prompt: Text description of the image you want to generate

Returns:
- A tuple containing:
- Raw image data (bytes)
- Path to the saved image file (str)

This dual return format allows AI assistants to either work with the image data directly or reference the saved file path.

Examples:
- "Generate an image of a sunset over mountains"
- "Create a photorealistic flying pig in a sci-fi city"

Example Output

This image was generated using the prompt:

"Hi, can you create a 3d rendered image of a pig with wings and a top hat flying over a happy futuristic scifi city with lots of greenery?"

Flying pig over sci-fi city

A 3D rendered pig with wings and a top hat flying over a futuristic sci-fi city filled with greenery

Known Issues

When using this MCP server with Claude Desktop Host:

1. Performance Issues: Using transform_image_from_encoded may take significantly longer to process compared to other methods. This is due to the overhead of transferring large base64-encoded image data through the MCP protocol.

2. Path Resolution Problems: There may be issues with correctly resolving image paths when using Claude Desktop Host. The host application might not properly interpret the returned file paths, making it difficult to access the generated images.

For the best experience, consider using alternative MCP clients or the transform_image_from_file method when possible.

2. transform_image_from_encoded

Transforms an existing image based on a text prompt using base64-encoded image data.

transform_image_from_encoded(encoded_image: str, prompt: str) -> Tuple[bytes, str]

Parameters:
- encoded_image: Base64 encoded image data with format header (must be in format: "data:image/[format];base64,[data]")
- prompt: Text description of how you want to transform the image

Returns:
- A tuple containing:
- Raw transformed image data (bytes)
- Path to the saved transformed image file (str)

Example:
- "Add snow to this landscape"
- "Change the background to a beach"

3. transform_image_from_file

Transforms an existing image file based on a text prompt.

transform_image_from_file(image_file_path: str, prompt: str) -> Tuple[bytes, str]

Parameters:
- image_file_path: Path to the image file to be transformed
- prompt: Text description of how you want to transform the image

Returns:
- A tuple containing:
- Raw transformed image data (bytes)
- Path to the saved transformed image file (str)

Examples:
- "Add a llama next to the person in this image"
- "Make this daytime scene look like night time"

Example Transformation

Using the flying pig image created above, we applied a transformation with the following prompt:

"Add a cute baby whale flying alongside the pig"

Before:
Flying pig over sci-fi city

After:
Flying pig with baby whale

The original flying pig image with a cute baby whale added flying alongside it

Setup

Prerequisites

- Python 3.11+
- Google AI API key (Gemini)
- MCP host application (Claude Desktop App, Cursor, or other MCP-compatible clients)

Getting a Gemini API Key

1. Visit Google AI Studio API Keys page
2. Sign in with your Google account
3. Click "Create API Key"
4. Copy your new API key for use in the configuration
5. Note: The API key provides a certain quota of free usage per month. You can check your usage in the Google AI Studio

Installation

Installing via Smithery

To install Gemini Image Generator MCP for Claude Desktop automatically via Smithery:

npx -y @smithery/cli install @qhdrl12/mcp-server-gemini-image-gen --client claude

Manual Installation

1. Clone the repository:
git clone https://github.com/your-username/mcp-server-gemini-image-generator.git
cd mcp-server-gemini-image-generator

2. Create a virtual environment and install dependencies:
```bash

No reviews yet — be the first

Sign in to leave a review

Use Google, GitHub, or an email account so ratings stay tied to real people.

Email sign in

No reviews posted yet.