Python Code Validator

by jkanselaar

Not rated
GitHub

About

Proves AI-generated Python does what you asked: state the intent as assertions and the server runs the code in a locked-down container, returning a fix only when every example passes.

Details

Author
jkanselaar
Categories
Developer Tools

Setup

Install Python Code Validator in your MCP client (Claude Desktop, Cursor, Windsurf, and others).

Repository: https://github.com/jkanselaar/python-code-validator

Follow the installation instructions in the repository README, then restart your MCP client.

An MCP server that validates, repairs and runs Python against the examples it is supposed to satisfy —validate_python,repair_pythonandexecute_pythonover HTTP athttps://api.statemind.ai/mcp, with a free key and no account.

A hosted service that proves AI-generated Python does what you asked. State the intent — assertions or doctest lines — and the code is run against it inside a container with no network and a read-only filesystem; a fix comes back only when every example passes. On the QuixBugs defects that is 41% repaired and 77% refused as not doing what they say, with no false alarms on the corrected programs — whereruffandmypyflag the defect in none of them (the numbers).

The checks that need no intent come with it: syntax and lint diagnostics, an AST security policy that also catches calls hidden behind dynamic imports and runtime attribute lookups, a bandit pass, a credential scan and deterministic repair — one verdict with a score. Asking the same question twice inside ten minutes is answered from the first answer and costs nothing (x-msvc-repeat: 1).

This repository holds the client side: the MCP configuration, the CI script and the pre-commit hook. The service itself runs athttps://api.statemind.ai, so there is nothing to install or host.

curl -s -X POST https://api.statemind.ai/v1/keys # {"api_key": "msvc_free_…", "tier": "free", "calls_per_day": 25, "modes": ["static"]}

25 static checks a day, metered per UTC day, and a few keys per address: enough to try it and to run it over a small project, not a supply. Every answer carries the state of the allowance (x-quota-remaining,x-quota-reset), so a client can back off before it is cut off.

Registered in the official MCP registry asai.statemind/python-code-validator, a name verified against the domain that serves it rather than a GitHub account. Any MCP client adds it with one block:

{ "mcpServers": { "python-code-validator": { "type": "http", "url": "https://api.statemind.ai/mcp", "headers": { "Authorization": "Bearer msvc_free_…" } } } }

- Claude Code:claude mcp add --transport http python-code-validator https://api.statemind.ai/mcp --header "Authorization: Bearer msvc_free_…"
- Cursor:~/.cursor/mcp.json, same block.
- VS Code / Copilot:.vscode/mcp.jsonunder"servers".

A client that only launches a command uses the stdio bridge in this repository instead, which forwards the same tool over HTTPS:

{ "mcpServers": { "python-code-validator": { "command": "python3", "args": ["/path/to/python-code-validator/mcp_stdio.py"] } } }

Or as a container, which theDockerfilehere builds:

docker build -t python-code-validator . docker run -i --rm -e VALIDATOR_API_KEY python-code-validator

Gemini CLI installs the same bridge as an extension, with the instruction file that makes it get used:

gemini extensions install jkanselaar/python-code-validator

Three tools, named after what they do to the code:

The old singlepython_code_validatortool, with itsmodeargument, still answers for clients that already configured it, but is no longer listed.

Every check above passes on a function that computes the wrong answer. The one thing that catches it is the intent, and the agent that asked for the code is the only one who has it — so pass it along:

{"code": "def bitcount(n): …", "mode": "execute", "options": {"examples": "assert bitcount(127) == 7"}}

Doctest lines (>>> bitcount(127)then7) work the same way, as do>>>examples already written in the source.execute_pythonruns them in the sandbox: one that does not hold is apython:example-mismatcherror, and the repair search returns a fix only when every example passes. On the QuixBugs defect set — real bugs, hidden test inputs deciding correctness — that repairs 41% and refuses 77% as not doing what they say, with no false alarms on the corrected programs.

Repeating a call costs nothing: the same key asking the same question — same mode, same code, same examples — is answered from the answer it already got, markedx-msvc-repeat: 1, so an agent that checks its work at every step is not billed for verdicts that cannot have changed.

An instruction can be ignored; a hook cannot. The plugin checks every Python file Claude Code writes or edits, in the turn it was written, and hands the errors back to the model instead of to you:

/plugin marketplace add jkanselaar/python-code-validator /plugin install python-code-validator@statemind

Nothing to configure: it mints and keeps its own free key on first use. A file that comes back accepted is silent, a rejected one stops the turn with the offending lines named, and an identical file is not asked about twice. It never ends a session over its own trouble — an unreachable service or a spent allowance lets the turn continue, and the allowance says how to raise it.

SetVALIDATOR_API_KEYto use a paid key instead of the free tier, andVALIDATOR_URLto point at your own deployment. The plugin also carries thevalidate-pythonskill, for the part a hook cannot do: stating the intent as examples and running the code against them.

The same script, wired to Cursor'spostToolUse, where the verdict comes back as context on the conversation instead of as an exit code:

mkdir -p .cursor/hooks base=https://raw.githubusercontent.com/jkanselaar/python-code-validator/main curl -sf $base/plugin/hooks/validate_written.py -o .cursor/hooks/validate_written.py curl -sf $base/cursor/hooks.json -o .cursor/hooks.json

Project hooks run from the project root, which is why the command incursor/hooks.jsonis a path relative to it. For a hook that applies to every project instead, put the script in~/.cursor/hooks/and the same block in~/.cursor/hooks.jsonwith the commandpython3 ./hooks/validate_written.py --cursor.

Configuring the server is not what gets it called: the instruction file is.AGENTS.mdin this repository is that text, written to be dropped into any project under whichever name the client reads:

mkdir -p .github curl -sf https://raw.githubusercontent.com/jkanselaar/python-code-validator/main/AGENTS.md \ | tee AGENTS.md CLAUDE.md GEMINI.md .github/copilot-instructions.md >/dev/null

Cursor reads rules with front matter instead, so that one is a separate file — copy.cursor/rules/python-code-validator.mdcinto.cursor/rules/of the project.

The short version, if you would rather add a line to instructions you already have:

Write what the code should do asassertexamples before writing the code, and pass them inoptions.examples. Callvalidate_pythonafter every edit andexecute_pythononce a function is finished, not again until what it does has changed. When a call returnsfixed_code, take it — the service ran it against your examples. Do not present code that came backvalid: false.

The service hands out the client, so a workflow needs no checkout of this repository and no secret:

- run: | curl -sf https://api.statemind.ai/v1/client -o validate.py python3 validate.py --changed-against "origin/${{ github.base_ref }}"
permissions: contents: read pull-requests: write # so the run can comment its result on the pull request steps: - uses: jkanselaar/python-code-validator@v1.22.0 with: api-key: ${{ secrets.VALIDATOR_API_KEY }} # optional; free tier without it

The changed Python is validated and offending lines are annotated on the diff, failing the job on syntax errors and unsafe patterns. Files the service refuses outright (over its 200 kB limit) are skipped with a warning rather than failing the run.

The run also leaves one comment on the pull request, edited in place on later pushes rather than repeated: what was accepted, what was repaired and how much of the day's allowance is left. Withoutpull-requests: writenothing is written and the job is unaffected;comment: "false"turns it off.

On the free tier the action keeps its key in the workflow cache, one per repository per day, so the allowance belongs to the repository rather than to the run. Withapi-keyset the cache is skipped.

repos: - repo: https://github.com/jkanselaar/python-code-validator rev: v1.22.0 hooks: - id: python-code-validator

validate.pyis standard library only, so it also works aspython validate.py file.pyin a Makefile, a git hook or a container:

$ python3 validate.py service.py ::error file=service.py,line=88,title=SyntaxError::invalid syntax FAIL service.py score=0.66 0/1 files accepted

VALIDATOR_API_KEYis used when set; otherwise the client mints a free key — keeping it inVALIDATOR_KEY_FILEwhen that names a path, which is how a series of runs shares one allowance.VALIDATOR_URLpoints it at another deployment.VALIDATOR_SOURCEnames the caller, which is only ever counted: a run inside a workflow saysgithub-actionby itself.

A repository whose Python is checked on every pull request can say so:

Python validated
curl -s https://api.statemind.ai/v1/validate \ -H "Authorization: Bearer $VALIDATOR_API_KEY" \ -H 'content-type: application/json' \ -d '{"code": "def f(:\n pass\n", "mode": "static"}'

modeisstatic,repairorexecute;repairandexecuteneed a configured key. Submitted code is not logged.

A refused call says what to do about it, so a caller with no operator to ask can resolve it itself:

{"error": "payment_required", "remedy": {"action": "upgrade_key", "hint": "A free key covers static only. …"}}

A free key covers 25 static checks a day, and one address gets a few keys a day, so the allowance is a trial rather than a supply. Beyond it a key carries credits: a static check costs 1, a repair 3 and a sandboxed run 10, and an identical call repeated within ten minutes is answered from the first one for free.

Credits are bought with a card, without an invoice or anyone to ask:

curl -s -X POST https://api.statemind.ai/v1/keys/checkout \ -H 'content-type: application/json' \ -d '{"api_key": "'"$VALIDATOR_API_KEY"'", "credits": 500}'

That answers with a Stripe Checkout page; the credits are on the key seconds after the card clears (500 credits is €10). An agent with a Gnosis wallet can instead pay in xDAI without a browser —GET /v1/pricingstates both routes.

examples/holds three files and the client to send them with: one that passes every check and still returns the wrong number, one the security policy refuses, and one that comes back accepted from the sandbox.

This is a web browser that enables your coding agent, such as Claude Code, to visit websites on your behalf and assist you in identifying bugs or creating UI test cases.

Create crafted UI components inspired by the best 21st.dev design engineers.

Bring agent evaluations, observability, and synthetic test set generation directly into your IDE for free with Galileo's new MCP server

An MCP server to help AI assistants to answer questions and generate AccelByte Extend SDK code more effectively .

MCP server for AI Diagram Maker — generate beautiful software engineering diagrams directly inside Cursor, Claude Desktop, Claude Code, or any MCP-compatible AI agent

ALAPI MCP Tools,Call hundreds of API interfaces via MCP

AI-powered SVG animation generator that transforms static files into animated SVG components using the Allyson platform

MCP server that gives AI assistants on-demand access to 1,500+ amCharts docs, ~300 code examples, and 1000+ class API references.

APIMatic MCP Server is used to validate OpenAPI specifications using APIMatic. The server processes OpenAPI files and returns validation summaries by leveraging APIMatic’s API.

One shared context layer for AI agents and humans — live API specs, DB schemas, and versioned contracts across repos so every agent and teammate works from the same source of truth.

Build and deploy full-stack Next.js apps with 98 tools for React, AWS, and MongoDB

No reviews yet — be the first

Sign in to leave a review

Use Google, GitHub, or an email account so ratings stay tied to real people.

Email sign in

No reviews posted yet.