Inferbench
About
InferBench's MCP server lets coding agents run, serve and benchmark local LLMs (text + image, llama.cpp + Stable Diffusion) on your own hardware on demand. Measures real tokens/sec, picks the optimal quant for your GPU, and exposes a 124-model catalog. Local-first, no cloud requi
Details
- Author
- JoniMartin27
- Downloads
- 318
- Categories
- Other, AI
Jump to
- Run, serve, and benchmark local LLMs on demand
- Measures real tokens per second throughput
- Selects optimal quantization for your GPU automatically
- Supports text and image models (llama.cpp + Stable Diffusion)
- Exposes a catalog of 124 models
- Operates entirely local-first, no cloud required
—
InferBench's MCP server lets coding agents run, serve and benchmark local LLMs (text + image, llama.cpp + Stable Diffusion) on your own hardware on demand. Measures real tokens/sec, picks the optimal quant for your GPU, and exposes a 124-model catalog. Local-first, no cloud required.
Sign in to leave a review
Use Google, GitHub, or an email account so ratings stay tied to real people.
No reviews posted yet.




