Inferbench

by JoniMartin27

318 downloads
Not rated
GitHub

About

InferBench's MCP server lets coding agents run, serve and benchmark local LLMs (text + image, llama.cpp + Stable Diffusion) on your own hardware on demand. Measures real tokens/sec, picks the optimal quant for your GPU, and exposes a 124-model catalog. Local-first, no cloud requi

Details

Author
JoniMartin27
Downloads
318
Categories
Other, AI

- Run, serve, and benchmark local LLMs on demand
- Measures real tokens per second throughput
- Selects optimal quantization for your GPU automatically
- Supports text and image models (llama.cpp + Stable Diffusion)
- Exposes a catalog of 124 models
- Operates entirely local-first, no cloud required

InferBench's MCP server lets coding agents run, serve and benchmark local LLMs (text + image, llama.cpp + Stable Diffusion) on your own hardware on demand. Measures real tokens/sec, picks the optimal quant for your GPU, and exposes a 124-model catalog. Local-first, no cloud required.

No reviews yet — be the first

Sign in to leave a review

Use Google, GitHub, or an email account so ratings stay tied to real people.

Email sign in

No reviews posted yet.