🎉 Smithery is now a part of Arcade.dev! Read more in our announcement here
Checks whether an open-weight LLM fits your GPU: VRAM for weights, KV cache and overhead at each quantisation, a tokens-per-second ceiling, the longest context that fits, and what would work instead...