Free local AI tools

Runyard tools
built around local LLM decisions.

Every tool here computes from the same arithmetic: weights are parameters × bits-per-weight ÷ 8, plus a KV-cache term for your context, plus runtime overhead. Throughput is memory bandwidth divided by bytes read per token. Nothing is a lookup table, and nothing asks you to sign up.

Priority tools

The highest-intent local LLM utilities are live as full tool pages or structured gateway pages.

54 ready
VRAM Calculator for LLMs — Runyard tool

VRAM Calculator for LLMs

Live

Find exactly how much VRAM any model needs at any quantization level and context length.

TurboQuant-aware sizingDense and MoE friendly
Open details ->
Context Window Memory Calculator — Runyard tool

Context Window Memory Calculator

Live

See how larger context windows change KV cache growth, memory pressure, and practical fit.

Built for long chat and RAG use casesExplains why 128K is expensive locally
Open details ->
Quantization Picker — Runyard tool

Quantization Picker

Live

Choose the right quant level for your hardware, quality target, and inference speed goals.

Friendly for first-time local usersBalances fit and answer quality
Open details ->
GPU-to-Model Fit Checker — Runyard tool

GPU-to-Model Fit Checker

Gateway

Match your GPU directly to the best-fit models and skip dead-end downloads.

Direct gateway to Model RadarUseful for both NVIDIA and Apple setups
Open gateway ->
Ollama Context Length Calculator — Runyard tool

Ollama Context Length Calculator

Live

Estimate what your chosen Ollama context setting means for memory use and local stability.

Explains the hidden cost of contextUseful for Modelfile tuning
Open details ->
Tokens-per-Second Estimator — Runyard tool

Tokens-per-Second Estimator

Live

Estimate likely local inference speed before you commit to a model and runtime stack.

Helpful before a long model downloadFrames speed as a decision input
Open details ->
Can My PC Run This AI Model? — Runyard tool

Can My PC Run This AI Model?

Gateway

A guided compatibility check that funnels users to Runyard home for the live answer.

Simple wording for broad search intentActs as a product gateway
Open gateway ->
Model Size to RAM / VRAM Converter — Runyard tool

Model Size to RAM / VRAM Converter

Live

Convert model parameter counts and quant formats into memory estimates you can reason about.

Good companion to model cardsMakes 7B versus 14B more tangible
Open details ->
CPU vs GPU Offload Calculator — Runyard tool

CPU vs GPU Offload Calculator

Live

Estimate whether partial GPU offload will help or just create a slow, awkward compromise.

Useful for borderline hardwareFrames latency costs clearly
Open details ->
GGUF Variant Chooser — Runyard tool

GGUF Variant Chooser

Live

Pick the right GGUF file variant without memorizing every quant suffix and packaging nuance.

Friendly for Hugging Face browsingReduces wasted downloads
Open details ->
Best Model for My Hardware Finder — Runyard tool

Best Model for My Hardware Finder

Gateway

A recommendation-first landing page that points users back to Runyard home for the live shortlist.

Designed for recommendation intentSearch-friendly framing
Open gateway ->
Prompt Token Counter — Runyard tool

Prompt Token Counter

Live

Count prompt size, estimate context consumption, and catch oversized inputs before they fail.

Useful for prompt-heavy workflowsGreat for debugging truncation
Open details ->
Model Comparison Matrix — Runyard tool

Model Comparison Matrix

Live

Compare model families across fit, speed, context, and practical use-case tradeoffs.

Built for decisions, not just statsUseful before longer tests
Open details ->
Local LLM Cost Savings Calculator — Runyard tool

Local LLM Cost Savings Calculator

Live

Estimate how much local inference can save compared with repeated paid API usage.

Useful for teams and solo buildersFrames costs in plain language
Open details ->
Power & Heat Calculator — Runyard tool

Power & Heat Calculator

Live

Estimate watts, monthly electricity cost, and BTU/hr heat for any local LLM rig.

Covers RTX 50/40/30, H100, MI300X, Apple SiliconDuty-cycle aware (small models on big GPUs)
Open details ->
API Cost vs Local Cost Calculator — Runyard tool

API Cost vs Local Cost Calculator

Live

Compare monthly OpenAI / Anthropic / Google API spend against a local GPU rig — find the break-even in months.

Covers GPT-4o, GPT-4.1, Claude Sonnet, Opus, Gemini, DeepSeekRecommendation banner: local wins or API wins
Open details ->
GGUF File Size Estimator — Runyard tool

GGUF File Size Estimator

Live

Estimate the on-disk size of any GGUF file by parameter count and quant format, plus VRAM at load.

14 quant formats including IQ4_XS, IQ3_XS, IQ2_XSIncludes 3% GGUF metadata overhead
Open details ->
OOM Fix Assistant — Runyard tool

OOM Fix Assistant

Live

Diagnose likely out-of-memory causes and suggest the fastest fixes for local inference failures.

Good for debugging under pressureUseful for Ollama and GGUF users
Open details ->
QLoRA VRAM Calculator — Runyard tool

QLoRA VRAM Calculator

Live

QLoRA, LoRA and full fine-tuning memory for any model size, against your card.

Fine-tuning2 inputs
Open details ->
LoRA vs QLoRA vs Full Fine-Tune — Runyard tool

LoRA vs QLoRA vs Full Fine-Tune

Live

Memory, fidelity and hardware for all three, at every model size.

Fine-tuning1 inputs
Open details ->
Fine-Tune Batch Size Calculator — Runyard tool

Fine-Tune Batch Size Calculator

Live

How much memory is left for activations after the model, and what that buys.

Fine-tuning3 inputs
Open details ->
Training Time Estimator — Runyard tool

Training Time Estimator

Live

Wall-clock estimate from dataset size, epochs and your hardware.

Fine-tuning4 inputs
Open details ->
llama.cpp Flags Generator — Runyard tool

llama.cpp Flags Generator

Live

A complete, copy-pasteable command sized to your card and context.

Runtime & errors4 inputs
Open details ->
vLLM Memory Calculator — Runyard tool

vLLM Memory Calculator

Live

Work out the fraction to set, and how many concurrent sequences it buys.

Runtime & errors4 inputs
Open details ->
Concurrent Requests Calculator — Runyard tool

Concurrent Requests Calculator

Live

Concurrency from KV cache headroom, and what it does to per-user speed.

Runtime & errors4 inputs
Open details ->
Multi-GPU Split Calculator — Runyard tool

Multi-GPU Split Calculator

Live

Cards required at each quantisation, and what that costs.

Hardware3 inputs
Open details ->
Ollama Storage Planner — Runyard tool

Ollama Storage Planner

Live

Disk for a library of models, before you run out mid-pull.

Runtime & errors4 inputs
Open details ->
GPU Upgrade Advisor — Runyard tool

GPU Upgrade Advisor

Live

What a bigger card actually unlocks, in models and in speed.

Hardware3 inputs
Open details ->
Local Cost per Million Tokens — Runyard tool

Local Cost per Million Tokens

Live

Electricity plus hardware amortisation, expressed the way APIs price.

Cost6 inputs
Open details ->
GPU Payback Period Calculator — Runyard tool

GPU Payback Period Calculator

Live

Break-even in months, from your real token volume.

Cost6 inputs
Open details ->
LLM Electricity Cost Calculator — Runyard tool

LLM Electricity Cost Calculator

Live

Monthly and yearly electricity, from power draw and hours.

Cost4 inputs
Open details ->
Cloud GPU vs Buy Calculator — Runyard tool

Cloud GPU vs Buy Calculator

Live

The hours per month at which renting stops being the cheaper option.

Cost5 inputs
Open details ->
API Token Cost Calculator — Runyard tool

API Token Cost Calculator

Live

Monthly spend from requests, prompt length and price — with caching.

Cost6 inputs
Open details ->
RAG vs Fine-Tune Cost — Runyard tool

RAG vs Fine-Tune Cost

Live

A one-off training cost against the per-request cost of retrieval.

Cost5 inputs
Open details ->
GPU Price per GB of VRAM — Runyard tool

GPU Price per GB of VRAM

Live

Every catalogued card ranked by cost per usable gigabyte.

Hardware2 inputs
Open details ->
Cheapest GPU for a Model — Runyard tool

Cheapest GPU for a Model

Live

The lowest-cost card that fits, at every quantisation level.

Hardware2 inputs
Open details ->
Mac Unified Memory Calculator — Runyard tool

Mac Unified Memory Calculator

Live

What each memory configuration actually runs, and how fast.

Hardware3 inputs
Open details ->
RAM Offload Speed Calculator — Runyard tool

RAM Offload Speed Calculator

Live

The real cost of running a model that does not quite fit.

Speed5 inputs
Open details ->
CPU-Only Inference Calculator — Runyard tool

CPU-Only Inference Calculator

Live

What CPU-only inference realistically gives you, at each model size.

Speed4 inputs
Open details ->
Laptop vs Desktop for Local AI — Runyard tool

Laptop vs Desktop for Local AI

Live

Capacity, speed and cost across both, for the model you want to run.

Hardware3 inputs
Open details ->
Used GPU Value Checker — Runyard tool

Used GPU Value Checker

Live

Price per usable gigabyte against every alternative at the same capability.

Hardware3 inputs
Open details ->
KV Cache Calculator — Runyard tool

KV Cache Calculator

Live

The part of your VRAM that grows with every token in the conversation.

Context3 inputs
Open details ->
Flash Attention Savings Calculator — Runyard tool

Flash Attention Savings Calculator

Live

What --flash-attn buys you in context length, on your card.

Context3 inputs
Open details ->
RAG Chunk Size Calculator — Runyard tool

RAG Chunk Size Calculator

Live

How many chunks fit in your context once the prompt takes its share.

Context5 inputs
Open details ->
Embedding Storage Calculator — Runyard tool

Embedding Storage Calculator

Live

Disk and memory for a vector index, from document count and dimension.

Context5 inputs
Open details ->
Vector Database Memory Calculator — Runyard tool

Vector Database Memory Calculator

Live

Whether the index fits in memory, and what happens when it does not.

Context5 inputs
Open details ->
Codebase Context Planner — Runyard tool

Codebase Context Planner

Live

Lines of code to tokens, against real context windows.

Context4 inputs
Open details ->
MoE vs Dense Memory Calculator — Runyard tool

MoE vs Dense Memory Calculator

Live

Capacity follows total parameters; speed follows the active ones.

Memory5 inputs
Open details ->
Model Download Time Calculator — Runyard tool

Model Download Time Calculator

Live

File size at each quantisation, and the wait on your connection.

Runtime & errors3 inputs
Open details ->
Speculative Decoding Speedup — Runyard tool

Speculative Decoding Speedup

Live

Expected speedup from a draft model, at your acceptance rate.

Speed5 inputs
Open details ->
GPU vs GPU for Local AI — Runyard tool

GPU vs GPU for Local AI

Live

Two cards compared on what actually matters: capacity and bandwidth.

Hardware3 inputs
Open details ->
Whisper VRAM Calculator — Runyard tool

Whisper VRAM Calculator

Live

Memory and real-time factor for each Whisper size, on your card.

Memory4 inputs
Open details ->
Team GPU Sizing Calculator — Runyard tool

Team GPU Sizing Calculator

Live

Cards required to serve N people, from concurrency and daily volume.

Runtime & errors6 inputs
Open details ->
Agent Loop Cost Calculator — Runyard tool

Agent Loop Cost Calculator

Live

Why a coding agent bills far more than the tokens you think you sent.

Cost6 inputs
Open details ->

In the workshop

Ideas on the list, not yet built. They appear here without links until they compute something real.

Soon

Tokens & Context Calculator

T-02

Estimate tokens for any prompt, check context window fit, and compare model limits.

Coming soon
Soon

Model Comparison Matrix

T-03

Side-by-side specs, performance, and cost for any two models you want to compare.

Coming soon
Soon

Inference Speed Estimator

T-04

Predict tokens/sec for your hardware given a model size and quantization.

Coming soon
Soon

Quantization Picker

T-05

Answer 3 questions about your GPU and use case and get the exact quant level to use.

Coming soon
Soon

Cloud GPU Cost Estimator

T-06

Compare RunPod, Lambda, Vast.ai and more and find the cheapest option for your workload.

Coming soon
Product-led tools

Need the actual hardware-fit answer?

GPU-to-Model Fit Checker, Can My PC Run This AI Model, and Best Model for My Hardware Finder all funnel into Model Radar because that is already the real interactive product.

Open Model Radar ->