Apps that run models on your own computer: supported systems, memory needs, GPU support and a download from our catalogue.
Updated
Ollama
Open source
Runs open-weight models on your own computer from the command line or a small desktop app, and exposes them through a local, OpenAI-compatible API. Open source under MIT.
Systems
Windows, macOS, Linux
GPU
Optional: CUDA, ROCm, Vulkan, Metal
OpenAI-compatible API
yes
License
MIT
GitHub stars
181,947
LM Studio
Free
Desktop app for finding, downloading and chatting with local models, with a built-in API server for developers. Free to use, though not open source.
Systems
Windows, macOS, Linux
Min. RAM
16 GB
GPU
Optional: 4 GB+ VRAM, Apple Silicon
OpenAI-compatible API
yes
License
Proprietary
Jan
Open source
Open-source chat app in the ChatGPT style that runs models offline and can switch to cloud APIs in the same window. It also starts a local OpenAI-compatible server.
Systems
Windows, macOS, Linux
Min. RAM
8 GB
GPU
Optional: NVIDIA, AMD, Intel Arc, Metal
OpenAI-compatible API
yes
License
Apache-2.0
GitHub stars
44,717
GPT4All
Open source
Nomic's open-source desktop app for local models on CPU or GPU, known for LocalDocs, which answers questions about your own files. New releases have been rare since February 2025.
Systems
Windows, macOS, Linux
Min. RAM
8 GB
GPU
Optional: Vulkan, Apple M-series
OpenAI-compatible API
yes
License
MIT
GitHub stars
77,386
AnythingLLM
Open source
All-in-one open-source workspace, on the desktop or in Docker, where documents, agents and local or cloud models come together. Suited to private work with your own files.
Systems
Windows, macOS, Linux
Min. RAM
16 GB
GPU
Optional: CUDA, ROCm, Vulkan, NPU
License
MIT
GitHub stars
66,618
Open WebUI
Open source
Self-hosted web interface for Ollama and OpenAI-compatible APIs, installed with Docker or pip. Several people can share one server through their browsers.
Systems
Windows, macOS, Linux
GPU
Not needed: uses Ollama or an API
OpenAI-compatible API
yes
License
Open WebUI License
GitHub stars
153,609
llama.cpp
Open source
Open-source C/C++ engine that runs GGUF models on CPU and GPU, with command-line tools and an OpenAI-compatible server. Many local AI apps are built on top of it.
Systems
Windows, macOS, Linux
GPU
Optional: CUDA, ROCm, Vulkan, SYCL, Metal
OpenAI-compatible API
yes
License
MIT
GitHub stars
129,950
KoboldCpp
Open source
Single-file program based on llama.cpp, with a web interface for chat and story writing and KoboldAI- and OpenAI-compatible APIs, favoured for fiction and role-play.
Systems
Windows, macOS, Linux
GPU
Optional: CUDA, Vulkan, hipBLAS, Metal
OpenAI-compatible API
yes
License
AGPL-3.0
GitHub stars
11,908
Running models on your own computer
A local LLM app downloads an open-weight model and runs it on your PC or laptop. Your prompts stay on the machine, there is no per-token bill, and it works offline once the model is downloaded. This page lists such apps: desktop programs (Ollama, LM Studio, Jan, GPT4All, AnythingLLM), a web interface for a home server (Open WebUI) and engines for people who want full control (llama.cpp, KoboldCpp).
What the cards show
supported systems, the minimum RAM and GPU support;
the licence and, for open-source projects, GitHub stars, refreshed weekly;
a Download button when the app is in our catalogue — those installers were checked for a valid digital signature.
Choosing an app
Want a chat window and nothing else? LM Studio and Jan have a built-in model browser.
Want a local API for your own tools? Ollama and llama.cpp serve one, and many other apps connect to them.
Want to ask questions about your documents? AnythingLLM and GPT4All are built around that.
Serving several people from one machine? Open WebUI runs in the browser on top of Ollama.
Memory decides what you can run
The app is free; the model is a separate download with its own size and licence. A model has to fit into RAM, or into the GPU's memory if you want the graphics card to do the work, so a bigger model needs a bigger machine. Without a suitable GPU, models run on the processor, which works but is slower. The cards list what the vendors state; we do not promise any particular speed or answer quality, because both depend on the model and the hardware.
When local is not enough
The strongest models are not published as open weights and run only in the cloud. If you need them, look at the API providers or the AI apps; several local apps, Jan and AnythingLLM among them, can also connect to a cloud API alongside local models.