Skip to content
fedi.software

Local LLM apps

Apps that run models on your own computer: supported systems, memory needs, GPU support and a download from our catalogue.

Updated

Ollama

Open source

Runs open-weight models on your own computer from the command line or a small desktop app, and exposes them through a local, OpenAI-compatible API. Open source under MIT.

Systems
Windows, macOS, Linux
GPU
Optional: CUDA, ROCm, Vulkan, Metal
OpenAI-compatible API
yes
License
MIT
GitHub stars
181,947

LM Studio

Free

Desktop app for finding, downloading and chatting with local models, with a built-in API server for developers. Free to use, though not open source.

Systems
Windows, macOS, Linux
Min. RAM
16 GB
GPU
Optional: 4 GB+ VRAM, Apple Silicon
OpenAI-compatible API
yes
License
Proprietary

Jan

Open source

Open-source chat app in the ChatGPT style that runs models offline and can switch to cloud APIs in the same window. It also starts a local OpenAI-compatible server.

Systems
Windows, macOS, Linux
Min. RAM
8 GB
GPU
Optional: NVIDIA, AMD, Intel Arc, Metal
OpenAI-compatible API
yes
License
Apache-2.0
GitHub stars
44,717

GPT4All

Open source

Nomic's open-source desktop app for local models on CPU or GPU, known for LocalDocs, which answers questions about your own files. New releases have been rare since February 2025.

Systems
Windows, macOS, Linux
Min. RAM
8 GB
GPU
Optional: Vulkan, Apple M-series
OpenAI-compatible API
yes
License
MIT
GitHub stars
77,386

AnythingLLM

Open source

All-in-one open-source workspace, on the desktop or in Docker, where documents, agents and local or cloud models come together. Suited to private work with your own files.

Systems
Windows, macOS, Linux
Min. RAM
16 GB
GPU
Optional: CUDA, ROCm, Vulkan, NPU
License
MIT
GitHub stars
66,618

Open WebUI

Open source

Self-hosted web interface for Ollama and OpenAI-compatible APIs, installed with Docker or pip. Several people can share one server through their browsers.

Systems
Windows, macOS, Linux
GPU
Not needed: uses Ollama or an API
OpenAI-compatible API
yes
License
Open WebUI License
GitHub stars
153,609

llama.cpp

Open source

Open-source C/C++ engine that runs GGUF models on CPU and GPU, with command-line tools and an OpenAI-compatible server. Many local AI apps are built on top of it.

Systems
Windows, macOS, Linux
GPU
Optional: CUDA, ROCm, Vulkan, SYCL, Metal
OpenAI-compatible API
yes
License
MIT
GitHub stars
129,950

KoboldCpp

Open source

Single-file program based on llama.cpp, with a web interface for chat and story writing and KoboldAI- and OpenAI-compatible APIs, favoured for fiction and role-play.

Systems
Windows, macOS, Linux
GPU
Optional: CUDA, Vulkan, hipBLAS, Metal
OpenAI-compatible API
yes
License
AGPL-3.0
GitHub stars
11,908

Running models on your own computer

A local LLM app downloads an open-weight model and runs it on your PC or laptop. Your prompts stay on the machine, there is no per-token bill, and it works offline once the model is downloaded. This page lists such apps: desktop programs (Ollama, LM Studio, Jan, GPT4All, AnythingLLM), a web interface for a home server (Open WebUI) and engines for people who want full control (llama.cpp, KoboldCpp).

What the cards show

  • supported systems, the minimum RAM and GPU support;
  • the licence and, for open-source projects, GitHub stars, refreshed weekly;
  • a Download button when the app is in our catalogue — those installers were checked for a valid digital signature.

Choosing an app

  1. Want a chat window and nothing else? LM Studio and Jan have a built-in model browser.
  2. Want a local API for your own tools? Ollama and llama.cpp serve one, and many other apps connect to them.
  3. Want to ask questions about your documents? AnythingLLM and GPT4All are built around that.
  4. Serving several people from one machine? Open WebUI runs in the browser on top of Ollama.

Memory decides what you can run

The app is free; the model is a separate download with its own size and licence. A model has to fit into RAM, or into the GPU's memory if you want the graphics card to do the work, so a bigger model needs a bigger machine. Without a suitable GPU, models run on the processor, which works but is slower. The cards list what the vendors state; we do not promise any particular speed or answer quality, because both depend on the model and the hardware.

When local is not enough

The strongest models are not published as open weights and run only in the cloud. If you need them, look at the API providers or the AI apps; several local apps, Jan and AnythingLLM among them, can also connect to a cloud API alongside local models.