Files
bdeb1337 bb71b49fde feat: localchat, a Go + Templ + HTMX chat app for local Gemma models
Streams replies over SSE from any OpenAI-compatible server (oMLX,
llama.cpp, Ollama, ...). Single binary that can install itself as an
OS service; docker compose bundles llama.cpp + Gemma 4 E2B.
2026-10-05 07:25:42 +02:00

128 lines
6.2 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# localchat
localchat gives you a browser chat UI for an open-weight AI model running on your own machine. It's a **Go + Templ + HTMX** app that talks to any OpenAI-compatible model server, and it defaults to **Gemma 4 E2B**, which runs on a normal laptop.
- 🔒 Your prompts and replies stay on your machine, and the app works offline once you have the model.
- 💸 You need no API key and pay nothing per token.
- 🔁 You can point it at oMLX, llama.cpp, Ollama, LM Studio, Lemonade or vLLM by changing one environment variable.
- 📦 You get a ~9 MB binary for Linux, macOS or Windows that can install itself as a background service. `docker compose up` starts the app, llama.cpp and Gemma together.
![screenshot](docs/screenshot.png)
## Quick start
### Option A: containers
```sh
docker compose up -d # first run downloads Gemma 4 E2B (~3 GB)
open http://localhost:3000
```
To use a different model, set `LLM_HF_REPO`:
```sh
LLM_HF_REPO=unsloth/gemma-4-E4B-it-GGUF:Q4_K_M docker compose up -d
```
### Option B: binary + your own model server
```sh
make build # or grab a binary from `make dist`
cp .env.example .env # point it at your model server
./bin/localchat # → http://127.0.0.1:3000
```
Set `LLM_BASE_URL` to match your server:
| Server | `LLM_BASE_URL` |
|---|---|
| llama.cpp `llama-server` | `http://127.0.0.1:8080/v1` |
| oMLX (Apple Silicon) | `http://127.0.0.1:8000/v1` |
| Ollama | `http://127.0.0.1:11434/v1` |
| LM Studio | `http://127.0.0.1:1234/v1` |
| Lemonade (AMD) | `http://127.0.0.1:13305/api/v1` |
If you don't run a model server yet, llama.cpp gets you one in two commands:
```sh
brew install llama.cpp # or a release from github.com/ggml-org/llama.cpp
llama-server -hf unsloth/gemma-4-E2B-it-GGUF:Q4_K_M --no-mmproj
```
### Option C: background service
You can register localchat with your OS service manager: systemd on Linux, launchd on macOS, or the Service Control Manager on Windows.
```sh
./bin/localchat install --user # per-user service; drop --user for a system-wide one (needs sudo/admin)
./bin/localchat start --user
./bin/localchat status --user
./bin/localchat stop --user && ./bin/localchat uninstall --user
```
`install` writes the absolute path of your `.env` (or `--env-file`) into the service definition, so the service runs with the settings you tested.
## Configuration
Set these as environment variables or in a `.env` file. If you set both, the environment variable wins.
| Variable | Default | |
|---|---|---|
| `LLM_BASE_URL` | `http://127.0.0.1:8000/v1` | OpenAI-compatible API root |
| `LLM_MODEL` | *(first model the server lists)* | model id |
| `LLM_API_KEY` | | bearer token, if your server needs one |
| `SYSTEM_PROMPT` | friendly, concise assistant | give it a personality or a purpose |
| `LOCALCHAT_TITLE` | `localchat` | name in the header and tab |
| `LOCALCHAT_ADDR` | `127.0.0.1:3000` | listen address (`0.0.0.0:3000` to share on your LAN) |
| `MAX_HISTORY` | `20` | past messages sent as context |
| `LLM_TIMEOUT` | `5m` | time limit for one reply |
## Architecture
```
Browser ── HTMX + SSE extension
│ POST /chat → returns the user bubble + an empty reply bubble
│ GET /chat/stream/{id} ← server-sent events: "token" (append) … "done" (swap in Markdown)
Go (net/http + Templ)
│ POST /v1/chat/completions {stream: true}
Model server (llama.cpp / oMLX / Ollama / …) ── Gemma 4 E2B
```
1. Your browser posts the form with `hx-post`. The server stores the message and returns two Templ fragments: your message and an assistant bubble with `sse-connect="/chat/stream/{id}"`.
2. The SSE handler claims the reply, sends the conversation to the model and forwards each chunk as an HTML-escaped `token` event. HTMX appends the chunks with `hx-swap="beforeend"`, so you see the reply appear as the model writes it.
3. Once the model finishes, the handler sends a `done` event. HTMX swaps the bubble for the reply rendered as Markdown (goldmark, raw HTML stripped), and removing the `sse-connect` element closes the stream.
4. Browsers reconnect a dropped EventSource on their own. If a browser reconnects for a reply that another request already claimed or finished, the handler returns the final state and leaves the model alone.
You won't find a JavaScript framework or a frontend build step. The JavaScript comes down to htmx, its SSE extension and about 40 lines of UX code, and localchat embeds all of it in the binary so the app runs offline.
## Project layout
```
cmd/localchat/ entry point, CLI + service install/start/stop
internal/config/ env + .env loading
internal/llm/ OpenAI-compatible streaming client (standard library)
internal/chat/ in-memory conversations, one per browser session
internal/web/ routes, SSE streaming, embedded static assets
internal/web/views/ Templ components (*.templ → generated *_templ.go)
scripts/e2e.py Playwright browser test against a real model
compose.yaml, Dockerfile
```
## Development
```sh
make test # unit + handler tests with a fake model server, -race
make run # build + run with ./.env
make e2e # headless browser test against the running app + real model
make dist # cross-compile linux/darwin/windows × amd64/arm64
```
After you edit a `.templ` file, run `make generate` (or `go tool templ generate --watch`). The repo includes the generated `*_templ.go` files, so you can run `go build` or a Docker build without Templ installed.
## Good to know
- localchat keeps conversations in memory, so a restart clears them. That keeps chats private; add SQLite if you want history to survive a restart.
- localchat has no authentication and listens on localhost by default. Put a reverse proxy with auth in front before you expose it.
- In Docker, speed depends on how many CPUs you give the Docker VM. With Docker Desktop set to 1 CPU, Gemma E2B needs about 20 s before it starts replying. On a Mac, a native model server (oMLX or `llama-server`, both using the GPU through Metal) sends the first token in about a second on an M3.
- The compose file runs the model on the CPU, which works on any platform and suits E2B. On Linux with an NVIDIA or AMD GPU, switch to the `server-cuda` or `server-rocm` llama.cpp image.