feat: localchat, a Go + Templ + HTMX chat app for local Gemma models
Streams replies over SSE from any OpenAI-compatible server (oMLX, llama.cpp, Ollama, ...). Single binary that can install itself as an OS service; docker compose bundles llama.cpp + Gemma 4 E2B.
This commit is contained in:
30 files changed
+2566
No files matched your search
@@ -0,0 +1,125 @@
|
||||
# localchat
|
||||
|
||||
A small, private chat app for a **local, open-weight AI model**: **Go + Templ + HTMX** on top of any OpenAI-compatible model server. It defaults to **Gemma 4 E2B**, which runs on a normal laptop.
|
||||
|
||||
- 🔒 **Private:** prompts and replies never leave your machine.
|
||||
- ✈️ **Works offline** once the model is downloaded.
|
||||
- 💸 **Free to run:** no API keys, no per-token billing.
|
||||
- 🔁 **Model-agnostic:** oMLX, llama.cpp, Ollama, LM Studio, Lemonade, vLLM… change one env var.
|
||||
- 📦 **One ~9 MB binary** for Linux, macOS and Windows. It can install itself as a background service.
|
||||
- 🐳 **Or `docker compose up`** for the app plus llama.cpp plus Gemma in one go.
|
||||
|
||||

|
||||
|
||||
## Quick start
|
||||
|
||||
### Option A: containers (easiest)
|
||||
|
||||
```sh
|
||||
docker compose up -d # first run downloads Gemma 4 E2B (~3 GB)
|
||||
open http://localhost:3000
|
||||
```
|
||||
|
||||
To use a different model, set `LLM_HF_REPO=unsloth/gemma-4-E4B-it-GGUF:Q4_K_M docker compose up -d`.
|
||||
|
||||
### Option B: binary + a model server you already run
|
||||
|
||||
```sh
|
||||
make build # or grab a binary from `make dist`
|
||||
cp .env.example .env # point it at your model server
|
||||
./bin/localchat # → http://127.0.0.1:3000
|
||||
```
|
||||
|
||||
Common `LLM_BASE_URL` values:
|
||||
|
||||
| Server | `LLM_BASE_URL` |
|
||||
|---|---|
|
||||
| llama.cpp `llama-server` | `http://127.0.0.1:8080/v1` |
|
||||
| oMLX (Apple Silicon) | `http://127.0.0.1:8000/v1` |
|
||||
| Ollama | `http://127.0.0.1:11434/v1` |
|
||||
| LM Studio | `http://127.0.0.1:1234/v1` |
|
||||
| Lemonade (AMD) | `http://127.0.0.1:13305/api/v1` |
|
||||
|
||||
For example, to get a model running quickly with llama.cpp:
|
||||
|
||||
```sh
|
||||
brew install llama.cpp # or a release from github.com/ggml-org/llama.cpp
|
||||
llama-server -hf unsloth/gemma-4-E2B-it-GGUF:Q4_K_M --no-mmproj
|
||||
```
|
||||
|
||||
### Option C: run it as a background service
|
||||
|
||||
The binary registers itself with the OS service manager (systemd, launchd or the Windows Service Control Manager):
|
||||
|
||||
```sh
|
||||
./bin/localchat install --user # per-user (launchd agent / systemd --user); drop --user for system-wide (needs sudo/admin)
|
||||
./bin/localchat start --user
|
||||
./bin/localchat status --user
|
||||
./bin/localchat stop --user && ./bin/localchat uninstall --user
|
||||
```
|
||||
|
||||
The service gets the absolute path of your `.env` (or `--env-file`), so it runs with the same settings you tested with.
|
||||
|
||||
## Configuration
|
||||
|
||||
Set these as environment variables or in a `.env` file. Real environment variables take precedence over the file.
|
||||
|
||||
| Variable | Default | |
|
||||
|---|---|---|
|
||||
| `LLM_BASE_URL` | `http://127.0.0.1:8000/v1` | OpenAI-compatible API root |
|
||||
| `LLM_MODEL` | *(first model the server lists)* | model id |
|
||||
| `LLM_API_KEY` | | bearer token, if your server needs one |
|
||||
| `SYSTEM_PROMPT` | friendly, concise assistant | give it a personality or a purpose |
|
||||
| `LOCALCHAT_TITLE` | `localchat` | name in the header and tab |
|
||||
| `LOCALCHAT_ADDR` | `127.0.0.1:3000` | listen address (`0.0.0.0:3000` to share on your LAN) |
|
||||
| `MAX_HISTORY` | `20` | past messages sent as context |
|
||||
| `LLM_TIMEOUT` | `5m` | maximum time for one reply |
|
||||
|
||||
## How it works
|
||||
|
||||
```
|
||||
Browser ── HTMX + SSE extension
|
||||
│ POST /chat → returns the user bubble + an empty reply bubble
|
||||
│ GET /chat/stream/{id} ← server-sent events: "token" (append) … "done" (swap in Markdown)
|
||||
Go (net/http + Templ)
|
||||
│ POST /v1/chat/completions {stream: true}
|
||||
Model server (llama.cpp / oMLX / Ollama / …) ── Gemma 4 E2B
|
||||
```
|
||||
|
||||
1. The form posts with `hx-post`. The server stores the message and returns two Templ fragments: your message, and an assistant bubble with `sse-connect="/chat/stream/{id}"`.
|
||||
2. The SSE handler claims that reply, sends the conversation to the model, and forwards each chunk as an HTML-escaped `token` event. HTMX appends each one (`hx-swap="beforeend"`), so the reply types out live.
|
||||
3. When the model finishes, a `done` event replaces the whole bubble with the reply rendered as Markdown (goldmark, with raw HTML stripped). That also removes the `sse-connect` element, which closes the stream.
|
||||
4. Browsers automatically reconnect an EventSource. A reconnect for a reply that is already claimed or finished gets the final state instead of a second generation.
|
||||
|
||||
There is no JavaScript framework and no build step. The only JS is htmx, its SSE extension and about 40 lines of UX glue. All of it is embedded in the binary, so the app works fully offline.
|
||||
|
||||
## Project layout
|
||||
|
||||
```
|
||||
cmd/localchat/ entry point, CLI + service install/start/stop
|
||||
internal/config/ env + .env loading
|
||||
internal/llm/ tiny OpenAI-compatible streaming client (stdlib only)
|
||||
internal/chat/ in-memory conversations, one per browser session
|
||||
internal/web/ routes, SSE streaming, embedded static assets
|
||||
internal/web/views/ Templ components (*.templ → generated *_templ.go)
|
||||
scripts/e2e.py Playwright browser test against a real model
|
||||
compose.yaml, Dockerfile
|
||||
```
|
||||
|
||||
## Development
|
||||
|
||||
```sh
|
||||
make test # unit + handler tests with a fake model server, -race
|
||||
make run # build + run with ./.env
|
||||
make e2e # headless browser test against the running app + real model
|
||||
make dist # cross-compile linux/darwin/windows × amd64/arm64
|
||||
```
|
||||
|
||||
After editing a `.templ` file, run `make generate` (or `go tool templ generate --watch`). The generated `*_templ.go` files are committed, so `go build` and Docker builds don't need Templ installed.
|
||||
|
||||
## Good to know
|
||||
|
||||
- Conversations live in memory and are lost on restart. That's on purpose for privacy, and simple enough to swap for SQLite.
|
||||
- There is no authentication. It listens on localhost by default. Put a reverse proxy with auth in front before exposing it.
|
||||
- Speed in Docker depends on how many CPUs the Docker VM gets. With Docker Desktop's 1-CPU setting, Gemma E2B takes about 20 s to start replying, so give it more cores. On a Mac, running the model natively (oMLX or `llama-server`, which use the GPU via Metal) is far faster: about 1.5 s to the first token on an M3.
|
||||
- GPU in Docker depends on the platform. The compose file runs on CPU everywhere, which E2B handles fine. On Linux with an NVIDIA or AMD GPU, use the `server-cuda` or `server-rocm` llama.cpp image tags.
|
||||
Reference in new issue
Block a user