Streams replies over SSE from any OpenAI-compatible server (oMLX, llama.cpp, Ollama, ...). Single binary that can install itself as an OS service; docker compose bundles llama.cpp + Gemma 4 E2B.
6.6 KiB
title, published, tags
| title | published | tags |
|---|---|---|
| localchat: a bare-basic Go + HTMX web app with a local Gemma model, built for a friend | false | hf26challenge, go, htmx, gemma |
This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend
What I Built
localchat is a small chat app for a local AI model: a Go web server, a browser UI and a small open-weight model running on your own machine. You type, and Google's Gemma 4 E2B answers token by token. No account, no API key, and nothing leaves the laptop.
Who it's for: a friend of mine wanted to see what a bare-basic Go web app with AI integration actually looks like. No frameworks hiding things, no SDK magic, just the essential pieces you can read in one sitting. I'd never built an AI app in Go either, so we treated this as a chance to get the hang of it together and make it really work.
I picked Gemma because I love that small models now run locally and are actually useful. A 2B-class model on a laptop gives a fast, private, free assistant, and it's exactly the kind of thing you can hand to a friend without asking them to sign up anywhere.
The problem it solves for my friend:
- One small, readable project that shows the whole path: an HTML form → Go handler → model → streamed reply in the browser.
- Runs on whatever they have. One ~9 MB binary for Linux, macOS or Windows, or
docker compose upwith the model included. - A starting point, not a demo. They can change the system prompt, swap the model, or install it as a background service, and build their own idea on top.
Demo
Run it yourself in one command (the first start downloads Gemma 4 E2B, ~3 GB):
docker compose up -d
open http://localhost:3000
On my M3 MacBook (24 GB) with Gemma 4 E2B (4-bit, via oMLX), the first token arrives in about a second. A short answer with a code block is done in 3–5 s.
Code
Repo: git.b0b.be/bdeb/localchat
cmd/localchat/ entry point: serve + install/start/stop as an OS service
internal/llm/ ~150-line OpenAI-compatible streaming client (stdlib only)
internal/chat/ in-memory conversations, one per browser session
internal/web/ routes, SSE streaming, embedded static assets
internal/web/views/ Templ components
scripts/e2e.py Playwright browser test against the real model
compose.yaml app + llama.cpp + Gemma 4 E2B
How I Built It
Open-source AI used:
- Model: Google Gemma 4 E2B (open weights), 4-bit quantized.
- Local inference: oMLX (Apple MLX) on my Mac during development, and llama.cpp (
llama-server, GGUF) in the container. My AMD GPU box runs Lemonade, which speaks the same API. - App stack: Go 1.27 · Templ · HTMX 2 + its SSE extension · goldmark · kardianos/service.
How it fits together:
Browser ── HTMX + SSE extension
│ POST /chat → user bubble + empty reply bubble (sse-connect)
│ GET /chat/stream/{id} ← "token" events (append) … "done" (swap in Markdown)
Go (net/http + Templ)
│ POST /v1/chat/completions {stream: true}
Local model server ── Gemma 4 E2B
The core idea is that streaming is just HTML over server-sent events:
- Sending a message returns two Templ fragments: your bubble, and an empty assistant bubble with
sse-connect="/chat/stream/{id}". - The Go handler streams the reply from the model and forwards every chunk as an HTML-escaped
<span>in atokenevent. HTMX appends them (hx-swap="beforeend"), so the answer types itself out with no custom JavaScript. - When the model finishes, a
doneevent replaces the bubble with server-rendered Markdown (goldmark, raw HTML stripped). Swapping out thesse-connectelement also closes the stream. - Browsers automatically reconnect an EventSource, so a reconnect for a finished reply gets the final HTML instead of a second run of the model. A test checks that the model is only called once.
Other parts I enjoyed:
- No AI SDK. OpenAI-compatible streaming is just a
bufio.Scannerreadingdata:lines, which is the "bare-basic" my friend asked for. - It runs as a real service.
localchat install --userregisters a launchd agent, a systemd unit or a Windows service that points at your.env. - Everything is embedded (htmx, the SSE extension, CSS, icon), so it works with the network cable pulled.
- Tested: handler tests against a fake OpenAI-style server (streaming, history, session isolation, XSS-safe Markdown, model offline), plus a Playwright browser test against the real Gemma model, both natively and in Docker.
Why Does Open Innovation Matter?
- Private by design. Gemma's weights sit on the laptop, so there's no third-party server to trust. You can paste work code or private notes, and conversations aren't even written to disk.
- Free to run, forever. No tokens, no subscription, no free tier that disappears. My friend can leave it running without thinking about a bill.
- Works offline. On a train, on a plane, or behind a strict corporate network.
- Swap anything. The app only speaks the OpenAI-compatible protocol that the open ecosystem standardized on. I developed against oMLX and shipped with llama.cpp: two different inference engines, the same binary, no code changes. Moving to Gemma 4 E4B or another open model is one environment variable. A closed API would have tied the app to one vendor, one price list and one data policy.
- Small open models are good enough now. Gemma 4 E2B at 4-bit fits in a few GB of RAM and writes code snippets and explanations quickly on a laptop with no GPU. That makes a "bare-basic" AI app something you can just hand to a friend.
My Agent Session
I built this with Claude Code as a pair-programmer: brainstorming ideas, scaffolding the Go/Templ/HTMX app, and testing it end to end against the local Gemma model.
Prize Categories
- Best Use of Gemma: the whole app is built around running Google's Gemma 4 E2B locally (MLX on Mac, GGUF via llama.cpp in Docker).

