Streams replies over SSE from any OpenAI-compatible server (oMLX, llama.cpp, Ollama, ...). Single binary that can install itself as an OS service; docker compose bundles llama.cpp + Gemma 4 E2B.
112 lines
6.6 KiB
Markdown
112 lines
6.6 KiB
Markdown
---
|
||
title: "localchat: a bare-basic Go + HTMX web app with a local Gemma model, built for a friend"
|
||
published: false
|
||
tags: hf26challenge, go, htmx, gemma
|
||
---
|
||
|
||
<!--
|
||
Before publishing:
|
||
- Upload docs/screenshot.png and docs/screenshot-dark.png to DEV and swap in the uploaded image URLs.
|
||
- Make https://git.b0b.be/bdeb/localchat public so judges can see the code.
|
||
- Optional: save the agent session with DevRelay and paste the embed into "My Agent Session".
|
||
- Deadline: Monday Oct 5, 08:59 Brussels time (06:59 UTC).
|
||
-->
|
||
|
||
*This is a submission for the [Hacktoberfest Weekend Challenge: Build for a Friend](https://dev.to/challenges/hacktoberfest-weekend-2026-10-01)*
|
||
|
||
## What I Built
|
||
|
||
**localchat** is a small chat app for a local AI model: a Go web server, a browser UI and a small open-weight model running on your own machine. You type, and **Google's Gemma 4 E2B** answers token by token. No account, no API key, and nothing leaves the laptop.
|
||
|
||
**Who it's for:** a friend of mine wanted to see what a *bare-basic* Go web app with AI integration actually looks like. No frameworks hiding things, no SDK magic, just the essential pieces you can read in one sitting. I'd never built an AI app in Go either, so we treated this as a chance to get the hang of it together and make it really work.
|
||
|
||
I picked Gemma because I love that small models now run *locally* and are actually useful. A 2B-class model on a laptop gives a fast, private, free assistant, and it's exactly the kind of thing you can hand to a friend without asking them to sign up anywhere.
|
||
|
||
The problem it solves for my friend:
|
||
|
||
- **One small, readable project** that shows the whole path: an HTML form → Go handler → model → streamed reply in the browser.
|
||
- **Runs on whatever they have.** One ~9 MB binary for Linux, macOS or Windows, or `docker compose up` with the model included.
|
||
- **A starting point, not a demo.** They can change the system prompt, swap the model, or install it as a background service, and build their own idea on top.
|
||
|
||
## Demo
|
||
|
||

|
||
|
||

|
||
|
||
Run it yourself in one command (the first start downloads Gemma 4 E2B, ~3 GB):
|
||
|
||
```sh
|
||
docker compose up -d
|
||
open http://localhost:3000
|
||
```
|
||
|
||
On my M3 MacBook (24 GB) with Gemma 4 E2B (4-bit, via oMLX), the first token arrives in **about a second**. A short answer with a code block is done in **3–5 s**.
|
||
|
||
## Code
|
||
|
||
**Repo:** [git.b0b.be/bdeb/localchat](https://git.b0b.be/bdeb/localchat)
|
||
|
||
```
|
||
cmd/localchat/ entry point: serve + install/start/stop as an OS service
|
||
internal/llm/ ~150-line OpenAI-compatible streaming client (stdlib only)
|
||
internal/chat/ in-memory conversations, one per browser session
|
||
internal/web/ routes, SSE streaming, embedded static assets
|
||
internal/web/views/ Templ components
|
||
scripts/e2e.py Playwright browser test against the real model
|
||
compose.yaml app + llama.cpp + Gemma 4 E2B
|
||
```
|
||
|
||
## How I Built It
|
||
|
||
**Open-source AI used:**
|
||
|
||
- **Model:** Google **Gemma 4 E2B** (open weights), 4-bit quantized.
|
||
- **Local inference:** **oMLX** (Apple MLX) on my Mac during development, and **llama.cpp** (`llama-server`, GGUF) in the container. My AMD GPU box runs **Lemonade**, which speaks the same API.
|
||
- **App stack:** Go 1.27 · Templ · HTMX 2 + its SSE extension · goldmark · kardianos/service.
|
||
|
||
**How it fits together:**
|
||
|
||
```
|
||
Browser ── HTMX + SSE extension
|
||
│ POST /chat → user bubble + empty reply bubble (sse-connect)
|
||
│ GET /chat/stream/{id} ← "token" events (append) … "done" (swap in Markdown)
|
||
Go (net/http + Templ)
|
||
│ POST /v1/chat/completions {stream: true}
|
||
Local model server ── Gemma 4 E2B
|
||
```
|
||
|
||
The core idea is that **streaming is just HTML over server-sent events**:
|
||
|
||
1. Sending a message returns two Templ fragments: your bubble, and an empty assistant bubble with `sse-connect="/chat/stream/{id}"`.
|
||
2. The Go handler streams the reply from the model and forwards every chunk as an HTML-escaped `<span>` in a `token` event. HTMX appends them (`hx-swap="beforeend"`), so the answer types itself out with no custom JavaScript.
|
||
3. When the model finishes, a `done` event replaces the bubble with server-rendered Markdown (goldmark, raw HTML stripped). Swapping out the `sse-connect` element also closes the stream.
|
||
4. Browsers automatically reconnect an EventSource, so a reconnect for a finished reply gets the final HTML instead of a second run of the model. A test checks that the model is only called once.
|
||
|
||
Other parts I enjoyed:
|
||
|
||
- **No AI SDK.** OpenAI-compatible streaming is just a `bufio.Scanner` reading `data:` lines, which is the "bare-basic" my friend asked for.
|
||
- **It runs as a real service.** `localchat install --user` registers a launchd agent, a systemd unit or a Windows service that points at your `.env`.
|
||
- **Everything is embedded** (htmx, the SSE extension, CSS, icon), so it works with the network cable pulled.
|
||
- **Tested:** handler tests against a fake OpenAI-style server (streaming, history, session isolation, XSS-safe Markdown, model offline), plus a Playwright browser test against the real Gemma model, both natively and in Docker.
|
||
|
||
## Why Does Open Innovation Matter?
|
||
|
||
- **Private by design.** Gemma's weights sit on the laptop, so there's no third-party server to trust. You can paste work code or private notes, and conversations aren't even written to disk.
|
||
- **Free to run, forever.** No tokens, no subscription, no free tier that disappears. My friend can leave it running without thinking about a bill.
|
||
- **Works offline.** On a train, on a plane, or behind a strict corporate network.
|
||
- **Swap anything.** The app only speaks the OpenAI-compatible protocol that the open ecosystem standardized on. I developed against oMLX and shipped with llama.cpp: two different inference engines, the same binary, no code changes. Moving to Gemma 4 E4B or another open model is one environment variable. A closed API would have tied the app to one vendor, one price list and one data policy.
|
||
- **Small open models are good enough now.** Gemma 4 E2B at 4-bit fits in a few GB of RAM and writes code snippets and explanations quickly on a laptop with no GPU. That makes a "bare-basic" AI app something you can just hand to a friend.
|
||
|
||
## My Agent Session
|
||
|
||
<!-- Paste the DevRelay agent_session embed here, or remove this section. -->
|
||
|
||
I built this with Claude Code as a pair-programmer: brainstorming ideas, scaffolding the Go/Templ/HTMX app, and testing it end to end against the local Gemma model.
|
||
|
||
## Prize Categories
|
||
|
||
- **Best Use of Gemma:** the whole app is built around running Google's Gemma 4 E2B locally (MLX on Mac, GGUF via llama.cpp in Docker).
|
||
|
||
<!-- Thanks for participating! -->
|