Streams replies over SSE from any OpenAI-compatible server (oMLX, llama.cpp, Ollama, ...). Single binary that can install itself as an OS service; docker compose bundles llama.cpp + Gemma 4 E2B.
15 lines
448 B
Bash
15 lines
448 B
Bash
# Where your OpenAI-compatible model server lives (llama.cpp, oMLX, Ollama, LM Studio, ...)
|
|
LLM_BASE_URL=http://127.0.0.1:8080/v1
|
|
# Leave empty to use the first model the server lists
|
|
LLM_MODEL=
|
|
# Only if your server requires one
|
|
LLM_API_KEY=
|
|
|
|
# Make it yours
|
|
LOCALCHAT_TITLE=localchat
|
|
# SYSTEM_PROMPT="You are a patient Go tutor. Keep answers short and show small code examples."
|
|
|
|
# LOCALCHAT_ADDR=127.0.0.1:3000
|
|
# MAX_HISTORY=20
|
|
# LLM_TIMEOUT=5m
|