Streams replies over SSE from any OpenAI-compatible server (oMLX, llama.cpp, Ollama, ...). Single binary that can install itself as an OS service; docker compose bundles llama.cpp + Gemma 4 E2B.
5 lines
24 B
Plaintext
5 lines
24 B
Plaintext
bin/
|
|
.env
|
|
*.local
|
|
dist/
|