Streams replies over SSE from any OpenAI-compatible server (oMLX, llama.cpp, Ollama, ...). Single binary that can install itself as an OS service; docker compose bundles llama.cpp + Gemma 4 E2B.
7 lines
35 B
Plaintext
7 lines
35 B
Plaintext
bin/
|
|
docs/
|
|
scripts/
|
|
.env
|
|
.git
|
|
*.md
|