cm-llm-manager apollo

Which model apollo is serving, and the only safe way to change it

Routes

RouteMethodsAuthPurpose
/ GET open Status page.
/active GET open What model is serving, and is it ready. The route other scripts poll.
/docs GET open Generated route list.
/health GET open Liveness of this app. 200 even when no model is running -- read /active for that.
/login POST open Exchange a token for the cookie the page's buttons use.
/port GET open Who holds port 11437 right now. The thing to curl when something looks wrong.
/repair POST token Clear a degraded state: stop everything and confirm the port came free.
/stacks GET open Every llama stack on this host, with why each one may or may not be used.
/stacks/<name> GET open One stack, as `docker compose config` resolves it.
/status GET open Everything in one object: active model, every stack, and recent switches.
/stop POST token Stop every managed stack and leave the GPU idle.
/switch GET open The switch in flight, if any, and the recent ones.
/switch POST token Stop whatever is running and bring up the named stack.
/switch/<job_id> GET open One switch, with its phases and timings.
/version GET open Which commit is actually running.

Switching from a script

curl -s http://192.168.10.124:5041/active curl -s -XPOST http://192.168.10.124:5041/switch \ -H "Authorization: Bearer $TOKEN" -H 'Content-Type: application/json' \ -d '{"stack":"llama-rocm-qwen38-27b-mtp"}' curl -s http://192.168.10.124:5041/switch/<job_id>
Use the LAN address, not the public name: Cloudflare answers repeated API calls with 403.