API access
OpenAI-compatible endpoint — point any client here:
Bearer token (send as Authorization: Bearer …).
Renewing invalidates the old token immediately:
••••••••••••••••••••
Model serving bind
Where the vLLM container listens.
Applies on the next model start. Wherever it binds, the model
port requires the same bearer token as the proxy (only
/health and /metrics stay open).
A renewed token reaches this port on the next model start.
Models
Live
No model running — live stats appear here once one is up.
History last hour · 5s samples
Charts fill in as samples arrive (one every 5 seconds while a model container exists).
Details
GPU
Effective launch flags
Logs
No model running
Messages go to whatever model is currently served
(as default). Attach text files to inline them, or
images for vision models.
ctx —