LLM Control Plane v2.71
Unified OpenAI proxy + MCP agent broker + vLLM administration

vLLM Model Runtime Bindings

A vLLM host is an Endpoint whose runtime type is vllm. Container, launch script, and SSH host-control binding are stored on the selected ModelProfile.

No container name is bound to this model profile.

Host Control Profiles

SSH connection records only. They no longer carry detached container/script allowlists; host actions use the exact container/script bound to the selected vLLM ModelProfile.

vLLM Endpoint Inventory

Only endpoints with RuntimeType=vllm are eligible vLLM hosts here.

Endpoint / hostURLOnlineModel profiles on this endpoint
GB10_1
gb10_1
GB10_1
http://192.168.10.1:8003/v1online
HTTP 200 OK
4
gb10_1_vllm_deepseek_ai_deepseek_v4_1_flash, gb10_1_vllm_deepseek_v41_flash_exl3, gb10_1_vllm_deepseek_v4_flash_dspark_abliterated, gb10_1_vllm_qwen3_8_flash_next

vLLM Model Binding Inventory

Model profileEndpointHost controlContainer/scriptLaunch shape
deepseek-ai/DeepSeek-V4.1-Flash
gb10_1_vllm_deepseek_ai_deepseek_v4_1_flash
served: deepseek-ai/DeepSeek-V4.1-Flash
GB10_1
http://192.168.10.1:8003/v1

port= ctx= seqs= batch= gpu= kv= quant=
deepseek-v41-flash-exl3
gb10_1_vllm_deepseek_v41_flash_exl3
served: deepseek-v41-flash-exl3
GB10_1
http://192.168.10.1:8003/v1

port= ctx= seqs= batch= gpu= kv= quant=
deepseek-v4-flash-dspark-abliterated
gb10_1_vllm_deepseek_v4_flash_dspark_abliterated
served: bonsai-27b-android-local
GB10_1
http://192.168.10.1:8003/v1

port= ctx= seqs= batch= gpu= kv= quant=
qwen3.8-flash-next
gb10_1_vllm_qwen3_8_flash_next
served: qwen3.8-flash-next
GB10_1
http://192.168.10.1:8003/v1

port= ctx= seqs= batch= gpu= kv= quant=