vLLM Model Runtime Bindings
A vLLM host is an Endpoint whose runtime type is vllm. Container, launch script, and SSH host-control binding are stored on the selected ModelProfile.
No container name is bound to this model profile.
Host Control Profiles
SSH connection records only. They no longer carry detached container/script allowlists; host actions use the exact container/script bound to the selected vLLM ModelProfile.
vLLM Endpoint Inventory
Only endpoints with RuntimeType=vllm are eligible vLLM hosts here.
| Endpoint / host | URL | Online | Model profiles on this endpoint |
|---|---|---|---|
GB10_1gb10_1GB10_1 | http://192.168.10.1:8003/v1 | online HTTP 200 OK | 4 gb10_1_vllm_deepseek_ai_deepseek_v4_1_flash, gb10_1_vllm_deepseek_v41_flash_exl3, gb10_1_vllm_deepseek_v4_flash_dspark_abliterated, gb10_1_vllm_qwen3_8_flash_next |
vLLM Model Binding Inventory
| Model profile | Endpoint | Host control | Container/script | Launch shape |
|---|---|---|---|---|
deepseek-ai/DeepSeek-V4.1-Flashgb10_1_vllm_deepseek_ai_deepseek_v4_1_flashserved: deepseek-ai/DeepSeek-V4.1-Flash | GB10_1http://192.168.10.1:8003/v1 | | | port= ctx= seqs= batch= gpu= kv= quant= |
deepseek-v41-flash-exl3gb10_1_vllm_deepseek_v41_flash_exl3served: deepseek-v41-flash-exl3 | GB10_1http://192.168.10.1:8003/v1 | | | port= ctx= seqs= batch= gpu= kv= quant= |
deepseek-v4-flash-dspark-abliteratedgb10_1_vllm_deepseek_v4_flash_dspark_abliteratedserved: bonsai-27b-android-local | GB10_1http://192.168.10.1:8003/v1 | | | port= ctx= seqs= batch= gpu= kv= quant= |
qwen3.8-flash-nextgb10_1_vllm_qwen3_8_flash_nextserved: qwen3.8-flash-next | GB10_1http://192.168.10.1:8003/v1 | | | port= ctx= seqs= batch= gpu= kv= quant= |