LLM Control Plane v2.71
Unified OpenAI proxy + MCP agent broker + vLLM administration

Usage Statistics

Historical request usage loaded from the Docker-mounted telemetry log. Choose a time window and the model profiles to include.

Models included (1 / 8 selected)
GB10_16 models
NVIDIA 50902 models
Showing 6,356 requests from 2026-09-15 05:32 through 2026-09-22 05:32. Averages exclude zero-valued observations/buckets. Dashboard Reset does not delete this retained history. Per-model Clear resets a live model's usage or removes a historical-only entry from the selector; the retained file and global all-time totals are never rewritten. History file: /persist/data/token-events.jsonl
Requests
6,356
54 errors ยท 64 missing usage
Input Tokens
830,339,536
prompt/input usage
Output Tokens
1,809,393
generated/completion usage
Avg Duration
11.67s
TTFT 4.88s average
Avg Output Rate
34.3
tokens/sec during active generation

Requests

Completed requests, errors, and requests without upstream usage data.

requestserrorsmissing usagerequest avg 288.91 reqerror avg 13.5 reqmissing avg 9.14 req
1.28K req0 req
09-15 05:3209-22 05:32

Input Tokens

Prompt/input tokens consumed in each time bucket.

input tokensavg 39.54M tok
104.45M tok0 tok
09-15 05:3209-22 05:32

Output Tokens

Generated/completion tokens produced in each time bucket.

output tokensavg 86.16K tok
185.88K tok0 tok
09-15 05:3209-22 05:32

Latency

Average end-to-end request duration and time to first token.

durationTTFTduration avg 11.67 sTTFT avg 4.88 s
21.13 s0 s
09-15 05:3209-22 05:32

Generation Throughput

Output tokens divided by active generation seconds for each bucket.

output tok/savg 34.3 tok/s
43.8 tok/s0 tok/s
09-15 05:3209-22 05:32

Breakdown by Model

SystemModelRequestsErrorsInputOutputTotalAvg durationAvg TTFTAvg tok/s
GB10_1qwen3.8-flash-next
gb10_1_vllm_qwen3_8_flash_next
6,35654830,339,5361,809,393832,148,92911.67s4.88s34.3