Usage Statistics
Historical request usage loaded from the Docker-mounted telemetry log. Choose a time window and the model profiles to include.
Models included (1 / 8 selected)
GB10_16 models
NVIDIA 50902 models
Showing 6,356 requests from 2026-09-15 05:32 through 2026-09-22 05:32. Averages exclude zero-valued observations/buckets. Dashboard Reset does not delete this retained history. Per-model Clear resets a live model's usage or removes a historical-only entry from the selector; the retained file and global all-time totals are never rewritten. History file:
/persist/data/token-events.jsonlRequests
6,356
54 errors ยท 64 missing usage
Input Tokens
830,339,536
prompt/input usage
Output Tokens
1,809,393
generated/completion usage
Avg Duration
11.67s
TTFT 4.88s average
Avg Output Rate
34.3
tokens/sec during active generation
Requests
Completed requests, errors, and requests without upstream usage data.
requestserrorsmissing usagerequest avg 288.91 reqerror avg 13.5 reqmissing avg 9.14 req
1.28K req0 req
09-15 05:3209-22 05:32
Input Tokens
Prompt/input tokens consumed in each time bucket.
input tokensavg 39.54M tok
104.45M tok0 tok
09-15 05:3209-22 05:32
Output Tokens
Generated/completion tokens produced in each time bucket.
output tokensavg 86.16K tok
185.88K tok0 tok
09-15 05:3209-22 05:32
Latency
Average end-to-end request duration and time to first token.
durationTTFTduration avg 11.67 sTTFT avg 4.88 s
21.13 s0 s
09-15 05:3209-22 05:32
Generation Throughput
Output tokens divided by active generation seconds for each bucket.
output tok/savg 34.3 tok/s
43.8 tok/s0 tok/s
09-15 05:3209-22 05:32
Breakdown by Model
| System | Model | Requests | Errors | Input | Output | Total | Avg duration | Avg TTFT | Avg tok/s |
|---|---|---|---|---|---|---|---|---|---|
| GB10_1 | qwen3.8-flash-nextgb10_1_vllm_qwen3_8_flash_next | 6,356 | 54 | 830,339,536 | 1,809,393 | 832,148,929 | 11.67s | 4.88s | 34.3 |