Datadog Dashboards

대시보드

총 3개
필터는 바로 적용됩니다
이름인기작성자변환 품질최근 수정
NVIDIA GPU for AI Workloads — Efficiency, Cost & Health (DCGM)gpunvidiadcgm

GPU monitoring for self-hosted AI: are your tensor cores actually working, what does each run cost in kWh, and is your VRAM degrading? Built and battle-tested on a local LLM rig (100B+ parameter models). Requires dcgm-exporter with DCP/profiling metrics enabled. Multi-GPU via UUID variable.

0
0
orangeoctopus1069좋음
NVIDIA GPU Operator & DCGM Observabilityk8snvidiagpu-operator

Live NVIDIA GPU Operator and DCGM exporter observability. Covers all 23 active DCGM metric families and GPU Operator health metrics; every panel has a Grafana hover-info description.

0
0
nafey1보통
GPU Simulation - DCGM Overviewgpudcgmsimulation

NVIDIA GPU overview from DCGM metrics: utilisation, framebuffer memory, temperature and power, one series per GPU. Works against a real dcgm-exporter unchanged; temperature and power ship as recording rules.

0
0
ChrisAdkin좋음