Datadog Dashboards

NVIDIA GPU for AI Workloads — Efficiency, Cost & Health (DCGM)

Good 910 downloads0 views0 commentsRevision 1Created Converted from Grafana dashboard #25526 by orangeoctopus1069
How to import

In Datadog: Dashboards → New Dashboard → ⚙ Configure → Import dashboard JSON → paste this JSON.

Reactions

GPU monitoring for self-hosted AI: are your tensor cores actually working, what does each run cost in kWh, and is your VRAM degrading? Built and battle-tested on a local LLM rig (100B+ parameter models). Requires dcgm-exporter with DCP/profiling metrics enabled. Multi-GPU via UUID variable.

Conversion quality

Good 91

Share of panels whose queries were translated without loss. Partial and unsupported panels keep the original PromQL in a note widget.

  • Native0
  • OpenMetrics21
  • Partial0
  • Unsupported2

Metrics without a native Datadog mapping

These metric names were kept in OpenMetrics naming. Adjust them if you collect with a native Datadog integration.

DCGM_FI_DEV_FB_USED ×3DCGM_FI_DEV_FB_FREE ×3DCGM_FI_DEV_GPU_TEMP ×2DCGM_FI_DEV_POWER_USAGE ×2DCGM_FI_DEV_GPU_UTIL ×2DCGM_FI_PROF_PIPE_TENSOR_ACTIVE ×2DCGM_FI_DEV_TOTAL_ENERGY_CONSUMPTION ×2DCGM_FI_PROF_GR_ENGINE_ACTIVEDCGM_FI_PROF_DRAM_ACTIVEDCGM_FI_DEV_MEM_COPY_UTILDCGM_FI_DEV_ENC_UTILDCGM_FI_DEV_DEC_UTILDCGM_FI_DEV_FB_RESERVEDDCGM_FI_DEV_MEMORY_TEMPDCGM_FI_DEV_SM_CLOCKDCGM_FI_DEV_MEM_CLOCKDCGM_FI_PROF_PCIE_TX_BYTESDCGM_FI_PROF_PCIE_RX_BYTESDCGM_FI_DEV_PCIE_REPLAY_COUNTERDCGM_FI_DEV_CORRECTABLE_REMAPPED_ROWSDCGM_FI_DEV_UNCORRECTABLE_REMAPPED_ROWSDCGM_FI_DEV_ROW_REMAP_FAILUREDCGM_FI_DEV_VGPU_LICENSE_STATUS

Unsupported panels

  • GPU Temp — distribución (histogram) histogram
  • Power — distribución (histogram) histogram

Revisions

  • Revision 1Converted from grafana.com revision 1Download JSON

Rate this dashboard

Sign in to rate.

0 comments

No comments yet.