Skip to main content
Enterprise feature. This is part of OpenLIT Enterprise and needs an enterprise license applied to your organisation; the community edition doesn’t include it. Book a call or email developers@openlit.io to get access to the enterprise build and a license.
The GPUs page brings your GPU fleet into one view: how many GPUs are active and effectively used, device health and errors, which workloads and kernels run on each GPU, and what the fleet costs, with Otter highlighting what needs attention. Open GPUs from the sidebar under Monitor (/gpus).

Availability

The GPUs page is part of the OpenLIT Enterprise build. It doesn’t need a separate license feature. The community edition collects GPU metrics with the GPU collector, which you can chart in dashboards.

Before you start

  • Collect GPU metrics with the OpenLIT GPU collector. If no GPU data is arriving, the page shows ready-to-copy Kubernetes, Docker, and Linux commands that send metrics to your project’s metrics backend.
  • Bind a metrics backend for the project and environment in signal routing. The GPUs page reads the metrics binding, which can be the built-in ClickHouse, Prometheus, Datadog, or New Relic. Tempo, Loki, and Jaeger can’t serve GPU metrics.
  • Enable eBPF in the collector for Kernel Insights, with the OTEL_GPU_EBPF_ENABLED collector setting. Kernel Insights works with ClickHouse and Prometheus, not Datadog or New Relic.
  • Configure Otter with an LLM provider key in Otter settings for GPU insights and recommendations.

Explore your GPUs

Use the filters to narrow by GPU, Region, Device, Cluster, and Provider, and the time picker to change the range. Select a device in a table to open its own view at /gpus/<device-id>.

Track GPU cost

The Costs tab prices the GPU hours seen in the selected range:
  • Cloud instances reporting an instance type (host.type), provider, and region are matched to an instance rate and billed per host-hour.
  • Local or on-prem GPUs without an instance type are matched by GPU name and billed per GPU-hour.
  • Idle cost is the share of time GPUs weren’t allocated.
  • GPUs with no matching rate appear under Unmatched SKUs and are left out of the total, not counted as $0. Select add on an unmatched SKU to price it.
OpenLIT includes on-demand list prices for AWS, Google Cloud, Azure, Oracle Cloud, and CoreWeave. Under Instance rates you can add a rate, edit the USD per hour, or delete a rate. A rate for a specific region takes precedence over a rate for all regions (*), and your custom rates take precedence over list prices.

Otter insights and recommendations

Each tab can show a short Otter headline and notes about the metrics on screen. The Recommendations tab lists Otter’s suggestions with a severity (Critical, Major, Minor, or Info), the suggested action, and a link to the tab to investigate. Dismiss a recommendation with a reason, such as “A fix has already been started”, to move it to Closed, and Reopen it later. Dismissals are shared with everyone working in the same project environment and recorded in Audit Logs. Dismissing and reopening need the gpu:configure permission.

Permissions

Changes to instance rates and recommendation dismissals are recorded in Audit Logs when audit logging is licensed.

GPU collector

Collect GPU metrics from Kubernetes, Docker, and Linux hosts

Signal routing

Choose which backend serves metrics for each environment