Skip to main content

NVIDIA Jetson

NVIDIA Jetson

Plugin: go.d.plugin Module: jetson

Maintained by Netdata

Overview​

Monitor GPU and memory controller utilization and clock frequencies on NVIDIA Jetson devices.

You get:

  • GPU utilization.
  • GPU clock frequency, as one value or, on devices that report it, one value per graphics processing cluster (GPC).
  • External memory controller (EMC) bandwidth utilization and clock frequency. The CPU and GPU share this memory, so high EMC utilization can reveal memory-bandwidth-bound workloads.

CPU, memory, and swap usage come from Netdata's standard Linux system monitoring; this collector adds the Jetson GPU and EMC readings those charts do not cover.

The collector runs NVIDIA's tegrastats --interval 1000 on the local host and reads the line of readings it prints every second. One tegrastats process runs for as long as the job runs, under the Netdata service account (usually netdata); the collector never elevates its privileges.

It stops only the tegrastats process it started; other instances on the host, such as one you run in a terminal, are not affected.

This collector is only supported on the following platforms:

  • Linux

This collector only supports collecting metrics from a single instance of this integration.

The Netdata service account must be able to run tegrastats. Because the collector never elevates privileges, a reading that your Jetson Linux release shows only to root is not collected.

Default Behavior​

Auto-Detection​

The stock configuration runs one job named localhost, which looks for tegrastats in Netdata's PATH and starts collecting when it is found. If tegrastats is missing when Netdata starts, the job does not start; restart Netdata after making it available.

Limits​

  • The readings depend on the Jetson module, the Jetson Linux release, and permissions. For example, tegrastats on Jetson Thor can report GPU clocks without GPU utilization. A reading that tegrastats does not report leaves a gap; it is never charted as zero.
  • If tegrastats stops printing readings, the charts show a gap after three seconds instead of repeating the last values.

Performance Impact​

Each job keeps one tegrastats --interval 1000 process running, the same load as running that command yourself. Netdata parses one line of its output per second. A longer update_every reduces how often Netdata stores readings, not how often tegrastats samples.

Setup​

You can configure the jetson collector in two ways:

MethodBest forHow to
UIFast setup without editing filesGo to Nodes → Configure this node → Collectors → Jobs, search for jetson, then click + to add a job.
FileIf you prefer configuring via file, or need to automate deployments (e.g., with Ansible)Edit go.d/jetson.conf and add a job.
important

UI configuration requires paid Netdata Cloud plan.

Prerequisites​

Make tegrastats available to Netdata​

The collector needs Netdata installed natively on the Jetson device and the tegrastats utility that NVIDIA Jetson Linux (installed with JetPack) provides. Netdata running in a container is not supported.

Check that the Netdata service account can run it:

sudo -u netdata tegrastats --interval 1000

Each line should contain a GR3D_FREQ (GPU) or EMC_FREQ (memory controller) reading; stop the command with Ctrl+C. Readings missing from this output are also missing from Netdata.

If the command is not found, restore tegrastats with NVIDIA's tegrastats deployment instructions. Netdata searches its own PATH, which adds /sbin, /usr/sbin, /usr/local/bin, and /usr/local/sbin to the PATH it starts with. If tegrastats is in another directory, set PATH in the [environment variables] section of netdata.conf to a list that includes it, then restart Netdata.

Configuration​

Options​

The option can be set globally or per job. One job covers the whole Jetson device.

Config options
OptionDescriptionDefaultRequired
update_everyData collection interval, in seconds. tegrastats samples once per second regardless; a longer interval charts the latest reading, not an average.1no

via UI​

Configure the jetson collector from the Netdata web interface:

  1. Go to Nodes.
  2. Select the node where you want the jetson data-collection job to run and click the ⚙ (Configure this node). That node will run the data collection.
  3. The Collectors → Jobs view opens by default.
  4. In the Search box, type jetson (or scroll the list) to locate the jetson collector.
  5. Click the + next to the jetson collector to add a new job.
  6. Fill in the job fields, then click Test to verify the configuration and Submit to save.
    • Test validates the provided settings and checks the collector's startup prerequisites. Successful validation does not guarantee that every metric will be available during collection.
    • If it fails, an error message appears with details (for example, connection refused, timeout, or command execution errors), so you can adjust and retest.

via File​

The configuration file name for this integration is go.d/jetson.conf.

The file format is YAML. Generally, the structure is:

update_every: 1
autodetection_retry: 0
jobs:
- name: some_name1
- name: some_name2

You can edit the configuration file using the edit-config script from the Netdata config directory.

cd /etc/netdata 2>/dev/null || cd /opt/netdata/etc/netdata
sudo ./edit-config go.d/jetson.conf
Examples​
Basic​

The stock configuration, with one job for the local Jetson device.

Config
jobs:
- name: localhost

Alerts​

There are no alerts configured by default for this integration.

Metrics​

Metrics grouped by scope.

The scope defines the instance that the metric belongs to. An instance is uniquely identified by a set of labels.

Charts show only readings present in the latest tegrastats line. A missing or invalid reading leaves a gap; a reported zero is charted as zero.

Per NVIDIA Jetson instance​

The Jetson device's integrated GPU and external memory controller (EMC).

This scope has no labels.

Metrics:

MetricDescriptionDimensionsUnit
jetson.gpu_utilizationGPU Utilizationutilizationpercentage
jetson.gpu_frequencyGPU FrequencyfrequencyMHz
jetson.emc_utilizationMemory Controller Bandwidth Utilization at Current Frequencyutilizationpercentage
jetson.emc_frequencyMemory Controller FrequencyfrequencyMHz

Per GPU graphics processing cluster​

One graphics processing cluster (GPC) of the GPU, present when tegrastats reports a clock for each GPC.

Labels:

LabelDescription
gpcPosition of the cluster in the tegrastats GPU clock list, starting at 0.

Metrics:

MetricDescriptionDimensionsUnit
jetson.gpu_gpc_frequencyGPU Graphics Processing Cluster FrequencyfrequencyMHz

Troubleshooting​

Diagnostics​

Debug Mode​

Important: Debug mode is not supported for data collection jobs created via the UI using the Dyncfg feature.

To troubleshoot issues with the jetson collector, run the go.d.plugin with the debug option enabled. The output should give you clues as to why the collector isn't working.

  • Navigate to the plugins.d directory, usually at /usr/libexec/netdata/plugins.d/. If that's not the case on your system, open netdata.conf and look for the plugins setting under [directories].

    cd /usr/libexec/netdata/plugins.d/
  • Switch to the netdata user.

    sudo -u netdata -s
  • Run the go.d.plugin to debug the collector:

    ./go.d.plugin -d -m jetson

    To debug a specific job:

    ./go.d.plugin -d -m jetson -j jobName

Getting Logs​

If you're encountering problems with the jetson collector, follow these steps to retrieve logs and identify potential issues:

  • Run the command specific to your system (systemd, non-systemd, or Docker container).
  • Examine the output for any warnings or error messages that might indicate issues. These messages should provide clues about the root cause of the problem.
System with systemd​

Use the following command to view logs generated since the last Netdata service restart:

journalctl _SYSTEMD_INVOCATION_ID="$(systemctl show --value --property=InvocationID netdata)" --namespace=netdata --grep jetson
System without systemd​

Locate the collector log file, typically at /var/log/netdata/collector.log, and use grep to filter for collector's name:

grep jetson /var/log/netdata/collector.log

Note: This method shows logs from all restarts. Focus on the latest entries for troubleshooting current issues.

Docker Container​

If your Netdata runs in a Docker container named "netdata" (replace if different), use this command:

docker logs netdata 2>&1 | grep jetson

Known Errors​

tegrastats executable not found in PATH​

Cause

No executable named tegrastats is in Netdata's PATH. This is expected on Linux hosts that are not Jetson devices, and when Netdata runs in a container.

Fix

On a Jetson device, make tegrastats available as described in Prerequisites, then restart Netdata. On other hosts, ignore the message or set jetson: no in the modules section of go.d.conf.

start tegrastats: error​

Cause

Netdata could not start tegrastats through its unprivileged command helper, nd-run.

Fix

Read the error text. If nd-run is missing or cannot be executed, repair or reinstall Netdata.

tegrastats source unavailable: error​

Cause

The running tegrastats exited or stopped printing readings. Netdata keeps restarting it, at most 30 seconds apart.

Fix

Run the check in Prerequisites and compare its output with the error text. tegrastats exited: exit status N means the utility failed; tegrastats stopped producing records means it kept running without printing complete lines.

no fresh tegrastats sample​

Cause

Netdata has not received a complete tegrastats line in the last three seconds, usually because tegrastats is starting, has stopped, or is being restarted.

Fix

If the error persists, look for tegrastats source unavailable messages in the collector log (see Diagnostics) and run the check in Prerequisites.

tegrastats sample contains no supported GPU or EMC readings​

Cause

The latest tegrastats line has neither a valid GR3D_FREQ (GPU) nor EMC_FREQ (memory controller) reading. Which readings tegrastats prints depends on the Jetson module, the Jetson Linux release, and the account that runs it.

Fix

Run the check in Prerequisites. If its output has no GR3D_FREQ or EMC_FREQ readings, this collector has nothing to chart on this device.


Do you have any feedback for this page? If so, you can open a new issue on our netdata/learn repository.