Skip to main content

Trace Storage and Retention

The Netdata Agent stores the OpenTelemetry traces it receives on its own disk, indexed, and serves them to the Traces tab of the dashboard. This page covers how spans are stored, how long they are kept, offloading to object storage, and sizing the receiving node. To send traces to the Agent, see OTLP Ingestion.

How spans are stored​

Spans received over OpenTelemetry are stored by the receiving Netdata Agent under base_dir (default /var/log/netdata/otel/v2 on package installs), in its traces/ directory. Incoming spans are appended to one write-ahead log. When it reaches traces.rotation.default.max_file_size (25MB) or traces.rotation.default.max_entries (50000 spans), or about 15 minutes after its first span, it is sealed into an indexed file and the write-ahead log is deleted. Each indexed file indexes the span names, kinds, and statuses and the span, resource, scope, event, and link attributes, and keeps an index of trace IDs so that a whole trace can be looked up by its ID. A span sent without a kind or a status is indexed with the OpenTelemetry defaults, UNSPECIFIED and UNSET, so those values can be filtered on like any other. Spans stored by an earlier version keep no kind or status for those defaults, so until their files leave retention a filter on UNSET, UNSPECIFIED, or a negation such as a status other than ERROR does not match them.

Retention​

Retention applies to sealed indexed files, oldest first, when any of three limits is exceeded:

OptionDefaultMeaning
traces.retention.default.max_files100000Maximum number of indexed files kept
traces.retention.default.max_total_size1GBMaximum total size of indexed files kept
traces.retention.default.max_age7 daysMaximum age of an indexed file, measured on the start time of its newest span

max_total_size is not a cap on the plugin's disk usage: the active write-ahead log (up to max_file_size), catalogs, and the download cache for offloaded data are additional. Retention runs when a file is sealed; when spans stop arriving, the last write-ahead log seals on idle (within about 15 minutes) and gets one final pass, and the remaining files are kept until spans arrive again or the Agent restarts.

Traces have their own settings, separate from logs. Set the limits under traces.retention.default; a user otel.yaml needs only the fields that change:

traces:
retention:
default:
max_total_size: "20GB"
max_age: "30 days"

Edit otel.yaml with edit-config and restart the Agent. The full option list is in the OpenTelemetry plugin reference.

Offloading to object storage​

Traces use the same remote_storage settings as logs: one object storage backend, and one download cache at <base_dir>/remote-read, bounded by remote_storage.read_cache_max_size and shared by both. With remote_storage.enabled: true, every sealed trace file is also uploaded, a local file is not deleted by retention until the remote confirms it, and a query that needs an offloaded file that is no longer local downloads it into the download cache. Setup, credentials, upload retries, and the download cache are described in Log Storage and Retention.

Reading traces back from object storage has these limits:

  • A query that needs more offloaded data than the download cache holds fails with a message to narrow the time range or raise remote_storage.read_cache_max_size. A search also reads up to 24 hours beyond each side of its window to complete the traces it finds.
  • Looking up one trace by its ID without a time range searches everything kept, local and offloaded, so it fails once the offloaded history exceeds the download cache.
  • A file that cannot be downloaded, or a catalog that cannot be read, makes the answer partial with the reason remote_unavailable; one missing file in a search's range marks every trace summary as possibly incomplete and empties searches that filter on the root service, root operation or trace duration. A file that downloads intact but cannot be read is reported as a failed source (source_failure).
  • Files are downloaded one at a time, each within its own time limit. While the remote is down, each file waits out its limit in turn (a file that another query is already downloading can cost up to about three limits), so a query over many offloaded files can take minutes.

Sizing the receiving node​

Run the senders for a full day, then measure du -sh on <base_dir>/traces/index/ and multiply by the retention you want locally; set max_total_size to that size, since the 1GB default is usually reached long before the 7-day max_age. For longer retention, keep max_age short, enable offloading, and size the object storage for one day's index size × the time you keep objects there; the Agent does not delete offloaded files, so expire them with the object storage's own lifecycle rules. Add headroom for the active write-ahead log (max_file_size) and, when offloading is enabled, for the download cache, which logs and traces share.

Exploring traces​

Stored traces are explored in the Traces tab, which queries the Agent's otel-traces Function. As with logs, viewing traces requires a signed-in Netdata Cloud user of the Agent's Space; trace data is not stored in Netdata Cloud.

Where to next​


Do you have any feedback for this page? If so, you can open a new issue on our netdata/learn repository.