How to Write a go.d Collector (V2)
This is the canonical starting point for new go.d collectors. New collectors MUST use framework V2. V1 collectors remain in the tree for compatibility and maintenance only.
For migrating an existing V1 collector, use src/go/plugin/go.d/docs/migrate-v1-to-v2.md instead. Migration is
compatibility work and has different rules from new collector authoring.
Use src/go/plugin/go.d/collector/cato_networks/ as the primary modern example. It is large, so copy the pattern, not
the whole shape. The useful references are called out below by responsibility.
Before Writing Code
Do the design work first. For a new collector, or a change to a public contract (config option, mode, metric meaning,
ownership of state, Functions, vnodes), fill the collector design note from
.agents/skills/collectors-go-design/SKILL.md in the SOW gate before code; the items below are its short form.
- Read the upstream API or protocol docs. Do not infer current behavior from memory or from generated SDK types alone.
- Check existing helper packages before implementing parser, HTTP, selector, command-execution, SQL, ping, log-reading,
or log-limiting plumbing. Start with
src/go/plugin/go.d/docs/helper-packages.md. - You MUST aim for the clean end state, not the smallest initial diff. If the clean collector design requires a
framework improvement, surface that as a design decision and follow
src/go/plugin/framework/docs/changing-framework-code.mdinstead of hiding it behind collector-local glue. - Decide the monitored entities and cardinality bounds. If one job collects remote resources that SHOULD be separate Netdata nodes, design V2 host scopes from the start.
- Decide the minimal public config surface. Public config is a compatibility contract. Use constants for internal tuning such as page limits, scan cadence, retry limits, fan-out concurrency, and cache TTLs unless the operator has a real decision to make. A proposed config option MUST name that concrete operator decision; "operators may want to tune it" is not enough.
- Decide whether the collector needs Functions or topology. Functions are interactive live/snapshot views; metrics are
time series. New topology producers MUST use
src/go/pkg/topology/v1and validate againstsrc/plugins.d/FUNCTION_TOPOLOGY_SCHEMA.json. - Plan collector consistency using
.agents/skills/integrations-lifecycle/consistency.md. Generated integration pages and README symlinks are outputs, not hand-authored sources. - Plan the first coherent batch and its boundaries. At each boundary, you MUST re-check whether new work has drifted out of scope; defer it or land it independently before continuing.
Source References
Primary V2 reference:
src/go/plugin/go.d/collector/cato_networks/
Read these files by responsibility:
collector.go: registration, defaults, public lifecycle methods,MetricStore(), the chart-provider getter, and Function wiring.config.go: config defaults, normalization, validation, and intentionally small public config.collect.go,collect_metrics.go,collect_bgp.go: collection orchestration and split domain operations.metrix.go,write_metrics.go,charts.yaml: typed instruments, metric writes, chart template,StateSet,instances.by_labels, andlabel_promotion. Audit everyinstances.by_labelsidentity choice instead of copying labels from the example blindly.host_scope.go: deterministic per-site V2 host scopes/vnodes.func_deps.go,catofunc/: Function subpackage boundary behind a narrow dependency interface.topology_store.go,topology.go,topology_test.go: immutable topology snapshot publishing and topology v1 schema validation.config_test.go,collector_lifecycle_test.go,collector_collect_test.go,charts_test.go: table-driven V2 tests and fixture validation.
Framework/API references:
src/go/plugin/framework/collectorapi/collector.gosrc/go/plugin/framework/docs/changing-framework-code.mdsrc/go/plugin/go.d/docs/helper-packages.mdsrc/go/pkg/metrix/README.mdsrc/go/plugin/framework/charttpl/README.mdsrc/go/plugin/framework/chartengine/README.mdsrc/go/plugin/framework/functions/README.mdsrc/go/tools/functions-validation/README.md.agents/skills/collectors-go-framework-v2/go-v2-host-scope.md.agents/skills/integrations-lifecycle/consistency.md
File Layout
Start with this layout and add focused files only when a responsibility needs its own boundary:
src/go/plugin/go.d/collector/<name>/
|-- collector.go # registration, New, public lifecycle, store/template
|-- init.go # Init helper methods for clients/matchers/state
|-- config.go # Config, defaults, validation
|-- collect.go # Collect orchestration
|-- metrix.go # typed metrix instruments built once in New
|-- write_metrics.go # normalized state -> metrix observations
|-- models.go # collector-local state/DTOs
|-- client.go # API/client boundary
|-- charts.yaml # V2 chart template
|-- config_schema.json # DYNCFG schema
|-- metadata.yaml # integration metadata source
|-- integrations/ # generated integration page
|-- README.md # symlink to generated integration page
|-- testdata/ # fixtures and config serialization files
`-- *_test.go # table-driven tests
Common optional splits:
Most collectors need none of these optional files. Add one only when the collector has the corresponding product surface
or state boundary; do not create empty host_scope.go, topology.go, or <name>func/ files just because Cato has
them.
init.gowhenInit()needs helper setup for clients, matchers, caches, or other persistent state. Keep the publicInit()method itself incollector.go; let it call focused helpers such asinitClient()orinitSiteSelector().collect_<operation>.gowhen the collector has multiple distinct collection operations, such as discovery, account metrics, BGP, or inventory.normalize_<operation>.gowhen API payload normalization would otherwise dominatecollect.go.host_scope.gowhen the collector emits generated vnodes.<name>func/plusfunc_deps.gowhen the collector exposes Functions.topology.goandtopology_store.gowhen the collector emits topology.
Avoid files whose names hide their responsibility. For example, a file named diagnostics.go SHOULD NOT contain only
error classification.
Registration And Lifecycle
New collectors MUST implement collectorapi.CollectorV2 from src/go/plugin/framework/collectorapi/collector.go and
register via CreateV2. In practice, collector.go should:
- embed
config_schema.jsonforJobConfigSchema; - provide exactly one chart capability: embedded
charts.yamlthroughChartTemplateYAML()for static definitions, orChartTemplateSet()for changing native entries; - expose
Config: func() any { return &Config{} }; - return a new collector from
CreateV2; - add
MethodsandMethodHandleronly when the collector has Functions.
New() SHOULD own defaults and test seams:
- create
metrix.NewCollectorStore(); - build typed instruments once from that store;
- set default config values;
- set injected seams such as client factories or clocks;
- create the Function router when Functions exist.
Public lifecycle and framework-contract methods MUST stay in collector.go:
Configuration() anyInit(context.Context) errorCheck(context.Context) errorCollect(context.Context) errorCleanup(context.Context)MetricStore() metrix.CollectorStoreChartTemplateYAML() stringorChartTemplateSet() *chartengine.TemplateSet
Native snapshot ownership, construction, replacement and fixed-policy rules are documented in
chartengine. Existing YAML collectors retain
their getter; implementing both providers is a startup error. collecttest.AssertChartCoverage supports either provider.
Collectors that need a receiver or long-running background loop MAY implement collectorapi.CollectorV2Runner with
Run(ctx context.Context, ready func()) error. Acquire exclusive runtime resources there, after predecessor cleanup,
then call ready() once all fallible startup prerequisites succeed and state is safe for concurrent Collect().
Readiness MUST NOT wait for a first packet or remote observation. Init() and Check() still run during configuration
validation while an incumbent may own the endpoint; DynCfg test never calls Run().
The framework waits for readiness before collecting. Startup errors follow the existing configured autodetection retry
policy; a collectorapi.PermanentError, unexpected nil return and recovered panic do not retry. Unexpected return after
readiness makes the job Failed and requires an explicit restart. The implementation MUST return promptly after
ctx.Done() and SHOULD make in-flight I/O cancellation-aware where the library permits. Cleanup waits for Run() to
return. See the
runtime readiness contract for
startup timeout, cancellation, output fencing and physical ownership semantics.
V2 collectors that need the Agent's existing first-sample storage behavior MAY set StoreFirst: true directly in
their collectorapi.Creator registration. This fixed collector-wide setting applies to every collector chart,
including automatic charts and later redefinitions. It defaults to false and does not change framework self-metrics
or preserve counter baselines across restarts. See
first-sample storage.
Init() validates config, prepares matchers/clients, and initializes persistent state. Explicit setup details SHOULD
live in helper methods, preferably in init.go, so the public method reads as the lifecycle sequence. Check() MUST be
a cheap auth/connectivity probe, not a full collection. Collect() MUST run the real write path through metrix.
Cleanup() closes idle connections and forwards Function cleanup.
An Init() or Check() error fails the job. By default, a Check() error is retried every autodetection_retry
seconds (never when it is 0, the framework default) and an Init() error is never retried. Classify the error when that
default is wrong:
collectorapi.PermanentError(err): a failure that retrying the same configuration cannot fix, such as an invalid option or an unknown named profile. The job is never retried.collectorapi.TemporaryError(err): a condition that may clear on its own, such as a dependency that is not ready yet. The job keeps theautodetection_retryschedule, even fromInit().
DynCfg update rejects every failed preflight regardless of autodetection_retry, preserving the incumbent:
422 for a permanent or unclassified collector failure, 503 for a temporary one. Successful replacement preflight
returns 202 before runtime startup completes. Non-running enable accepts intent with 202 before probing;
later failures are reported through job status and use the retry policy above. add is passive until enabled.
A classified failure keeps a failed stock job listed in DynCfg instead of removing it. Leave the expected absence of a
service at a stock endpoint unclassified. For all command outcomes, including restart and test, see
the Job Manager reply contract.
Config
Config SHOULD stay small and operator-oriented:
- connection identity and credentials;
- endpoint and standard HTTP/TLS/proxy fields when applicable;
update_every,timeout, andvnodewhen relevant;- selectors that let users intentionally scope cardinality.
Every configuration path (stock and user files, service discovery, DynCfg) applies confgroup.Config.ApplyDefaults
before the collector sees the job, replacing a non-positive update_every, autodetection_retry or priority with
the module default. A collector therefore always receives a positive update_every and priority and a non-negative
autodetection_retry, where zero is the valid "no retry" value. Collectors SHOULD NOT re-check those bounds; such
checks and their troubleshooting entries are unreachable.
Implementation tuning SHOULD use constants:
- discovery refresh cadence;
- page sizes and maximum pages;
- per-cycle fan-out concurrency;
- cache TTLs;
- retry/backoff internals;
- API batching constraints.
Durations are confopt.Duration (or confopt.LongDuration) written as 30m, never *_ms integers; do not write
custom parsing for units. You MUST NOT add a config option just because it is easy to expose. Once shipped, it is hard
to remove and MUST stay
synchronized across Config, config_schema.json, stock .conf, metadata, generated docs, and tests. A proposed
config option MUST name the concrete operator decision it enables; "operators may want to tune it" is not enough.
For SaaS/API credentials, examples SHOULD prefer secret indirection such as ${env:COLLECTOR_API_KEY} or
${file:/run/secrets/collector_api_key} instead of realistic-looking inline credentials. Schema fields that carry
secrets MUST use "ui:widget": "password" (the only signal that redacts the UI's YAML preview); writing the schema
file is covered by .agents/skills/collectors-go-design/config-schema.md.
Selectors SHOULD use existing matcher packages such as src/go/pkg/matcher unless the upstream API forces a different
grammar. Document the exact matching input, for example "site name when present, otherwise site ID."
Collect Flow
Collect() SHOULD stay orchestration, not a large parser. A typical flow is:
- ensure the client is initialized;
- refresh stable discovery only when needed;
- fetch the current snapshot/state needed for this cycle;
- enrich with optional or slower data;
- normalize API payloads into collector-local state;
- publish any immutable Function/topology snapshot;
- write metrics to
metrix.
When the collector performs several upstream calls or collection operations, those operations SHOULD be split into
focused files named by operation, for example collect_metrics.go or collect_bgp.go. collect.go SHOULD explain the
cycle; the operation files SHOULD own the operation-specific API calls, fail-soft behavior, and merge rules.
Fail-soft behavior MUST be used only when partial data is still truthful. If one optional operation fails, log a rate-limited warning and omit or preserve only values that remain honest. If the core operation fails, return an error with context.
Collect() MUST preserve context cancellation. If the context is canceled during a partial path, Collect() MUST
return the context error so the runtime aborts the cycle instead of committing a stale or partial frame.
Metrics And Charts
Metric instruments SHOULD be built once in New() when the metric surface is known. Use a typed collector metrics
struct so write code is a value mapping, not repeated dynamic instrument lookup. For input-defined metric names and
churning labels, construct direct snapshot handles inside the collection cycle or retain them only with their owning
series. A permanent Vec cache is unbounded; follow the descriptor-retention contract in
src/go/pkg/metrix/README.md#consumers-that-cache-per-name-state for any per-name cache that survives cycles.
Use the right instrument:
SnapshotGaugeVecfor labeled current values with a bounded, stable identity set; dynamic ingestion can use direct labeledSnapshotMeter.Gaugehandles with the ownership rules above;Counter.ObserveTotal()for source counters;StateSetSHOULD be used for fixed mutually exclusive states, such as connected vs disconnected or up vs down.
Use exactly one static charts.yaml provider or an immutable native TemplateSet provider as the chart contract.
Native providers reuse their prepared pointer until content changes; see
src/go/plugin/framework/chartengine/README.md#named-active-template-sets. Every static template MUST define:
version: v1;context_namespace;
Templates SHOULD also group charts by operational area, use instances.by_labels for stable instance identity when
charts are entity-scoped, and keep the default lifecycle unless a concrete reason exists to override it. Defaults,
families, ordering, statesets, values, labels and shared contexts are owned by
.agents/skills/collectors-go-framework-v2/chart-template.md.
Metric labels and chart instance labels MUST be bounded and stable. Use IDs for identity. Which labels to attach
and when to list label_promotion is owned by the chart template topic linked above. Do not blindly copy
instances.by_labels from Cato or any other example; audit every label used for chart identity and record why it is
stable enough for that collector.
Host Scopes And Vnodes
Use metrix.HostScope when one job emits data for remote entities that SHOULD appear as separate Netdata nodes. Labels
alone are not enough for that product semantics.
Rules:
ScopeKeyandGUIDare deterministic and based on stable IDs.- Hostname may use a human-readable name when safe, with stable fallback.
- Add
_vnode_type=<source>and useful source labels. - Route every metric for that remote entity through the same host scope.
- Keep the default host scope empty unless the metric truly belongs to the agent/job host.
Use .agents/skills/collectors-go-framework-v2/go-v2-host-scope.md for the framework contract.
Functions
Functions MUST NOT freely access collector internals. Put Function code in a dedicated subpackage, for example
<name>func/, with a narrow Deps interface declared by that subpackage.
Non-topology Function responses MUST conform to src/plugins.d/FUNCTION_UI_SCHEMA.json; topology Function responses
MUST conform to src/plugins.d/FUNCTION_TOPOLOGY_SCHEMA.json. Function payloads MUST be validated with
src/go/tools/functions-validation/ or an equivalent schema-validation test, and the validation method must be
recorded.
Pattern:
- collector package owns state and implements a small adapter in
func_deps.go; - Function package owns method IDs, router, handlers, presentation, and tests;
- Function package MUST import only framework/function types and other allowed dependencies, not the collector package;
- the Function package
Depsinterface MUST expose only the methods the Function needs and MUST NOT expose, return, or embed*Collector; Collector.Cleanup()forwards cleanup to the Function router.
Use catofunc/ as the primary example. Test the Function package with fake deps so the boundary is compile-enforced.
Topology
New topology producers MUST use src/go/pkg/topology/v1. The non-v1 root src/go/pkg/topology payload model has been
retired and MUST NOT be reintroduced for topology payloads.
Rules:
- build topology from normalized collector state, not directly from raw API payloads;
- publish immutable snapshots for Function readers;
- MUST NOT mutate a published topology value;
- use
src/go/plugin/go.d/collector/cato_networks/topology.goas the concrete construction reference for actors, links, detail tables, and telemetry fields; - MUST validate topology payloads in tests with both
topologyv1.ValidateDecodedDataandsrc/plugins.d/FUNCTION_TOPOLOGY_SCHEMA.json; seesrc/go/plugin/go.d/collector/cato_networks/topology_test.govalidateCatoTopologyV1Datafor the full marshal/decode/schema check shape; - follow
.agents/skills/topology-authoring/SKILL.mdfor actor/link/table design.
Repository Wiring
For a new collector <name>:
- Add the collector package under
src/go/plugin/go.d/collector/<name>/. - Import it in
src/go/plugin/go.d/collector/init.go. - Add the default stock config:
src/go/plugin/go.d/config/go.d/<name>.conf. - Add the module toggle in
src/go/plugin/go.d/config/go.d.conf. - Add or update
src/go/plugin/go.d/README.md. - Add health alerts under
src/health/health.d/<name>.confonly when alerts are useful and backed by emitted chart contexts. - If adding or changing service-discovery rules under
src/go/plugin/go.d/config/go.d/sd/orsdext, update generated service-discovery documentation through the integrations lifecycle recipe. - Write
metadata.yamlwith.agents/skills/collectors-metadata-yaml/SKILL.mdopen (one contract per field), then generateintegrations/<slug>.mdand the README symlink from it. Single-integration collector directories normally use the symlinked README. Multi-integration plugin directories may keep a hand-authored umbrella README; follow.agents/skills/integrations-lifecycle/consistency.md.
Use .agents/skills/integrations-lifecycle/recipes/add-go-collector.md for the integration-generation commands.
A collector that works only on some operating systems (its metadata.yaml supported_platforms) MUST NOT register
elsewhere, or its stock job starts and fails there:
- Put the matching build constraint, such as
//go:build linux, on every source and test file of the collector package exceptdoc.go, after the SPDX line with a blank line before and after. - Add an untagged
doc.gocontaining only the package clause and its doc comment, so theinit.goimport still compiles on every platform. - Platform-neutral subpackages such as
<name>func/andinternal/need no constraint; only the tagged files link them. - From
src/go, run the package tests on a supported platform, then check a platform thatsupported_platformsexcludes (for exampleGOOS=windowsfor a Linux-only collector): with thatGOOS,go build -o /dev/null ./cmd/godpluginsucceeds andgo list -f '{{.GoFiles}}' ./plugin/go.d/collector/<name>prints only[doc.go].
Examples: ap, zfspool, smbios_memory.
The PR description or design note MUST enumerate the relevant collector consistency artifacts and justify every artifact that did not need a matching change. Most of this is not CI-enforced; it must be reviewer-visible.
Tests
Tests SHOULD be table-driven with map[string]struct{} when cases share setup and assertion shape.
Recommended test coverage:
- config JSON/YAML serialization with
collecttest.TestConfigurationSerialize; - config validation, including required credentials and unsafe URLs;
Init, cheapCheck,Collect,Cleanup,MetricStore;- hard failure, partial failure, context cancellation, and recovery behavior;
- chart-template schema validation with
collecttest.AssertChartTemplateSchemaand chart-template compile validation through the chartengine path used by nearby V2 collectors; - post-collect chart coverage with
collecttest.AssertChartCoverage; - artifact drift checks between
metadata.yaml,charts.yamlandhealth.dwith thecollecttestchecks named in.agents/skills/collectors-go-framework-v2/chart-template.md#tests; never restate template contexts or dimensions; - state-set values for every known state and unknown fallback;
- host-scope routing when scopes/vnodes are used;
- Function handler tests with fake deps when Functions exist, plus the Live Data drift check
collecttest.AssertMetadataDocumentsFunctionsagainstmetadata.yaml; - topology schema validation when topology exists;
- fixture validity and attribution when fixtures come from public third-party projects.
Do not let tests depend on real credentials or live services unless the test is explicitly an integration test gated outside the default unit-test path.
Validate Locally
From src/go, run the narrow collector tests:
go test -count=1 ./plugin/go.d/collector/<name>/...
Verify that go.d can load the module:
timeout 15s go run ./cmd/godplugin -m <name> -d
Success means the module is registered, a job starts, and the command keeps running until the timeout stops it. Treat
unknown module, no jobs started, config-load errors, or an immediate exit before the timeout as failures. Use
-c <config-dir> when the test config lives outside the normal go.d config search path.
When the collector uses concurrency or Functions, also run:
go test -race -count=1 ./plugin/go.d/collector/<name>/...
When integration metadata, generated pages, or health alerts change, run the relevant integrations pipeline
checks from .agents/skills/integrations-lifecycle/.
Do not claim full-project validation from a narrow collector command. State exactly what was run.
Anti-Patterns
- New collector using
Collect() map[string]int64. - Full live collection from
Check(). - Starting operational background polling from
Check()orInit(); use the optional V2 runner hook when polling must be tied to the running job lifecycle. - Public config knobs for internal implementation details.
- Custom selector or retry framework when existing package/framework behavior is enough.
- Collector-local singleton, adapter, or glue code that substitutes for a missing shared framework capability.
- Per-cycle warning/error logs for recoverable partial failures.
- Metric charts for collector internals when logs are enough.
- Mutable names used as chart or vnode identity.
- Function package holding
*Collector. - New topology producer using legacy topology payloads.
- Hand-written
README.mdwhen the integration page should be generated.
Do you have any feedback for this page? If so, you can open a new issue on our netdata/learn repository.