Files
lostakj 92e54418bc
Build Docker image on push / docker (push) Successful in 21s
Build and push Docker image on tag / docker (push) Successful in 22s
Align metric names with Prometheus naming conventions
Use base units, export cumulative values as _total counters, update docs, dashboards and rules.
2026-08-24 19:13:05 +02:00

3.7 KiB

Architecture and Adding Collectors

Collection flow

config.yaml
    |
    v
Application -> PveApiClient -> Proxmox VE API host(s)
    |               |
    |               +-> liveness checks, round-robin selection, short API cache
    v
PveMetricsManager -> enabled collectors -> Prometheus registry -> /metrics
                              |
                              +-> TTL cleanup for stale label sets

Packages

Package Responsibility
application Loads configuration, creates dependencies, registers /metrics, and starts the HTTP server.
configuration YAML models and startup validation.
proxmox HTTP client, Proxmox-specific API methods, response models, and numeric state conversion.
metrics Collector implementations, periodic scheduling, Prometheus metric definitions, and TTL cleanup.
utils Shared formatting helpers.

API client

PveApiClient wraps the generic ApiClient with typed methods such as cluster status, node status, guest status, storage, disk, and ZFS calls.

The generic client:

  • normalizes API hosts so a trailing slash is optional;
  • checks every host at startup and every five seconds;
  • selects reachable hosts in round-robin order;
  • caches a successful method/path response briefly to avoid duplicate API calls;
  • uses API-token authorization and a ten-second HTTP timeout;
  • accepts self-signed certificates by disabling TLS verification.

Configured hosts must represent the same cluster because cached results are shared by method and path rather than host.

Metrics manager

PveMetricsManager creates only the collectors enabled in configuration. It runs one full collection immediately, then repeats it every proxmox.interval seconds. Collectors execute sequentially.

Each collector implements:

type PveMetricsCollector interface {
    CollectMetrics() error
    GetName() string
}

Successful execution time is recorded in pve_metrics_collection_duration_seconds. A collector error is logged and the manager continues with the next collector.

TTL metrics

Gauges use TTLGaugeVec, a wrapper around Prometheus GaugeVec. Cumulative values use TTLCounterVec, which exports the absolute value reported by the PVE API as a Prometheus counter (GaugeVec cannot be used because a counter name must carry a _total suffix, and CounterVec cannot be used because it only supports Inc/Add). Every label set records its last update. TTLRegistry checks registered metrics every five seconds and removes series that have not been updated for five minutes.

This prevents deleted guests or storage resources from remaining indefinitely, while allowing short API failures to preserve the last known values.

Adding a collector

  1. Add typed response models and API methods under proxmox.
  2. Add a collector implementing PveMetricsCollector under metrics.
  3. Create metrics with NewTTLGaugeVec or, for cumulative values, with NewTTLCounterVec, choose stable labels, and register each metric with the shared TTLRegistry.
  4. Add a boolean switch to PveMetricsConfiguration and config.example.yaml.
  5. Register the collector in NewPveMetricsManager.
  6. Add parsing, state-mapping, and topology tests as appropriate.
  7. Document the switch, labels, units, and state values in configuration.md and metrics.md.

Metric names must follow the Prometheus naming conventions; they are verified by TestPveMetricNamesFollowPrometheusConventions, which lints every registered metric with promlint.

Avoid labels containing changing messages, timestamps, or other unbounded values. For recursive resources such as ZFS, include a stable path so repeated component names remain distinguishable.