Configuration
August 28, 2026 ยท View on GitHub
This guide is the practical starting point for OTel Arrow Dataflow Engine
configuration. It walks from the file loaded by df_engine to the root
structure, runtime defaults, pipeline topology, node configuration, policies,
topics, observability, and validation workflow.
Read it sequentially if you are configuring the engine for the first time, or use the section headings as a checklist when reviewing YAML. For exact field semantics, defaults, precedence rules, and validation behavior, use the configuration model reference; for node-specific config payloads, use the core node catalog, development node catalog, and contrib node catalog.
Warning
This project is experimental. The configuration format is not yet stable and can change at any moment, including incompatible changes between releases.
Configuration Location
The df_engine binary reads one root configuration through --config.
If you need to build or run df_engine locally first, start with the
Development Setup section in the main README.
If --config is omitted, the engine looks for config.yaml in the current
working directory.
Supported config sources:
| Source | Description |
|---|---|
/path/to/config.yaml | Bare local path. Treated the same as file:. |
file:/path/to/config.yaml | Local file path. .json files parse as JSON; other files parse as YAML. |
env:MY_VAR | Environment variable containing the full configuration text. |
yaml:<content> | Inline YAML. :: expands nested keys for small test fragments. |
http://host/path | Unauthenticated HTTP GET. JSON is detected from Content-Type; otherwise YAML is assumed. |
The http: provider retries failed fetches with exponential backoff. https:,
authenticated config sources, and multi-file merge are not implemented.
Run with a config file:
cargo run -- --config configs/otlp-otlp.yaml
Validate a file before running it:
cargo run -- --config configs/otlp-otlp.yaml --validate-and-exit
Validation parses YAML or JSON, validates the root model, checks graph references, checks that every node type is registered in the binary, and runs node-specific config validation when the component provides it.
After loading, the CLI can override selected engine-level settings:
--num-cores--core-id-range--http-admin-bind
Raw configuration text supports environment substitution before parsing:
${env:VAR}: replace with$VAR; error if the variable is unset.${env:VAR:-default}: replace with$VAR, ordefaultwhen unset.${env:VAR:-}: replace with$VAR, or the empty string when unset.$$: literal$; use it only when escaping a sequence that would otherwise be treated as environment substitution, such as$${env:VAR}.
Configuration Structure
Every runtime file is a single root document that describes the engine process:
version: otel_dataflow/v1
policies: {}
topics: {}
engine: {}
groups:
default:
topics: {}
policies: {}
pipelines:
main:
type: otap
policies: {}
extensions: {}
nodes: {}
connections: []
Required root fields:
version: must beotel_dataflow/v1.groups: pipeline groups keyed by group id. A single engine can run many groups, and each group can run many pipelines.
Optional root fields:
policies: top-level policy defaults.topics: global topic declarations.engine: engine-wide settings.
Most simple configurations only need version and groups.
Minimal OTLP receive, batch, and export pipeline:
version: otel_dataflow/v1
groups:
default:
pipelines:
main:
nodes:
otlp/ingest:
type: receiver:otlp
config:
protocols:
grpc:
listening_addr: "127.0.0.1:4317"
batch:
type: processor:batch
config: {}
otlp/export:
type: exporter:otlp_grpc
config:
grpc_endpoint: "http://192.0.2.10:4317"
connections:
- from: otlp/ingest
to: batch
- from: batch
to: otlp/export
Receivers, Processors, and Exporters
Receivers, processors, and exporters are configured as pipeline nodes. Each
node has a pipeline-local id, such as otlp/ingest, batch, or otlp/export
in the example below, and a type:
nodes:
otlp/ingest:
type: receiver:otlp
config: {}
batch:
type: processor:batch
config: {}
otlp/export:
type: exporter:otlp_grpc
config: {}
The type can use either the OTel shortcut form or the full URN:
type: receiver:otlp
type: urn:otel:receiver:otlp
The node kind is inferred from the type. The engine does not use separate
Collector-style receivers, processors, and exporters maps.
Common node fields:
type: required node implementation URN or shortcut.description: optional human-readable description.config: node-specific configuration owned by the selected component.outputs: optional declared output ports for multi-output receivers or processors.default_output: optional default output port used by nodes that emit without selecting a port.capabilities: optional bindings from capability name to pipeline extension.entity: optional node entity enrichment metadata.header_capture: receiver-only transport header capture override.header_propagation: exporter-only transport header propagation override.
Core node types are listed in the core-node catalog. Each node links to a README beside its implementation with node-specific configuration examples, limits, stability, and telemetry.
Optional contrib nodes are listed in the contrib-node catalog. They are registered only when the corresponding crate features are enabled in the binary you run.
Pipeline Groups and Pipelines
The engine is designed to run and manage many pipelines in parallel inside one process. Groups are logical containers for related pipelines and can be mapped to operational boundaries such as a team, project, tenant, environment, or deployment slice.
groups:
ingest:
policies: {}
topics: {}
pipelines:
traces:
nodes: {}
connections: []
A pipeline is an executable graph:
type: pipeline data type. Defaults tootap.nodes: receivers, processors, and exporters in the data path.extensions: long-lived components available to nodes through capabilities.connections: explicit graph wiring.policies: optional pipeline-level policy overrides.
An otap pipeline is multi-signal by default. Logs, metrics, and traces can
move through the same pipeline graph, unlike the Collector model where
pipelines are usually split by signal type. You do not need a connector just to
move data from a traces path into a metrics path; model the routing or
conversion you need with nodes and explicit connections.
The engine does not infer pipeline order from node roles. If data should flow
between two nodes, add a connections entry.
Policies are scoped by hierarchy. For regular pipelines, precedence is:
- Pipeline-level
policies - Group-level
policies - Top-level
policies
Policy overrides apply by policy family rather than by deep-merging every
nested field. The process-wide memory limiter is only supported at top-level
policies.resources.memory_limiter.
Pipeline Core Allocation
Use policies.resources.core_allocation to select the worker cores assigned to
each regular pipeline. The policy can be set at the top-level, group, or
pipeline scope using the policy precedence described above. The default is
all_cores.
Use every process-visible core:
policies:
resources:
core_allocation:
type: all_cores
Have the controller select a fixed number of cores:
policies:
resources:
core_allocation:
type: core_count
count: 4
Or assign explicit inclusive core ranges:
policies:
resources:
core_allocation:
type: core_set
set:
- start: 0
end: 1
- start: 4
end: 5
The allocation strategies have different sharing behavior:
all_coresuses every core visible to the process. It neither reserves cores from other pipelines nor observes their reservations.core_countis controller-selected and exclusive. The controller excludes cores assigned tocore_countandcore_setpipelines and rejects the configuration if the requested positive count cannot be satisfied. A count of0, or an omitted count accepted from programmatic configuration, means all currently unreserved visible cores and fails if none remain.core_setuses exactly the configured ranges. Core IDs must be visible to the process, ranges must be valid and non-overlapping within the allocation, and the set cannot be empty. Explicitcore_setallocations may overlap one another, but their cores are unavailable to controller-selectedcore_countallocations.
On Linux, the controller discovers the CPUs available through process affinity
and cgroup restrictions and their NUMA topology. For core_count, it
deterministically prefers enough cores from one NUMA node, then falls back to
ascending visible core IDs when one node cannot satisfy the request. On
platforms where NUMA topology is unavailable, selection remains deterministic
by ascending visible core ID.
The controller resolves and validates all regular-pipeline placements before starting pipeline workers. Resource policies on the system observability pipeline are rejected. Its worker runs on one core that is not reserved and may overlap any regular-pipeline allocation.
For detailed policy guides, see:
Connections and Output Ports
Connections define the pipeline graph:
connections:
- from: otlp/ingest
to: batch
- from: batch
to: otlp/export
Connection sources must be receivers or processors. Exporters are sinks and
cannot be used as from endpoints.
Because connections are explicit, the configuration can describe topologies beyond a simple receiver-processor-exporter chain, including fan-in, fan-out, named output ports, competing consumers, topic bridges, and observability pipelines.
Fan-in and fan-out are explicit:
connections:
- from: [ingest/a, ingest/b]
to: batch
- from: batch
to: [worker/a, worker/b]
policies:
dispatch: one_of
dispatch: one_of sends each item to one destination. With multiple
destinations, the destinations act as competing consumers. The broadcast
dispatch policy is parsed but is not currently supported for multi-destination
connections.
Most nodes use the default output. Multi-output processors can expose named ports:
nodes:
router:
type: processor:type_router
outputs: ["logs", "metrics", "traces"]
config: {}
connections:
- from: router["logs"]
to: logs/export
- from: router["metrics"]
to: metrics/export
- from: router["traces"]
to: traces/export
If from omits a port selector, the engine selects the default output. When
a node declares outputs, any selected source port must be listed there.
For details, see Output Ports.
Topics
Topics are named in-process communication points. Use them when one pipeline should publish data that another pipeline consumes without direct pipeline-to-pipeline wiring.
Declare a global topic:
topics:
raw_signals:
description: "raw ingest stream"
backend: in_memory
Declare a group-local topic:
groups:
ingest:
topics:
raw_signals:
description: "ingest-local raw stream"
For a pipeline in a group, group-local topics override global topics with the same local name.
Publish to a topic with exporter:topic:
type: exporter:topic
config:
topic: raw_signals
Consume from a topic with receiver:topic:
type: receiver:topic
config:
topic: raw_signals
Use backend: in_memory for current runtime configurations. The quiver
backend is reserved in the schema and rejected by the current runtime.
Topic policies control balanced queue capacity, broadcast lag behavior, and Ack/Nack propagation across topic hops. For exact topic policy fields and limits, see Topic Declarations.
Engine Section
The optional engine section controls engine-wide settings:
http_admin: HTTP admin server bind address.telemetry: telemetry backend configuration shared across pipelines.observed_state: observed-state store settings.topics: engine-wide topic runtime defaults.observability: dedicated internal observability pipeline.custom: ignored by the engine and reserved for embedding applications.
HTTP admin bind example:
engine:
http_admin:
bind_address: "127.0.0.1:8080"
An observability pipeline reads internal telemetry and exports it like any other
pipeline. The engine installs one by default: metrics use the noop exporter,
and logs explicitly configured to use its use the console exporter. Global
and engine logs retain their console_async defaults and therefore bypass this
pipeline unless configured otherwise. Override it to send either signal
elsewhere:
engine:
observability:
pipeline:
nodes:
internal:
type: receiver:internal_telemetry
config: {}
otlp:
type: exporter:otlp_grpc
config: {}
connections:
- from: internal
to: otlp
Observability pipelines use the same node and connection model as regular
pipelines. They support channel_capacity, health, and telemetry policies,
but resource policies are intentionally not supported there. The pipeline is
mandatory and must contain exactly one connected internal telemetry receiver.
The receiver defaults to signals: [logs, metrics], while either signal can be
selected independently. Logs must remain enabled
when a log provider uses its. Optional metrics.interval and metrics.views
fields customize periodic export when metrics are selected. A logs-only
receiver drains the private ITS metric accumulator without converting or
emitting OTLP metrics, preserving registry cleanup and admin endpoint data.
The previous engine.telemetry.metrics SDK configuration is no longer
supported. For Prometheus scraping, bind engine.http_admin and use the fixed
/api/v1/metrics path. This endpoint does not apply receiver views and does not
reset the independent ITS export accumulator.
Migrating Internal Metrics from the Rust OpenTelemetry SDK
Move periodic export and views from engine.telemetry.metrics into the engine
observability pipeline. For example, this former SDK configuration exported
viewed metrics over OTLP/gRPC and exposed the unviewed metrics to Prometheus:
engine:
telemetry:
metrics:
readers:
- periodic:
interval: 60s
exporter:
type: otlp
config:
protocol: grpc/protobuf
endpoint: http://localhost:50066
temporality: delta
- pull:
exporter:
type: prometheus
config:
host: 0.0.0.0
port: 9091
path: /metrics
views:
- selector:
scope_name: engine
instrument_name: memory.rss
stream:
name: process_memory_usage
description: Process resident memory usage.
The equivalent native ITS configuration is:
engine:
http_admin:
bind_address: "0.0.0.0:9091"
observability:
pipeline:
nodes:
internal:
type: receiver:internal_telemetry
config:
signals: [metrics]
metrics:
interval: 60s
views:
- selector:
scope_name: engine
instrument_name: memory.rss
stream:
name: process_memory_usage
description: Process resident memory usage.
otlp:
type: exporter:otlp_grpc
config:
grpc_endpoint: http://localhost:50066
connections:
- from: internal
to: otlp
Apply these mappings when migrating:
- Remove
metrics.provider; ITS is now the only internal metrics path. - Move
periodic.intervalto the receiver'smetrics.interval. The receiver has one emission interval, so multiple periodic readers must share a cadence. Set it explicitly to6sif the former periodic-reader default was required. - Move
metrics.viewsunchanged to the receiver'smetrics.viewslist. - Replace a
grpc/protobufSDK exporter withexporter:otlp_grpc, changingendpointtogrpc_endpoint. - Replace an
http/protobufSDK exporter withexporter:otlp_http. The native exporter accepts a baseendpointand optional full signal-specific endpoint overrides. The native exporter does not support the formerhttp/jsonmode. - Remove exporter
temporality. Native ITS currently emits its canonical low-memory representation: synchronous counters and histograms are delta, while up-down and observed counters are cumulative. Instrument types describe how component code records measurements; they do not select a consumer-specific export preference. The pipeline cannot yet convert this representation to another temporality, andprocessor:temporal_reaggregationdoes not perform that conversion. Consequently, a former explicitcumulativeordeltapreference has no exact native equivalent yet. This limitation is tracked in #3543. - Replace a console reader with an
exporter:consolenode. Fan out the internal receiver to multiple exporters when they share the same emission interval. - Replace a Prometheus pull reader with
engine.http_admin.bind_addressand scrape/api/v1/metrics. The path is fixed, receiver views are not applied, and the bind address exposes the complete admin API rather than a dedicated metrics-only server. - Use
signals: [metrics]while log providers retain their defaults. Omitsignals, or includelogs, when a global, engine, or admin log provider usesits. To suppress metric conversion and export, select onlylogsand route that signal to a sink.
If engine.telemetry.metrics was omitted and no internal metrics were exported,
no migration is required: the built-in observability pipeline consumes metrics
with a noop exporter.
For exact engine-level fields, see Engine Section.
Other Information
Useful adjacent docs:
- Core-node catalog: node list and links to per-node configuration docs beside the implementation.
- Contrib-node catalog: optional feature-gated node implementations.
- Configuration model reference: exact field semantics, defaults, precedence, and validation behavior.
- Node and flow metrics: configure and interpret node item metrics and processor flow metrics.
- URN reference: node type and extension type syntax.
- Processor behavior taxonomy: processor behavior categories.
- Transport header policies: inbound header capture and outbound header propagation.
- TLS examples:
test-tls-only.yamlandtest-mtls.yaml. - Proxy support: outbound proxy behavior.
For agent consumption, start from this page, then follow the core-node and
contrib-node catalogs. Per-node READMEs use predictable headings such as
Metadata, Overview, Configuration, Examples, Telemetry, Limits, and
Related Docs.
Validate and Troubleshoot
Configuration parsing is strict. Unknown fields are rejected so typos fail early.
Common validation checks include:
versionmust beotel_dataflow/v1.- Connections must reference existing nodes.
- Connection sources must be receivers or processors.
- Graphs must not contain cycles.
- Referenced output ports must exist when
outputsis declared. - Channel and topic capacities must be non-zero.
groups.systemis reserved for engine-managed pipelines.backend: quivertopics are rejected by the current runtime.- Node
configfields must match the selected node type. - Node types must be registered in the
df_enginebinary. - Node-level
header_captureis receiver-only. - Node-level
header_propagationis exporter-only.
Use --validate-and-exit while editing:
cargo run -- --config path/to/config.yaml --validate-and-exit
If validation fails inside a node config, open that node's README from the core-node catalog or contrib-node catalog. If validation fails in the root model, use the configuration model reference.