Datadog Agent & Vector
Shipping logs, metrics and traces from the Datadog Agent and Vector's datadog_* sinks to Berserk's Datadog-compatible intake
Berserk's ingest service can serve a Datadog-compatible intake. If you already ship with the Datadog Agent or Vector's datadog_logs and datadog_metrics sinks, you can send logs, metrics and APM traces to Berserk by changing the endpoint and the API key. You don't need to rewrite the pipeline or reshape your events.
Agent 7.84+ behind Vector: traces are lost in Vector
Datadog Agent 7.84 sends traces in a new indexed format, which Vector's datadog_agent source (as of Vector 0.58) cannot read. It accepts the request and forwards nothing. Berserk reads both formats. With Agent 7.84 or later, either send traces from the Agent straight to Berserk rather than through Vector, or set apm_config.features: ["disable-convert-traces"] on the Agent to keep the classic format (Datadog plans to remove that switch), or stay on Agent 7.83.
The intake accepts logs (/api/v2/logs), metrics (/api/v2/series, /api/intake/metrics/v3/series, /api/beta/sketches) and APM traces (/api/v0.2/traces).
For the general ingest contract (tokens, size limits, what an acknowledgment means), see Ingestion. Coming from Datadog's UI? Datadog → Berserk field mapping says where each facet and attribute lands, with query translations.
Endpoint and API Key
- Endpoint. The intake listens on the ingest service's Datadog port,
9552, which your deployment exposes behind its ingest host (e.g.https://<your-ingest-host>). It is a different port from OTLP's4317/4318. Self-hosting? See Enabling the intake. - API key. Use a Berserk ingest token (
ing_…) wherever Datadog expects an API key. It is sent in theDD-API-KEYheader, as a Datadog client does by default. The token decides which table the logs land in. - Key check. Both the Agent and Vector check their key at startup against
/api/v1/validate. Berserk answers200 {"valid":true}for a token it accepts and403for one it does not.
Vector
Point the datadog_logs sink at Berserk. The events need no remapping: the sink sends the same Datadog payload it would send to Datadog.
secret:
berserk:
type: file
path: /etc/vector/berserk.json # {"ingest_token": "ing_…"}
sinks:
berserk:
type: datadog_logs
inputs: [my_logs]
endpoint: https://<your-ingest-host>
default_api_key: SECRET[berserk.ingest_token]
compression: zstd # the default; keep it
request:
concurrency: 32
timeout_secs: 60
batch:
timeout_secs: 1
buffer:
type: disk
max_size: 1073741824 # 1 GiB
when_full: blockWhy These Settings
The key comes from a secret backend. Since Vector 0.57, ${VAR} in a config file is no longer replaced by the environment variable unless Vector runs with --dangerously-allow-env-var-interpolation. Without that flag, default_api_key: ${BERSERK_INGEST_TOKEN} sends the literal text as the key, and Berserk answers 403. A secret backend works on every version that has one: file reads a JSON object, and directory, exec and aws_secrets_manager are the alternatives.
request.concurrency: 32 is the setting that matters most.
- Berserk acknowledges a request only once its batch is durably on object storage, and batches for up to 2 seconds. A request therefore takes about 1–3 seconds, against about 0.1 seconds at Datadog.
- The sink caps each request at 1000 events. That is a hard limit: a larger
batch.max_eventsis a configuration error, not something Vector rounds down. So throughput is roughly concurrency × 1000 events / 2 seconds. - Vector's default adaptive concurrency starts at 1 request in flight and adds at most one per round trip. That means a few hundred events per second until it ramps up, and every restart starts over.
- A fixed value avoids the ramp. Many requests in flight cost Berserk almost nothing, because they land in the same batch. Raise it for a busy aggregator: about 32 per 15k events/s.
batch.timeout_secs: 1. Berserk batches for 2 seconds anyway; the sink's default of 5 seconds only adds latency.
request.timeout_secs: 60 (the default). Never set it below 30. Berserk can hold a request for up to 25 seconds while it retries object storage. A shorter client timeout turns that into a resend.
buffer.type: disk. While Berserk is throttling or briefly unavailable, Vector's buffer holds the data. The default 500-event memory buffer does not survive a restart.
Retries need no configuration. Vector retries 408, 429 and 5xx indefinitely, and drops other 4xx. It does not read Retry-After; it backs off on its own schedule, up to 30 seconds.
End-to-end acknowledgements. If you enable acknowledgements with a datadog_agent source, the Agent waits for Berserk's acknowledgment. Keep batch.timeout_secs low, or the chain can exceed the Agent's 10-second request timeout.
Metrics
The datadog_metrics sink needs the same endpoint and key. Berserk reads both protobuf series formats, v2 and v3. Pin series_api_version: v2 anyway: Vector's next release switches its default to v3, and Vector's v3 encoder has not been tested against Berserk yet.
sinks:
berserk_metrics:
type: datadog_metrics
inputs: [my_metrics]
endpoint: https://<your-ingest-host>
default_api_key: SECRET[berserk.ingest_token]
series_api_version: v2
request:
concurrency: 16Vector sends distributions (histogram metrics) as sketches, which Berserk stores as histograms (below).
Traces
The datadog_traces sink relays what Vector's datadog_agent source receives from your Agents:
sinks:
berserk_traces:
type: datadog_traces
inputs: [agent.traces]
endpoint: https://<your-ingest-host>
default_api_key: SECRET[berserk.ingest_token]
compression: zstd
request:
concurrency: 64
batch:
timeout_secs: 2compression: zstd: the sink's default is no compression.concurrency: the sink batches per host, so an aggregator in front of many Agents sends many small requests. Size it at about number of hosts ÷ batch timeout × 2.5 s.batch.timeout_secs: 2: the default of 10 s can exceed the Agent's own 10 s request timeout when end-to-end acknowledgements are on.- Mind the Agent 7.84 warning at the top of this page.
Datadog Agent
The Agent can send logs to an additional endpoint with its own API key, while the rest of its data keeps going to Datadog. That is the low-risk way to start: dual-ship, compare, then cut over.
logs_enabled: true
logs_config:
force_use_http: true
additional_endpoints:
- api_key: "ing_xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx"
Host: "<your-ingest-host>"
Port: 443
is_reliable: trueforce_use_http: trueis required: the intake speaks the Agent's HTTP logs protocol, not its TCP one.
To dual-ship metrics too, add a top-level additional endpoint. It must be https://:
additional_endpoints:
"https://<your-ingest-host>":
- ing_xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxis_reliable: truemakes the Agent treat Berserk like its primary endpoint: it retries and applies backpressure instead of dropping data when Berserk is slow.
To dual-ship APM traces, add an additional trace endpoint (Agent 6.7+; the environment variable needs 7.19+):
apm_config:
additional_endpoints:
"https://<your-ingest-host>":
- ing_xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxThe Agent sends each log line as-is in message, with status set to its default for the source (usually info). Datadog's intake then parses JSON lines; Berserk does the same (see mapping). So a JSON line's level becomes the severity and its dd.trace_id stays searchable, exactly as in Datadog.
How Metrics Are Mapped
Each Datadog point becomes one row, with the standard metric columns. Tags become attributes (the env/version/service/host promotions above apply); the host resource becomes resource['host.name']. Series v2 and v3 map identically. scope_name is datadog, and scope_version records the API a point arrived on: v2, v3, or beta for sketches.
| Datadog type | Berserk row | Notes |
|---|---|---|
| GAUGE | gauge | |
| COUNT | sum, DELTA | start_time is timestamp − interval when the sender gives an interval |
| RATE | sum, DELTA | The per-second value × interval, so it sums like a count. attributes['datadog.metric_type'] = "rate". DogStatsD counters arrive this way |
| RATE with no interval | gauge | The per-second value, marked the same way |
| Distribution (sketch) | exponential_histogram, DELTA | count, sum, min and max are exact; buckets are re-binned at scale 5 |
Totals of counts and rates. Counts and rates are delta sums, so otel_increase($raw) totals them and otel_rate($raw) rates them, whether or not the sender gave an interval (Vector's counts carry none). Datadog counts can also go down (DogStatsD decrement), so they are stored as non-monotonic; for a count that does, use otel_delta($raw), since otel_increase warns about a negative increment.
Histogram (|h) metrics arrive already aggregated by the Agent, as .avg, .median, .max, .95percentile gauges and a .count rate.
Percentiles from distributions. Use otel_histogram_percentile, which merges the delta histograms in each group:
my_metrics
| where metric_name == "checkout.latency"
| summarize p50 = otel_histogram_percentile($raw, 50), p99 = otel_histogram_percentile($raw, 99) by bin(timestamp, 1m)Re-binning Datadog's buckets onto the histogram's costs some precision: a percentile can be off by up to about 3.3%, where Datadog's own sketch is within about 0.8%.
How Traces Are Mapped
Each Datadog span becomes one row, with the standard span columns. Both trace payload formats (the Agent's classic one and 7.84's indexed one) map identically.
| Datadog | Berserk | Notes |
|---|---|---|
resource (e.g. GET /pay) | span_name | Falls back to the operation name if empty |
name, the operation (e.g. http.request) | attributes['dd.operation_name'] | |
type (web, sql, …) | attributes['dd.span_type'] | |
service | resource['service.name'] | Per span |
env, version | resource['deployment.environment.name'], resource['service.version'] | |
| trace id | trace_id (128-bit hex) | High bits from the indexed payload or _dd.p.tid; zero for 64-bit traces |
| trace id, low 64 bits | attributes['dd.trace_id'] (decimal) | The key that matches logs |
| span id, parent id | span_id, parent_span_id (hex), and attributes['dd.span_id'] (decimal) | |
start, duration | start_time, end_time, duration | |
error | status_code = ERROR, status_description from error.message | Otherwise UNSET, unless the span recorded an OTel status (below) |
span.kind | span_kind | Missing → INTERNAL, as the Agent itself converts it |
meta, metrics | attributes | Strings and numbers, verbatim |
| span links and events | links, events | Native ones, and the legacy _dd.span_links / events JSON |
| tracer language, tracer version | resource['telemetry.sdk.language'], resource['telemetry.sdk.version']; telemetry.sdk.name = "datadog" | |
| runtime id | resource['service.instance.id'] | |
component (the integration, e.g. flask) | scope_name, with scope_version = tracer version | Spans without one have scope_name = "datadog" |
otel.scope.name, otel.scope.version (or the older otel.library.*) | scope_name, scope_version | Written by the Agent for spans it received as OTLP; they outrank component, which then stays in attributes |
otel.status_code, otel.status_description | status_code, status_description | Also from OTLP spans: Ok gives OK, Error gives ERROR even without the error flag |
Joining logs and traces. A log's trace_id and span_id lead to its span; Correlating logs and traces lists the tracer versions that need a setting for that.
APM stats (/api/v0.2/stats) are acknowledged and discarded: they can be computed from the spans.
How Fields Are Mapped
Each Datadog log entry becomes one row, with the standard log columns:
| Datadog field | Berserk column | Notes |
|---|---|---|
message | body | If it is a JSON object, its fields are expanded (below) |
status | severity_text, severity_number | trace, debug, info/ok, notice, warn, error, critical, alert, emergency, or syslog numbers 0–7. Unknown values keep their text with number 0 |
timestamp | timestamp | Epoch milliseconds or RFC 3339. If absent, the row uses observed_time |
hostname | resource['host.name'] | |
service | resource['service.name'] | |
ddtags | attributes | env:prod,team:a,team:b gives attributes.env = "prod" and attributes.team = ["a","b"]. env and version are also copied to resource['deployment.environment.name'] and resource['service.version'] |
ddsource | attributes['ddsource'] | |
dd.trace_id, dd.span_id | attributes['dd.trace_id'], attributes['dd.span_id'] | As decimal strings. trace_id/span_id are set too: a 128-bit dd.trace_id as is, a 64-bit one joined with the high 64 bits in _dd.p.tid (16 hex digits) if present, else with the high 64 bits zero. See Correlating logs and traces |
| anything else | attributes | Nested objects flatten to dotted keys (http.status_code) |
JSON log lines. When message is a JSON object, its fields become attributes, and Datadog's reserved names are remapped the way Datadog's own pipelines do it:
| Target | First key found among |
|---|---|
| body | message, msg, log |
| severity | status, severity, level, syslog.severity |
| timestamp | timestamp, date, _timestamp, Timestamp, eventTime, published_date |
| trace id | dd.trace_id, contextMap.dd.trace_id |
A severity key only counts if its value is a recognised severity. In an access log, "status": "410" is an HTTP status, so the envelope's status stays and 410 remains an attribute.
Correlating Logs and Traces
Berserk's trace views, the jump from a log to its trace, and trace-find all use the OpenTelemetry columns trace_id, span_id and parent_span_id. Every Datadog span and every log with a dd.trace_id has them. They lead from a log to its spans when the log's trace_id matches the spans' trace_id, which depends on the tracer version.
Datadog tracers began generating 128-bit trace ids by default in November 2023. From the versions below, they also write 128-bit ids into logs by default, as 32 hex digits. They switch to decimal only for a trace that really is 64-bit. Berserk reads a decimal id as a 64-bit trace, with the high 64 bits zero, which is how the spans of such a trace are stored too.
| Tracer | Works from | Released | Versions that need a setting |
|---|---|---|---|
Java (dd-trace-java) | 1.48.0 | 2025-04-09 | 1.24 – 1.47 |
Python (ddtrace) | 2.12.0 | 2024-09-09 | 2.3 – 2.11 |
Go (dd-trace-go) | v2.0.0 | 2025-06-05 | v1.58 and later v1 releases |
Node.js (dd-trace) | 5.38.0 | 2025-02-24 | 5.0 – 5.37, and the 4.x and 3.x lines from 4.19 and 3.40 |
.NET (Datadog.Trace) | 3.13.0 | 2025-03-24 | 2.42 – 3.12 |
Ruby (datadog, formerly ddtrace) | 2.13.0 | 2025-04-02 | 1.17 – 2.12 |
PHP (dd-trace-php) | 1.8.0 | 2025-04-02 | 0.94 – 1.7 |
NGINX (nginx-datadog), $datadog_trace_id in the log format | 1.6.0 | 2025-04-07 | earlier releases write the low 64 bits in decimal |
- From the "Works from" version on, nothing needs configuring: 128-bit log injection is on by default (Python 2.12 removed the setting).
- Versions in the last column generate 128-bit trace ids but, by default, write only the low 64 bits into logs. Their logs get a
trace_idthat matches no span, so a log's trace link finds nothing. SetDD_TRACE_128_BIT_TRACEID_LOGGING_ENABLED=trueon those services, or upgrade. With it, Python 2.3 – 2.11 writes the 128-bit id as one long decimal number, which Berserk reads too. - Tracers older than the last column generate 64-bit trace ids, so logs and spans agree, and linking works.
- The Datadog Agent and Vector versions do not matter here. Both pass log trace ids through unchanged. For traces sent through Vector, see the warning at the top of this page.
dd.trace_id and dd.span_id work for every version, including the ones in the last column. They hold the low 64 bits in decimal on spans and logs alike, as Datadog shows them:
my_table
| where attributes['dd.trace_id'] == "9532127138774266268"Responses
The intake follows Datadog clients' retry rules, which differ from OTLP's in two places:
| Response | Meaning |
|---|---|
202 {} | The batch is durably stored, as with any ingest acknowledgment |
503 + Retry-After | Backpressure. 503 rather than 429, because the Agent drops some payloads on 429 instead of retrying |
| 403 | Missing or rejected token. 403 rather than 401, because the Agent keeps retrying a 401 instead of reporting a bad key |
| 400 / 413 | Malformed payload, or over the size limit. Permanent |
| 404 | A Datadog endpoint Berserk does not serve, such as /api/v1/series (JSON) or /api/v1/distribution_points |
Service checks (/api/v1/check_run), host metadata (/intake/, /api/v1/metadata) and APM stats (/api/v0.2/stats) are acknowledged and discarded.
Enabling the Intake (Self-Hosted)
The intake is off by default. In the Helm chart, enable the listener and route it:
ingest:
config:
datadogEnabled: true
datadogIngress:
enabled: true
className: <your-ingress-class>
hosts:
- host: dd-intake.example.com
paths:
- path: /
pathType: Prefix
portName: datadog
tls:
- hosts: [dd-intake.example.com]
secretName: dd-intake-tlsA host of its own is simplest: Datadog clients append many paths to their endpoint. To share a host with other services instead, route these Exact paths to portName: datadog:
/api/v2/logs, /api/v2/series, /api/intake/metrics/v3/series, /api/beta/sketches, /api/v0.2/traces, /api/v1/validate, /_health, /api/v1/check_run, /intake/, /api/v1/metadata, /api/v0.2/stats
Tested With
The intake is tested against real payloads captured from these senders, and Vector runs end to end in CI:
| Sender | Version | Logs | Metrics | Traces |
|---|---|---|---|---|
| Datadog Agent | 7.84.1 | ✓ | ✓ (series v2 and v3, sketches) | ✓ (indexed format) |
| Datadog Agent | 7.83.3 | ✓ (classic format, direct and through Vector) | ||
| Vector | 0.58.0 | ✓ datadog_logs | ✓ datadog_metrics | ✓ datadog_traces (from Agent 7.83.3) |
Version Compatibility
Berserk accepts series in API v2 and v3 (both protobuf), but not v1 (JSON). Most version boundaries that matter are about the series format, and about traces through Vector:
| Sender | Versions | What to do |
|---|---|---|
| Datadog Agent | < 7.41 | Sends v1 (JSON) series by default. Set use_v2_api.series: true (7.33+), or upgrade |
| Datadog Agent | 7.81+ | Sends v3 series: 7.81.0 and 7.81.1 to every endpoint, 7.81.2+ only to Datadog's own hostnames (Berserk gets v2). Berserk reads both |
| Datadog Agent | any | Set logs_config.force_use_http: true. Otherwise a failed HTTP check at startup falls back to TCP, which Berserk does not serve |
| Datadog Agent | 7.37+ | Logs additional_endpoints are reliable: a slow Berserk applies backpressure to the Datadog leg too |
| Datadog Agent | 7.84+ behind Vector | Traces are lost inside Vector; see the warning at the top |
| Vector | < 0.55 | datadog_metrics sends v1 series only; upgrade |
| Vector | 0.55 – 0.58 | Works as-is (v2 is the default; these versions have no v3) |
| Vector | next release (expected 0.59) | datadog_metrics defaults to v3 series. Berserk reads v3, but Vector's encoder is untested: set series_api_version: v2 |
Troubleshooting
| Symptom | Cause |
|---|---|
Vector: Healthcheck failed … 403 Forbidden | The default_api_key is not a valid ingest token. On Vector 0.57+ a ${VAR} key is sent literally; use a secret backend. With --require-healthy, Vector exits; without it, it keeps retrying the data |
| Agent: API key reported invalid at startup | Same: the key in additional_endpoints must be an ingest token |
| Throughput stuck at a few hundred events/s | Vector's adaptive concurrency against Berserk's ~2 s acknowledgment. Set a fixed request.concurrency |
| A log's trace link finds no trace, though its spans are there | The tracer generates 128-bit ids but logs only the low 64 bits. See Correlating logs and traces for the versions and the setting that fixes it; dd.trace_id still matches |
| Traces sent through Vector never arrive, with no error | Agent 7.84+ behind Vector: see the warning at the top of this page. Send traces from the Agent directly |
| 404 in the shipper's log | An endpoint Berserk does not serve; see Responses |
400 on /api/v2/series with a JSON body | A sender using the JSON series format (API v1). Berserk reads protobuf (series API v2 or v3); see Version compatibility |
404 on /api/intake/metrics/v3/series | An ingest older than series v3 support; upgrade it, or pin the sender to v2 |
| Data accepted but not in the table you query | The token is bound to a different table. See Verifying Ingestion |
Then confirm end to end:
bzrk search "<your table> | where scope_name == 'datadog' | take 10" --since "5m ago"