Berserk Docs

Datadog Agent & Vector

Shipping logs, metrics and traces from the Datadog Agent and Vector's datadog_* sinks to Berserk's Datadog-compatible intake

Berserk's ingest service can serve a Datadog-compatible intake. If you already ship with the Datadog Agent or Vector's datadog_logs and datadog_metrics sinks, you can send logs, metrics and APM traces to Berserk by changing the endpoint and the API key. You don't need to rewrite the pipeline or reshape your events.

Agent 7.84+ behind Vector: traces are lost in Vector

Datadog Agent 7.84 sends traces in a new indexed format, which Vector's datadog_agent source (as of Vector 0.58) cannot read. It accepts the request and forwards nothing. Berserk reads both formats. With Agent 7.84 or later, either send traces from the Agent straight to Berserk rather than through Vector, or set apm_config.features: ["disable-convert-traces"] on the Agent to keep the classic format (Datadog plans to remove that switch), or stay on Agent 7.83.

The intake accepts logs (/api/v2/logs), metrics (/api/v2/series, /api/intake/metrics/v3/series, /api/beta/sketches) and APM traces (/api/v0.2/traces).

For the general ingest contract (tokens, size limits, what an acknowledgment means), see Ingestion. Coming from Datadog's UI? Datadog → Berserk field mapping says where each facet and attribute lands, with query translations.

Endpoint and API Key

  • Endpoint. The intake listens on the ingest service's Datadog port, 9552, which your deployment exposes behind its ingest host (e.g. https://<your-ingest-host>). It is a different port from OTLP's 4317/4318. Self-hosting? See Enabling the intake.
  • API key. Use a Berserk ingest token (ing_…) wherever Datadog expects an API key. It is sent in the DD-API-KEY header, as a Datadog client does by default. The token decides which table the logs land in.
  • Key check. Both the Agent and Vector check their key at startup against /api/v1/validate. Berserk answers 200 {"valid":true} for a token it accepts and 403 for one it does not.

Vector

Point the datadog_logs sink at Berserk. The events need no remapping: the sink sends the same Datadog payload it would send to Datadog.

secret:
  berserk:
    type: file
    path: /etc/vector/berserk.json            # {"ingest_token": "ing_…"}
sinks:
  berserk:
    type: datadog_logs
    inputs: [my_logs]
    endpoint: https://<your-ingest-host>
    default_api_key: SECRET[berserk.ingest_token]
    compression: zstd                         # the default; keep it
    request:
      concurrency: 32
      timeout_secs: 60
    batch:
      timeout_secs: 1
    buffer:
      type: disk
      max_size: 1073741824                    # 1 GiB
      when_full: block

Why These Settings

The key comes from a secret backend. Since Vector 0.57, ${VAR} in a config file is no longer replaced by the environment variable unless Vector runs with --dangerously-allow-env-var-interpolation. Without that flag, default_api_key: ${BERSERK_INGEST_TOKEN} sends the literal text as the key, and Berserk answers 403. A secret backend works on every version that has one: file reads a JSON object, and directory, exec and aws_secrets_manager are the alternatives.

request.concurrency: 32 is the setting that matters most.

  • Berserk acknowledges a request only once its batch is durably on object storage, and batches for up to 2 seconds. A request therefore takes about 1–3 seconds, against about 0.1 seconds at Datadog.
  • The sink caps each request at 1000 events. That is a hard limit: a larger batch.max_events is a configuration error, not something Vector rounds down. So throughput is roughly concurrency × 1000 events / 2 seconds.
  • Vector's default adaptive concurrency starts at 1 request in flight and adds at most one per round trip. That means a few hundred events per second until it ramps up, and every restart starts over.
  • A fixed value avoids the ramp. Many requests in flight cost Berserk almost nothing, because they land in the same batch. Raise it for a busy aggregator: about 32 per 15k events/s.

batch.timeout_secs: 1. Berserk batches for 2 seconds anyway; the sink's default of 5 seconds only adds latency.

request.timeout_secs: 60 (the default). Never set it below 30. Berserk can hold a request for up to 25 seconds while it retries object storage. A shorter client timeout turns that into a resend.

buffer.type: disk. While Berserk is throttling or briefly unavailable, Vector's buffer holds the data. The default 500-event memory buffer does not survive a restart.

Retries need no configuration. Vector retries 408, 429 and 5xx indefinitely, and drops other 4xx. It does not read Retry-After; it backs off on its own schedule, up to 30 seconds.

End-to-end acknowledgements. If you enable acknowledgements with a datadog_agent source, the Agent waits for Berserk's acknowledgment. Keep batch.timeout_secs low, or the chain can exceed the Agent's 10-second request timeout.

Metrics

The datadog_metrics sink needs the same endpoint and key. Berserk reads both protobuf series formats, v2 and v3. Pin series_api_version: v2 anyway: Vector's next release switches its default to v3, and Vector's v3 encoder has not been tested against Berserk yet.

sinks:
  berserk_metrics:
    type: datadog_metrics
    inputs: [my_metrics]
    endpoint: https://<your-ingest-host>
    default_api_key: SECRET[berserk.ingest_token]
    series_api_version: v2
    request:
      concurrency: 16

Vector sends distributions (histogram metrics) as sketches, which Berserk stores as histograms (below).

Traces

The datadog_traces sink relays what Vector's datadog_agent source receives from your Agents:

sinks:
  berserk_traces:
    type: datadog_traces
    inputs: [agent.traces]
    endpoint: https://<your-ingest-host>
    default_api_key: SECRET[berserk.ingest_token]
    compression: zstd
    request:
      concurrency: 64
    batch:
      timeout_secs: 2
  • compression: zstd: the sink's default is no compression.
  • concurrency: the sink batches per host, so an aggregator in front of many Agents sends many small requests. Size it at about number of hosts ÷ batch timeout × 2.5 s.
  • batch.timeout_secs: 2: the default of 10 s can exceed the Agent's own 10 s request timeout when end-to-end acknowledgements are on.
  • Mind the Agent 7.84 warning at the top of this page.

Datadog Agent

The Agent can send logs to an additional endpoint with its own API key, while the rest of its data keeps going to Datadog. That is the low-risk way to start: dual-ship, compare, then cut over.

logs_enabled: true
logs_config:
  force_use_http: true
  additional_endpoints:
    - api_key: "ing_xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx"
      Host: "<your-ingest-host>"
      Port: 443
      is_reliable: true
  • force_use_http: true is required: the intake speaks the Agent's HTTP logs protocol, not its TCP one.

To dual-ship metrics too, add a top-level additional endpoint. It must be https://:

additional_endpoints:
  "https://<your-ingest-host>":
    - ing_xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx
  • is_reliable: true makes the Agent treat Berserk like its primary endpoint: it retries and applies backpressure instead of dropping data when Berserk is slow.

To dual-ship APM traces, add an additional trace endpoint (Agent 6.7+; the environment variable needs 7.19+):

apm_config:
  additional_endpoints:
    "https://<your-ingest-host>":
      - ing_xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx

The Agent sends each log line as-is in message, with status set to its default for the source (usually info). Datadog's intake then parses JSON lines; Berserk does the same (see mapping). So a JSON line's level becomes the severity and its dd.trace_id stays searchable, exactly as in Datadog.

How Metrics Are Mapped

Each Datadog point becomes one row, with the standard metric columns. Tags become attributes (the env/version/service/host promotions above apply); the host resource becomes resource['host.name']. Series v2 and v3 map identically. scope_name is datadog, and scope_version records the API a point arrived on: v2, v3, or beta for sketches.

Datadog typeBerserk rowNotes
GAUGEgauge
COUNTsum, DELTAstart_time is timestamp − interval when the sender gives an interval
RATEsum, DELTAThe per-second value × interval, so it sums like a count. attributes['datadog.metric_type'] = "rate". DogStatsD counters arrive this way
RATE with no intervalgaugeThe per-second value, marked the same way
Distribution (sketch)exponential_histogram, DELTAcount, sum, min and max are exact; buckets are re-binned at scale 5

Totals of counts and rates. Counts and rates are delta sums, so otel_increase($raw) totals them and otel_rate($raw) rates them, whether or not the sender gave an interval (Vector's counts carry none). Datadog counts can also go down (DogStatsD decrement), so they are stored as non-monotonic; for a count that does, use otel_delta($raw), since otel_increase warns about a negative increment.

Histogram (|h) metrics arrive already aggregated by the Agent, as .avg, .median, .max, .95percentile gauges and a .count rate.

Percentiles from distributions. Use otel_histogram_percentile, which merges the delta histograms in each group:

my_metrics
| where metric_name == "checkout.latency"
| summarize p50 = otel_histogram_percentile($raw, 50), p99 = otel_histogram_percentile($raw, 99) by bin(timestamp, 1m)

Re-binning Datadog's buckets onto the histogram's costs some precision: a percentile can be off by up to about 3.3%, where Datadog's own sketch is within about 0.8%.

How Traces Are Mapped

Each Datadog span becomes one row, with the standard span columns. Both trace payload formats (the Agent's classic one and 7.84's indexed one) map identically.

DatadogBerserkNotes
resource (e.g. GET /pay)span_nameFalls back to the operation name if empty
name, the operation (e.g. http.request)attributes['dd.operation_name']
type (web, sql, …)attributes['dd.span_type']
serviceresource['service.name']Per span
env, versionresource['deployment.environment.name'], resource['service.version']
trace idtrace_id (128-bit hex)High bits from the indexed payload or _dd.p.tid; zero for 64-bit traces
trace id, low 64 bitsattributes['dd.trace_id'] (decimal)The key that matches logs
span id, parent idspan_id, parent_span_id (hex), and attributes['dd.span_id'] (decimal)
start, durationstart_time, end_time, duration
errorstatus_code = ERROR, status_description from error.messageOtherwise UNSET, unless the span recorded an OTel status (below)
span.kindspan_kindMissing → INTERNAL, as the Agent itself converts it
meta, metricsattributesStrings and numbers, verbatim
span links and eventslinks, eventsNative ones, and the legacy _dd.span_links / events JSON
tracer language, tracer versionresource['telemetry.sdk.language'], resource['telemetry.sdk.version']; telemetry.sdk.name = "datadog"
runtime idresource['service.instance.id']
component (the integration, e.g. flask)scope_name, with scope_version = tracer versionSpans without one have scope_name = "datadog"
otel.scope.name, otel.scope.version (or the older otel.library.*)scope_name, scope_versionWritten by the Agent for spans it received as OTLP; they outrank component, which then stays in attributes
otel.status_code, otel.status_descriptionstatus_code, status_descriptionAlso from OTLP spans: Ok gives OK, Error gives ERROR even without the error flag

Joining logs and traces. A log's trace_id and span_id lead to its span; Correlating logs and traces lists the tracer versions that need a setting for that.

APM stats (/api/v0.2/stats) are acknowledged and discarded: they can be computed from the spans.

How Fields Are Mapped

Each Datadog log entry becomes one row, with the standard log columns:

Datadog fieldBerserk columnNotes
messagebodyIf it is a JSON object, its fields are expanded (below)
statusseverity_text, severity_numbertrace, debug, info/ok, notice, warn, error, critical, alert, emergency, or syslog numbers 0–7. Unknown values keep their text with number 0
timestamptimestampEpoch milliseconds or RFC 3339. If absent, the row uses observed_time
hostnameresource['host.name']
serviceresource['service.name']
ddtagsattributesenv:prod,team:a,team:b gives attributes.env = "prod" and attributes.team = ["a","b"]. env and version are also copied to resource['deployment.environment.name'] and resource['service.version']
ddsourceattributes['ddsource']
dd.trace_id, dd.span_idattributes['dd.trace_id'], attributes['dd.span_id']As decimal strings. trace_id/span_id are set too: a 128-bit dd.trace_id as is, a 64-bit one joined with the high 64 bits in _dd.p.tid (16 hex digits) if present, else with the high 64 bits zero. See Correlating logs and traces
anything elseattributesNested objects flatten to dotted keys (http.status_code)

JSON log lines. When message is a JSON object, its fields become attributes, and Datadog's reserved names are remapped the way Datadog's own pipelines do it:

TargetFirst key found among
bodymessage, msg, log
severitystatus, severity, level, syslog.severity
timestamptimestamp, date, _timestamp, Timestamp, eventTime, published_date
trace iddd.trace_id, contextMap.dd.trace_id

A severity key only counts if its value is a recognised severity. In an access log, "status": "410" is an HTTP status, so the envelope's status stays and 410 remains an attribute.

Correlating Logs and Traces

Berserk's trace views, the jump from a log to its trace, and trace-find all use the OpenTelemetry columns trace_id, span_id and parent_span_id. Every Datadog span and every log with a dd.trace_id has them. They lead from a log to its spans when the log's trace_id matches the spans' trace_id, which depends on the tracer version.

Datadog tracers began generating 128-bit trace ids by default in November 2023. From the versions below, they also write 128-bit ids into logs by default, as 32 hex digits. They switch to decimal only for a trace that really is 64-bit. Berserk reads a decimal id as a 64-bit trace, with the high 64 bits zero, which is how the spans of such a trace are stored too.

TracerWorks fromReleasedVersions that need a setting
Java (dd-trace-java)1.48.02025-04-091.24 – 1.47
Python (ddtrace)2.12.02024-09-092.3 – 2.11
Go (dd-trace-go)v2.0.02025-06-05v1.58 and later v1 releases
Node.js (dd-trace)5.38.02025-02-245.0 – 5.37, and the 4.x and 3.x lines from 4.19 and 3.40
.NET (Datadog.Trace)3.13.02025-03-242.42 – 3.12
Ruby (datadog, formerly ddtrace)2.13.02025-04-021.17 – 2.12
PHP (dd-trace-php)1.8.02025-04-020.94 – 1.7
NGINX (nginx-datadog), $datadog_trace_id in the log format1.6.02025-04-07earlier releases write the low 64 bits in decimal
  • From the "Works from" version on, nothing needs configuring: 128-bit log injection is on by default (Python 2.12 removed the setting).
  • Versions in the last column generate 128-bit trace ids but, by default, write only the low 64 bits into logs. Their logs get a trace_id that matches no span, so a log's trace link finds nothing. Set DD_TRACE_128_BIT_TRACEID_LOGGING_ENABLED=true on those services, or upgrade. With it, Python 2.3 – 2.11 writes the 128-bit id as one long decimal number, which Berserk reads too.
  • Tracers older than the last column generate 64-bit trace ids, so logs and spans agree, and linking works.
  • The Datadog Agent and Vector versions do not matter here. Both pass log trace ids through unchanged. For traces sent through Vector, see the warning at the top of this page.

dd.trace_id and dd.span_id work for every version, including the ones in the last column. They hold the low 64 bits in decimal on spans and logs alike, as Datadog shows them:

my_table
| where attributes['dd.trace_id'] == "9532127138774266268"

Responses

The intake follows Datadog clients' retry rules, which differ from OTLP's in two places:

ResponseMeaning
202 {}The batch is durably stored, as with any ingest acknowledgment
503 + Retry-AfterBackpressure. 503 rather than 429, because the Agent drops some payloads on 429 instead of retrying
403Missing or rejected token. 403 rather than 401, because the Agent keeps retrying a 401 instead of reporting a bad key
400 / 413Malformed payload, or over the size limit. Permanent
404A Datadog endpoint Berserk does not serve, such as /api/v1/series (JSON) or /api/v1/distribution_points

Service checks (/api/v1/check_run), host metadata (/intake/, /api/v1/metadata) and APM stats (/api/v0.2/stats) are acknowledged and discarded.

Enabling the Intake (Self-Hosted)

The intake is off by default. In the Helm chart, enable the listener and route it:

ingest:
  config:
    datadogEnabled: true
  datadogIngress:
    enabled: true
    className: <your-ingress-class>
    hosts:
      - host: dd-intake.example.com
        paths:
          - path: /
            pathType: Prefix
            portName: datadog
    tls:
      - hosts: [dd-intake.example.com]
        secretName: dd-intake-tls

A host of its own is simplest: Datadog clients append many paths to their endpoint. To share a host with other services instead, route these Exact paths to portName: datadog:

/api/v2/logs, /api/v2/series, /api/intake/metrics/v3/series, /api/beta/sketches, /api/v0.2/traces, /api/v1/validate, /_health, /api/v1/check_run, /intake/, /api/v1/metadata, /api/v0.2/stats

Tested With

The intake is tested against real payloads captured from these senders, and Vector runs end to end in CI:

SenderVersionLogsMetricsTraces
Datadog Agent7.84.1✓✓ (series v2 and v3, sketches)✓ (indexed format)
Datadog Agent7.83.3✓ (classic format, direct and through Vector)
Vector0.58.0✓ datadog_logs✓ datadog_metrics✓ datadog_traces (from Agent 7.83.3)

Version Compatibility

Berserk accepts series in API v2 and v3 (both protobuf), but not v1 (JSON). Most version boundaries that matter are about the series format, and about traces through Vector:

SenderVersionsWhat to do
Datadog Agent< 7.41Sends v1 (JSON) series by default. Set use_v2_api.series: true (7.33+), or upgrade
Datadog Agent7.81+Sends v3 series: 7.81.0 and 7.81.1 to every endpoint, 7.81.2+ only to Datadog's own hostnames (Berserk gets v2). Berserk reads both
Datadog AgentanySet logs_config.force_use_http: true. Otherwise a failed HTTP check at startup falls back to TCP, which Berserk does not serve
Datadog Agent7.37+Logs additional_endpoints are reliable: a slow Berserk applies backpressure to the Datadog leg too
Datadog Agent7.84+ behind VectorTraces are lost inside Vector; see the warning at the top
Vector< 0.55datadog_metrics sends v1 series only; upgrade
Vector0.55 – 0.58Works as-is (v2 is the default; these versions have no v3)
Vectornext release (expected 0.59)datadog_metrics defaults to v3 series. Berserk reads v3, but Vector's encoder is untested: set series_api_version: v2

Troubleshooting

SymptomCause
Vector: Healthcheck failed … 403 ForbiddenThe default_api_key is not a valid ingest token. On Vector 0.57+ a ${VAR} key is sent literally; use a secret backend. With --require-healthy, Vector exits; without it, it keeps retrying the data
Agent: API key reported invalid at startupSame: the key in additional_endpoints must be an ingest token
Throughput stuck at a few hundred events/sVector's adaptive concurrency against Berserk's ~2 s acknowledgment. Set a fixed request.concurrency
A log's trace link finds no trace, though its spans are thereThe tracer generates 128-bit ids but logs only the low 64 bits. See Correlating logs and traces for the versions and the setting that fixes it; dd.trace_id still matches
Traces sent through Vector never arrive, with no errorAgent 7.84+ behind Vector: see the warning at the top of this page. Send traces from the Agent directly
404 in the shipper's logAn endpoint Berserk does not serve; see Responses
400 on /api/v2/series with a JSON bodyA sender using the JSON series format (API v1). Berserk reads protobuf (series API v2 or v3); see Version compatibility
404 on /api/intake/metrics/v3/seriesAn ingest older than series v3 support; upgrade it, or pin the sender to v2
Data accepted but not in the table you queryThe token is bound to a different table. See Verifying Ingestion

Then confirm end to end:

bzrk search "<your table> | where scope_name == 'datadog' | take 10" --since "5m ago"

On this page