Skip to main content

Parse K3s Log Metadata

Extract pod and container metadata from kubectl log prefixes and attach the configured namespace for filtering and searching.

Prerequisites and collection scope​

Install kubectl on the edge node and configure a kubeconfig with permission to list pods and read pods/log in production. Replace production and app=web-app below with your namespace and workload label. The command follows the matching pods available when it starts, with at most 10 concurrent log streams; it does not discover new pods continuously. For fleet-wide collection across pod churn, use a Kubernetes log collector. Restarting this command can replay log lines; design downstream storage for duplicates.

These are pipeline configuration fragments. Put input, pipeline, and output under config in a job with name and type: pipeline, as shown in the quickstart.

Pipeline​

input:
subprocess:
name: kubectl
args:
- logs
- --all-containers=true
- --prefix=true
- --follow
- --namespace=production
- --selector=app=web-app
- --max-log-requests=10
codec: lines
restart_on_exit: true

pipeline:
processors:
# kubectl prefixes lines with [pod/POD_NAME/CONTAINER_NAME].
- mapping: |
root.raw_log = content().string()
root.timestamp = now()
root.namespace = "production"
meta namespace = "production"

let parts = content().string().re_find_all_submatch("^\\[pod/([^/]+)/([^\\]]+)\\] (.*)$")
root.pod = $parts.index(0).index(1)
root.container = $parts.index(0).index(2)
root.message = $parts.index(0).index(3)

# Add context
root.node_id = env("NODE_ID")
root.location = env("LOCATION")
root.cluster = env("CLUSTER_NAME")

output:
aws_s3:
bucket: edge-k3s-logs
path: 'logs/${! env("NODE_ID") }/${! now().ts_format("2006-01-02") }/${! metadata("namespace") }/${! uuid_v4() }.jsonl'
batching:
count: 1000
period: 1m
processors:
- archive:
format: lines

What This Does​

  • Parses kubectl prefix: Extracts pod and container from [pod/POD_NAME/CONTAINER_NAME]; namespace comes from the command configuration
  • Separates message: Stores the actual log message separately from metadata
  • Adds location context: Includes node ID, location, and cluster name
  • Organizes by namespace: S3 path includes namespace for easy filtering

Example Output​

Input (kubectl log line):

[pod/web-app-7d8f9c/app] Request processed in 45ms

Output (structured JSON):

{
"namespace": "production",
"pod": "web-app-7d8f9c",
"container": "app",
"message": "Request processed in 45ms",
"node_id": "edge-site-42",
"location": "chicago",
"cluster": "k3s-chicago",
"timestamp": "2024-11-09T10:30:45Z"
}

Prefix and namespace​

The S3 archive combines each batch into JSONL. Its path uses namespace metadata retained from the first record plus a UUID, so multiple batches do not overwrite one another. Keep the namespace metadata aligned with the command.

The first prefix segment is the literal pod, followed by the pod name and container name. It does not contain the namespace. Keep root.namespace aligned with --namespace when copying this example. The mapping expects this prefix on every input line; inspect processor errors if a different command or collector changes the format.

Use Cases​

Search by namespace: Query S3 for all logs from production namespace

Filter by pod: Find all logs from a specific pod across time

Container-level debugging: Isolate logs from sidecar containers

Multi-cluster aggregation: Compare logs from same namespace across different edge locations

Next Steps​