Parse K3s Log Metadata
Extract pod and container metadata from kubectl log prefixes and attach the configured namespace for filtering and searching.
Prerequisites and collection scope
Install kubectl on the edge node and configure a kubeconfig with permission to list pods and read pods/log in production. Replace production and app=web-app below with your namespace and workload label. The command follows the matching pods available when it starts, with at most 10 concurrent log streams; it does not discover new pods continuously. For fleet-wide collection across pod churn, use a Kubernetes log collector. Restarting this command can replay log lines; design downstream storage for duplicates.
These are pipeline configuration fragments. Put input, pipeline, and output under config in a job with name and type: pipeline, as shown in the quickstart.
Pipeline
input:
subprocess:
name: kubectl
args:
- logs
- --all-containers=true
- --prefix=true
- --follow
- --namespace=production
- --selector=app=web-app
- --max-log-requests=10
codec: lines
restart_on_exit: true
pipeline:
processors:
# kubectl prefixes lines with [pod/POD_NAME/CONTAINER_NAME].
- mapping: |
root.raw_log = content().string()
root.timestamp = now()
root.namespace = "production"
meta namespace = "production"
let parts = content().string().re_find_all_submatch("^\\[pod/([^/]+)/([^\\]]+)\\] (.*)$")
root.pod = $parts.index(0).index(1)
root.container = $parts.index(0).index(2)
root.message = $parts.index(0).index(3)
# Add context
root.node_id = env("NODE_ID")
root.location = env("LOCATION")
root.cluster = env("CLUSTER_NAME")
output:
aws_s3:
bucket: edge-k3s-logs
path: 'logs/${! env("NODE_ID") }/${! now().ts_format("2006-01-02") }/${! metadata("namespace") }/${! uuid_v4() }.jsonl'
batching:
count: 1000
period: 1m
processors:
- archive:
format: lines
What This Does
- Parses kubectl prefix: Extracts pod and container from
[pod/POD_NAME/CONTAINER_NAME]; namespace comes from the command configuration - Separates message: Stores the actual log message separately from metadata
- Adds location context: Includes node ID, location, and cluster name
- Organizes by namespace: S3 path includes namespace for easy filtering
Example Output
Input (kubectl log line):
[pod/web-app-7d8f9c/app] Request processed in 45ms
Output (structured JSON):
{
"namespace": "production",
"pod": "web-app-7d8f9c",
"container": "app",
"message": "Request processed in 45ms",
"node_id": "edge-site-42",
"location": "chicago",
"cluster": "k3s-chicago",
"timestamp": "2024-11-09T10:30:45Z"
}
Prefix and namespace
The S3 archive combines each batch into JSONL. Its path uses namespace metadata retained from the first record plus a UUID, so multiple batches do not overwrite one another. Keep the namespace metadata aligned with the command.
The first prefix segment is the literal pod, followed by the pod name and container name. It does not contain the namespace. Keep root.namespace aligned with --namespace when copying this example. The mapping expects this prefix on every input line; inspect processor errors if a different command or collector changes the format.
Use Cases
Search by namespace: Query S3 for all logs from production namespace
Filter by pod: Find all logs from a specific pod across time
Container-level debugging: Isolate logs from sidecar containers
Multi-cluster aggregation: Compare logs from same namespace across different edge locations
Next Steps
- Multiple Destinations: Send parsed logs to OpenSearch for real-time search
- Filter by Log Level: Combine with log level filtering
- Best Practices: Learn about efficient log handling