Skip to main content

Basic K3s Log Collection

Stream logs from selected pods in one namespace to S3 with batching.

Prerequisites and collection scope​

Install kubectl on the edge node and configure a kubeconfig with permission to list pods and read pods/log in production. Replace production and app=web-app below with your namespace and workload label. The command follows the matching pods available when it starts, with at most 10 concurrent log streams; it does not discover new pods continuously. For fleet-wide collection across pod churn, use a Kubernetes log collector. Restarting this command can replay log lines; design downstream storage for duplicates.

These are pipeline configuration fragments. Put input, pipeline, and output under config in a job with name and type: pipeline, as shown in the quickstart.

Pipeline​

input:
subprocess:
name: kubectl
args:
- logs
- --all-containers=true
- --prefix=true
- --follow
- --tail=-1
- --namespace=production
- --selector=app=web-app
- --max-log-requests=10
codec: lines
restart_on_exit: true

pipeline:
processors:
- mapping: |
root.raw_log = content().string()
root.timestamp = now()
root.node_id = env("NODE_ID")
root.cluster = env("CLUSTER_NAME").or("k3s-edge")

output:
aws_s3:
bucket: edge-k3s-logs
path: 'logs/${! env("NODE_ID") }/${! timestamp_unix() }-${! uuid_v4() }.jsonl'
batching:
count: 1000
period: 1m
processors:
- archive:
format: lines

What This Does​

  • Follows logs from all containers in the selected pods using kubectl logs --follow
  • Adds metadata: node identifier, cluster name, and timestamp to each log entry
  • Batches logs: Collects 1000 logs or waits 1 minute before writing
  • Writes to S3: Organizes logs by node and timestamp for easy retrieval
  • Auto-restarts: If kubectl process exits, it automatically restarts

Key Configuration​

--all-containers=true: Includes logs from all containers in each pod

--prefix=true: Adds a pod and container prefix to each log line

--follow: Continuously streams new logs (like tail -f)

restart_on_exit: true: Ensures log collection continues even if kubectl crashes

Batching: Reduces S3 API calls by writing 1000 logs at once instead of individual files

Environment Variables​

Set these environment variables where Expanso runs:

  • NODE_ID: Unique identifier for this edge node (e.g., edge-site-42)
  • CLUSTER_NAME: Name of the K3s cluster (optional, defaults to k3s-edge)

Next Steps​