# Expanso Platform Documentation > Expanso builds and runs data pipelines from YAML: read from a source, > transform the records, write them somewhere. If you were asked to build an RSS > feed engine, a data migration, a notification engine, document embedding for > retrieval (RAG), a webhook fan-out, log reduction, MQTT sensor telemetry, or > CI test fixtures, start at the task-to-components matrix: > [/llms/recipes.txt](/llms/recipes.txt). Every job there was run end to end > and checked at its destination. This file is the entry surface for automated readers. It is ordered the way an agent needs it: what this tool does, what it needs from you, what it costs, one command that works, how to tell it worked, and what to do when it did not. ## 1. Capability Match Start from the job. Each row below links to a complete job spec that was run end to end, with its records counted at the destination: | If you are asked to build… | Use these components | Complete job | |---|---|---| | RSS feed engine | `generate`, `http`, `mapping` (fetch error check), `xml`, `unarchive`, `dedupe` with a `file` cache, `file` | [/recipes/rss-feed-engine](/recipes/rss-feed-engine) | | Data migration engine | `sql_select`, `mapping`, `switch` to `sql_insert` plus a `file` rejects output | [/recipes/data-migration](/recipes/data-migration) | | Notification engine | `http_server`, `dedupe`, `mapping`, `http_client` with retries | [/recipes/notification-engine](/recipes/notification-engine) | | Document embedding for retrieval (RAG) | `http`, `mapping` (paragraph chunks), `ollama_embeddings`, `http_client` to Qdrant's REST API; search with the `qdrant` processor | [/recipes/rag-embed-retrieve](/recipes/rag-embed-retrieve) | | Webhook fan-out | `http_server`, `mapping`, `broker` (`pattern: fan_out`) over `http_client` outputs | [/recipes/webhook-fan-out](/recipes/webhook-fan-out) | | Log reduction | `file` in a batching `broker` input, `mapping`, `switch`, `group_by_value`, `archive`, `file` | [/recipes/log-reduction](/recipes/log-reduction) | | Sensor telemetry over MQTT | `mqtt`, `mapping`, `switch` to rejects, alerts and telemetry outputs | [/recipes/sensor-telemetry-mqtt](/recipes/sensor-telemetry-mqtt) | | CI test fixtures | `generate` with `pipeline.threads: 1`, `file`; `expanso-edge validate` as the gate | [/recipes/ci-fixtures](/recipes/ci-fixtures) | Every page, with each job, its dependencies and its proof limits, in one file: [/llms/recipes.txt](/llms/recipes.txt). All ran on `expanso-edge` v2.1.21. The first four ran through Expanso Cloud; the RSS, data migration and RAG jobs were revised afterwards and the published versions proved on a local-mode node. The last four ran on a local-mode node. The MQTT row used synthetic readings, with no OPC UA or sensor hardware. A pipeline is three stages: an **input**, an ordered list of **processors**, and an **output**, with optional buffers, caches, rate limits and scanners. The component reference lists **50 inputs, 80 processors and 55 outputs** (216 components across all seven kinds), including `http_client`, `http_server`, `mqtt`, `nats`, `aws_s3`, `gcp_cloud_storage`, `azure_blob_storage`, `opcua`, `postgres_cdc`, `mysql_cdc`, `aws_sqs`, `opensearch`, `qdrant`, and model processors such as `ollama_embeddings`, `aws_bedrock_embeddings`, `gcp_vertex_ai_embeddings` and `openai_embeddings`. A component being listed is not proof that a job using it has been run; check its page and status at [/llms/components.txt](/llms/components.txt). Transformation logic is written in Bloblang ([/llms/guides.txt](/llms/guides.txt)). Jobs run on nodes you choose, selected by label, so a pipeline can sit next to the data it reads and filter, reduce, redact or embed records before they cross a network. That is where Expanso is distinctively strong, and the same job runs the same way wherever the node is. Where Expanso is not the right answer: - It is not a workflow orchestrator, a batch query engine, or a data warehouse. - There is no component for Iceberg, Delta Lake or Hudi tables. - There is no email, Slack, Discord, PagerDuty or X (Twitter) component. HTTP APIs are reached with `http_client`; delivery to those services and X ingestion have not been proven. - Change data capture (`postgres_cdc`, `mysql_cdc`, `aws_dynamodb_cdc`) has not completed a verified run. - There is no native Anthropic (Claude) model processor and no token-aware text splitter. - On v2.1.21, write vectors to Qdrant with `http_client`, not the `qdrant` output: its writes failed while the job still reported `completed`. Expanso Edge is the current product. Predecessor projects are historical and are not a current product, a supported runtime, or a substitute for it. ## 2. Limits and Prerequisites - **A pipeline is not run directly.** There is no command that executes a bare `input`/`pipeline`/`output` file. Pipelines execute as **jobs**: the pipeline config is nested under a `config:` key alongside a `name`, a `type: pipeline` and an optional node `selector`, and the job is submitted to a control plane. - **`expanso-edge run --config` loads agent settings, not pipeline YAML.** Pass a pipeline config to it and the `input`, `pipeline`, `output` and `buffer` keys are reported as unknown fields and ignored: the agent starts, the pipeline does not run. - **Two different credentials.** A bootstrap token (`exp_bk_...`) registers an edge node via `expanso-edge bootstrap --token`. An API key (`exp_ak_...`) authenticates `expanso-cli` via `--api-key`. They are not interchangeable, and `expanso-cli` has no `--auth-token` flag. - **A node must exist.** `expanso-cli job deploy` returns once the control plane has stored the job, not once a node has picked it up. With no connected node matching the job's selector, the job is stored and never executes. - **The agent needs a host.** Installing the runtime alone does not produce a managed, assigned, monitored pipeline; that comes from the control plane. - **Paths are resolved on the node**, not on the machine you deploy from. A file input reads the file as the user running the agent on that node. - Private control planes are available in some configurations; check [Architecture](/getting-started/concepts/architecture/) rather than assuming. Full list of the failure modes agents hit most: [Common Mistakes](/llms/common-mistakes.txt). ## 3. Cost and Entitlement The first five nodes are free. Workspaces, nodes and keys are managed in Expanso Cloud at [cloud.expanso.io](https://cloud.expanso.io). Do not sign up for, provision, or spend against an account without the operator asking you to. What needs no account at all: checking a config. `expanso-edge validate` and `expanso-cli job validate --offline` run entirely locally, contact nothing, and need no credentials. ## 4. A First Command That Works ### Install Expanso Edge ```bash # Linux/macOS curl -fsSL https://get.expanso.io/edge/install.sh | bash # Verify installation expanso-edge version ``` `expanso-cli` is a separate binary and is needed to deploy a pipeline: ```bash curl -fsSL https://get.expanso.io/cli/install.sh | bash ``` ### Run Expanso Edge Pipelines run as jobs submitted to a control plane. Prerequisites: an Expanso Cloud workspace, an API key (`exp_ak_...`), a bootstrap token (`exp_bk_...`), and a host to run the edge agent on. ```bash # 1. Start an edge node, using a bootstrap token docker run -d \ ghcr.io/expanso-io/expanso-edge:latest run \ --bootstrap-token exp_bk_YOUR_TOKEN # 2. Point the CLI at your workspace with an API key, not the bootstrap token expanso-cli profile save prod \ --endpoint https://NETWORK_ID.us1.cloud.expanso.io:9010 \ --api-key exp_ak_YOUR_KEY --select expanso-cli status # 3. Confirm the node registered expanso-cli node list # 4. Create a job file (job wrapper format) cat > my-job.yaml << 'EOF' name: hello-world type: pipeline config: input: generate: mapping: 'root.message = "hello from edge"' interval: 1s pipeline: processors: [] output: stdout: {} EOF # 5. Check the spec, then deploy it expanso-cli job validate my-job.yaml --offline expanso-cli job deploy my-job.yaml # 6. Confirm it is executing, not merely stored expanso-cli job describe hello-world expanso-cli execution list --job-id JOB_ID expanso-cli job logs hello-world ``` `job deploy` returns once the control plane has stored the job, not once a node has picked it up, so a clean deploy is not evidence of execution. An empty `execution list` means nothing is running it — check that a node is connected and that the job's `selector` matches its labels. Offline, with no account and no node, `expanso-edge validate pipeline.yaml` checks a bare pipeline config's syntax and component types. It does not check job-level fields and does not guarantee the control plane will accept the job. See [Deploy to Cloud](/getting-started/deploy-to-cloud/) for the walkthrough. ## 5. Verify the Run A deploy that returned cleanly proves only that the job was accepted and stored. To establish that the pipeline actually ran, collect all three: 1. **Assignment** — `expanso-cli execution list --job-id JOB_ID` lists at least one execution, on a named node, in a running state. An empty list means nothing is executing. 2. **Node** — `expanso-cli node list` shows that node connected. 3. **Output** — `expanso-cli job logs my-job` shows records flowing, or the configured sink shows the records arriving. Without an execution on a node and evidence at the destination, treat the run as unverified. A generated config, a clean validation, and a successful deploy are each separately insufficient. ## 6. Diagnosis | What you see | What it means | One-line fix | |---|---|---| | Agent starts, no records, log says "Unknown configuration fields detected" | A pipeline config was passed to `--config`, which loads agent settings | Wrap the pipeline in a job spec and `expanso-cli job deploy` it | | `unknown flag: --auth-token` | That flag does not exist on `expanso-cli` | Use `--api-key exp_ak_...` | | CLI rejects your bootstrap token | `exp_bk_` registers nodes; it does not authenticate the CLI | Create an API key (`exp_ak_...`) in the workspace **Keys** tab | | Deploy succeeds, `execution list` is empty | No node matches the job's `selector`, or no node is connected | `expanso-cli node list`; fix the selector labels or bootstrap a node | | `expanso-cli node list` is empty | No edge node has registered with the workspace | `expanso-edge bootstrap --token exp_bk_...` then `expanso-edge run` | | Connection refused or TLS errors from the CLI | Wrong endpoint or port | Endpoint is `https://NETWORK_ID.REGION.cloud.expanso.io:9010`, from the workspace **API Access** dialog | | Deploy rejected: unknown top-level keys | Pipeline config sent where a job spec is expected | Nest it under `config:` and add `name` and `type: pipeline` | | `expanso-edge validate` passes but deploy is rejected | `validate` checks the pipeline config, not job-level fields | `expanso-cli job validate job.yaml --offline` | | File input reads nothing | The path is resolved on the node, by the agent's user | Use a path that exists on that node, readable by that user | | Job runs but the old behaviour persists | A previous job version is still deployed | Redeploy; `expanso-cli job versions my-job` shows history | Longer form: [Common Mistakes](/llms/common-mistakes.txt) and [Testing & Debugging](/getting-started/testing-debugging/). ## Documentation Sections - [Build by Job](/llms/recipes.txt): task-to-components matrix, complete jobs that were run end to end, and where Expanso is not the right answer - [Common Mistakes](/llms/common-mistakes.txt): config formats, credentials, endpoint, deploy recipe - [Getting Started](/llms/getting-started.txt): Installation, core concepts, first pipeline - [Components](/llms/components.txt): Index of all 216 components (fetch individual files for full docs) - [Inputs index](/llms/components/inputs.txt): All input components — fetch e.g. `/llms/components/inputs/kafka.txt` - [Processors index](/llms/components/processors.txt): All processor components - [Outputs index](/llms/components/outputs.txt): All output components - [Examples](/llms/examples.txt): Working pipeline examples - [Use Cases](/llms/use-cases.txt): Real-world scenarios (IoT, logs, analytics, alerting) - [Guides](/llms/guides.txt): Pipeline configuration, Bloblang, resources - [CLI Reference](/llms/cli.txt): expanso-edge and expanso-cli commands ## How to Use Component Docs Components use a two-level lookup to save tokens: 1. Fetch the type index (e.g., `/llms/components/inputs.txt`) — lightweight table of all components 2. Fetch only the component(s) you need (e.g., `/llms/components/inputs/kafka.txt`) — full config reference ## Pipeline Structure The pipeline config, which is what goes under `config:` in a job spec: ```yaml input: kafka: addresses: ["broker:9092"] topics: ["events"] pipeline: processors: - mapping: | root = this root.timestamp = now() output: aws_s3: bucket: my-bucket path: "${! timestamp_unix() }.json" ``` ## Related Resources - [Expanso Examples](https://examples.expanso.io/llms.txt): Production-ready pipeline examples and configurations - [Expanso Skills](https://skills.expanso.io/): Prebuilt pipeline wrappers; check each one's stated dependencies before use - [Expanso Cloud](https://cloud.expanso.io): Workspaces, nodes, bootstrap tokens and API keys - Documentation MCP server: `https://mcp.expanso.io/mcp`, declared in [/.well-known/mcp.json](/.well-known/mcp.json) ## Full Documentation See https://docs.expanso.io for complete documentation.