aws_bedrock_embeddings
Computes vector embeddings on text, using the AWS Bedrock API.
- Common
- Advanced
# Common config fields, showing default values
pipeline:
processors:
- label: ""
aws_bedrock_embeddings:
model: "" # No default (required)
text: "" # No default (optional)
input_type: "" # No default (optional)
# All config fields, showing default values
pipeline:
processors:
- label: ""
aws_bedrock_embeddings:
region: "" # No default (optional)
endpoint: "" # No default (optional)
tcp:
connect_timeout: "0s"
keep_alive:
idle: "15s"
interval: "15s"
count: 9
tcp_user_timeout: "0s"
credentials:
profile: "" # No default (optional)
id: "" # No default (optional)
secret: "" # No default (optional)
token: "" # No default (optional)
from_ec2_role: false # No default (optional)
role: "" # No default (optional)
role_external_id: "" # No default (optional)
model: "" # No default (required)
text: "" # No default (optional)
input_type: "" # No default (optional)
This processor sends text to your chosen large language model (LLM) and computes vector embeddings, using the AWS Bedrock API. For more information, see the AWS Bedrock documentation.
Examples
Store embedding vectors in Clickhouse
Compute embeddings for some generated data and store it within Clickhouse
input:
generate:
interval: 1s
mapping: |
root = {"text": fake("paragraph")}
pipeline:
processors:
- branch:
request_map: |
root = this.text
processors:
- aws_bedrock_embeddings:
model: amazon.titan-embed-text-v1
result_map: |
root.embeddings = this
output:
sql_insert:
driver: clickhouse
dsn: "clickhouse://localhost:9000"
table: searchable_text
columns: ["id", "text", "vector"]
args_mapping: "root = [uuid_v4(), this.text, this.embeddings]"
Fields
region
The AWS region to target.
Type: string
endpoint
Allows you to specify a custom endpoint for the AWS API.
Type: string
tcp
TCP socket configuration.
Type: object
tcp.connect_timeout
Maximum amount of time a dial will wait for a connect to complete. Zero disables.
Type: string
Default: "0s"
tcp.keep_alive
TCP keep-alive probe configuration.
Type: object
tcp.keep_alive.idle
Duration the connection must be idle before sending the first keep-alive probe. Zero defaults to 15s. Negative values disable keep-alive probes.
Type: string
Default: "15s"
tcp.keep_alive.interval
Duration between keep-alive probes. Zero defaults to 15s.
Type: string
Default: "15s"
tcp.keep_alive.count
Maximum unanswered keep-alive probes before dropping the connection. Zero defaults to 9.
Type: int
Default: 9
tcp.tcp_user_timeout
Maximum time to wait for acknowledgment of transmitted data before killing the connection. Linux-only (kernel 2.6.37+), ignored on other platforms. When enabled, keep_alive.idle must be greater than this value per RFC 5482. Zero disables.
Type: string
Default: "0s"
credentials
Optional manual configuration of AWS credentials to use. More information can be found in this document.
Type: object
credentials.profile
A profile from ~/.aws/credentials to use.
Type: string
credentials.id
The ID of credentials to use.
Type: string
credentials.secret
The secret for the credentials being used.
This field contains sensitive information. Use a secret reference rather than a literal value.
Type: string
credentials.token
The token for the credentials being used, required when using short term credentials.
This field contains sensitive information. Use a secret reference rather than a literal value.
Type: string
credentials.from_ec2_role
Use the credentials of a host EC2 machine configured to assume an IAM role associated with the instance.
Type: bool
credentials.role
A role ARN to assume.
Type: string
credentials.role_external_id
An external ID to provide when assuming a role.
Type: string
model
The model ID to use. For a full list see the AWS Bedrock documentation.
Type: string
text
The prompt you want to generate a response for. By default, the processor submits the entire payload as a string.
Type: string
input_type
Specifies the type of input passed to the model. Required by Cohere embedding models; ignored by Amazon Titan models.
Type: string
| Option | Summary |
|---|---|
classification | Used for embeddings passed through a text classifier. |
clustering | Used for the embeddings run through a clustering algorithm. |
search_document | Used for embeddings stored in a vector database for search use-cases. |
search_query | Used for embeddings of search queries run against a vector DB to find relevant documents. |