Create an Anthropic inference endpoint | Elasticsearch API documentation (v9)

Path parameters

task_typestring Required
The task type. The only valid task type for the model to perform is completion.
Value is completion.
anthropic_inference_idstring Required
The unique identifier of the inference endpoint.

application/json

Body

chunking_settingsobject
Hide chunking_settings attributes Show chunking_settings attributes object
- max_chunk_sizenumber
  The maximum size of a chunk in words. This value cannot be higher than 300 or lower than 20 (for sentence strategy) or 10 (for word strategy).
- overlapnumber
  The number of overlapping words for chunks. It is applicable only to a word chunking strategy. This value cannot be higher than half the max_chunk_size value.
- sentence_overlapnumber
  The number of overlapping sentences for chunks. It is applicable only for a sentence chunking strategy. It can be either 1 or 0.
- strategystring
  The chunking strategy: sentence or word.
servicestring Required
Value is anthropic.
service_settingsobject Required
Hide service_settings attributes Show service_settings attributes object
- api_keystring Required
  A valid API key for the Anthropic API.
- model_idstring Required
  The name of the model to use for the inference task. Refer to the Anthropic documentation for the list of supported models.
- rate_limitobject
  Hide rate_limit attribute Show rate_limit attribute object
  requests_per_minutenumber
  The number of requests allowed per minute.
task_settingsobject
Hide task_settings attributes Show task_settings attributes object
- max_tokensnumber Required
  For a completion task, it is the maximum number of tokens to generate before stopping.
- temperaturenumber
  For a completion task, it is the amount of randomness injected into the response. For more details about the supported range, refer to Anthropic documentation.
  External documentation
- top_knumber
  For a completion task, it specifies to only sample from the top K options for each subsequent token. It is recommended for advanced use cases only. You usually only need to use temperature.
- top_pnumber
  For a completion task, it specifies to use Anthropic's nucleus sampling. In nucleus sampling, Anthropic computes the cumulative distribution over all the options for each subsequent token in decreasing probability order and cuts it off once it reaches the specified probability. You should either alter temperature or top_p, but not both. It is recommended for advanced use cases only. You usually only need to use temperature.

Responses

200 application/json
Hide response attributes Show response attributes object
- chunking_settingsobject
  Hide chunking_settings attributes Show chunking_settings attributes object
  max_chunk_sizenumber
  The maximum size of a chunk in words. This value cannot be higher than 300 or lower than 20 (for sentence strategy) or 10 (for word strategy).
  overlapnumber
  The number of overlapping words for chunks. It is applicable only to a word chunking strategy. This value cannot be higher than half the max_chunk_size value.
  sentence_overlapnumber
  The number of overlapping sentences for chunks. It is applicable only for a sentence chunking strategy. It can be either 1 or 0.
  strategystring
  The chunking strategy: sentence or word.
- servicestring Required
  The service type
- service_settingsobject Required
- task_settingsobject
- inference_idstring Required
  The inference Id
- task_typestring Required
  Value is completion.

PUT /_inference/{task_type}/{anthropic_inference_id}

PUT _inference/completion/anthropic_completion
{
    "service": "anthropic",
    "service_settings": {
        "api_key": "Anthropic-Api-Key",
        "model_id": "Model-ID"
    },
    "task_settings": {
        "max_tokens": 1024
    }
}

curl \
 --request PUT 'http://api.example.com/_inference/{task_type}/{anthropic_inference_id}' \
 --header "Authorization: $API_KEY" \
 --header "Content-Type: application/json" \
 --data '"{\n    \"service\": \"anthropic\",\n    \"service_settings\": {\n        \"api_key\": \"Anthropic-Api-Key\",\n        \"model_id\": \"Model-ID\"\n    },\n    \"task_settings\": {\n        \"max_tokens\": 1024\n    }\n}"'

Request example

Run `PUT _inference/completion/anthropic_completion` to create an inference endpoint that performs a completion task.

{
    "service": "anthropic",
    "service_settings": {
        "api_key": "Anthropic-Api-Key",
        "model_id": "Model-ID"
    },
    "task_settings": {
        "max_tokens": 1024
    }
}