Configuration options
Configuring with context managers
The methods use_configuration(...) and use_model(...) are designed to work in context managers. Once the context manager is left, old config values are restored.
from inference_sdk import InferenceHTTPClient, InferenceConfiguration
image_url = "https://source.roboflow.com/pwYAXv9BTpqLyFfgQoPZ/u48G0UpWfk8giSw7wrU8/original.jpg"
custom_configuration = InferenceConfiguration(confidence_threshold=0.8)
# Replace ROBOFLOW_API_KEY with your Roboflow API Key
CLIENT = InferenceHTTPClient(
api_url="http://localhost:9001",
api_key="ROBOFLOW_API_KEY"
)
with CLIENT.use_configuration(custom_configuration):
_ = CLIENT.infer(image_url, model_id="soccer-players-5fuqs/1")
with CLIENT.use_model("soccer-players-5fuqs/1"):
_ = CLIENT.infer(image_url)
# after leaving context manager - changes are reverted and `model_id` is still required
_ = CLIENT.infer(image_url, model_id="soccer-players-5fuqs/1")As you can see, model_id is required for a prediction method only when a default model is not configured.
The model ID is composed of the string <project_id>/<version_id>. See Workspace and Project IDs to find these pieces of information.
Setting the configuration once and using it until the next change
The methods configure(...) and select_model(...) alter the client state and the change is preserved until the next change.
from inference_sdk import InferenceHTTPClient, InferenceConfiguration
image_url = "https://source.roboflow.com/pwYAXv9BTpqLyFfgQoPZ/u48G0UpWfk8giSw7wrU8/original.jpg"
custom_configuration = InferenceConfiguration(confidence_threshold=0.8)
# Replace ROBOFLOW_API_KEY with your Roboflow API Key
CLIENT = InferenceHTTPClient(
api_url="http://localhost:9001",
api_key="ROBOFLOW_API_KEY"
)
CLIENT.configure(custom_configuration)
CLIENT.infer(image_url, model_id="soccer-players-5fuqs/1")
# custom configuration still holds
CLIENT.select_model(model_id="soccer-players-5fuqs/1")
_ = CLIENT.infer(image_url)
# custom configuration and selected model - still holds
_ = CLIENT.infer(image_url)You may also initialise in chain mode:
from inference_sdk import InferenceHTTPClient, InferenceConfiguration
# Replace ROBOFLOW_API_KEY with your Roboflow API Key
CLIENT = InferenceHTTPClient(api_url="http://localhost:9001", api_key="ROBOFLOW_API_KEY") \
.select_model("soccer-players-5fuqs/1")Overriding model_id for a specific call
model_id can be overridden for a specific call:
from inference_sdk import InferenceHTTPClient
image_url = "https://source.roboflow.com/pwYAXv9BTpqLyFfgQoPZ/u48G0UpWfk8giSw7wrU8/original.jpg"
# Replace ROBOFLOW_API_KEY with your Roboflow API Key
CLIENT = InferenceHTTPClient(api_url="http://localhost:9001", api_key="ROBOFLOW_API_KEY") \
.select_model("soccer-players-5fuqs/1")
_ = CLIENT.infer(image_url, model_id="another-model/1")Details about client configuration
InferenceHTTPClient provides the InferenceConfiguration dataclass to hold the full configuration.
from inference_sdk import InferenceConfigurationOverriding fields in this config changes the behaviour of the client (and of the API serving the model). Specific fields are used in specific contexts. In particular:
Classification model
visualize_predictions: flag to enable / disable visualisationconfidence_thresholdasconfidencestroke_width: width of stroke in visualisationdisable_preproc_auto_orientation,disable_preproc_contrast,disable_preproc_grayscale,disable_preproc_static_cropto alter server-side pre-processingdisable_active_learningto prevent the Active Learning feature from registering the datapoint (can be useful, for instance, while testing a model)active_learning_target_dataset- when making inference from a specific model (let's sayproject_a/1) and you want to save data in another projectproject_b, the latter should be pointed to by this parameter. Note that you cannot use different types of models inproject_aandproject_b; if that is the case, data will not be registered.source: optional string that sets a "source" attribute on the inference call. If using model monitoring, this is logged with the inference request so you can filter or query inference requests coming from a particular source, for example to identify which application, system, or deployment is making the request.source_info: optional string that sets an additional "source_info" attribute on the inference call, for example to identify a sub-component in an app.
Object detection model
visualize_predictions: flag to enable / disable visualisationvisualize_labels: flag to enable / disable label visualisation if visualisation is enabledconfidence_thresholdasconfidenceclass_filterto filter out a list of classesclass_agnostic_nms: flag to control whether NMS is class-agnosticfix_batch_sizeiou_threshold: to dictate the NMS IoU thresholdstroke_width: width of stroke in visualisationmax_detections: max detections to return from the modelmax_candidates: max candidates for post-processing from the modeldisable_preproc_auto_orientation,disable_preproc_contrast,disable_preproc_grayscale,disable_preproc_static_cropto alter server-side pre-processingdisable_active_learning,active_learning_target_dataset,source,source_info- as described above
Keypoint detection model
visualize_predictions: flag to enable / disable visualisationvisualize_labels: flag to enable / disable label visualisation if visualisation is enabledconfidence_thresholdasconfidencekeypoint_confidence_threshold(askeypoint_confidence) to filter out detected keypoints based on model confidenceclass_filterto filter out a list of object classesclass_agnostic_nms: flag to control whether NMS is class-agnosticfix_batch_sizeiou_threshold: to dictate the NMS IoU thresholdstroke_width: width of stroke in visualisationmax_detections: max detections to return from the modelmax_candidates: max candidates for post-processing from the modeldisable_preproc_auto_orientation,disable_preproc_contrast,disable_preproc_grayscale,disable_preproc_static_cropto alter server-side pre-processingdisable_active_learning,active_learning_target_dataset,source,source_info- as described above
Instance segmentation model
visualize_predictions: flag to enable / disable visualisationvisualize_labels: flag to enable / disable label visualisation if visualisation is enabledconfidence_thresholdasconfidenceclass_filterto filter out a list of classesclass_agnostic_nms: flag to control whether NMS is class-agnosticfix_batch_sizeiou_threshold: to dictate the NMS IoU thresholdstroke_width: width of stroke in visualisationmax_detections: max detections to return from the modelmax_candidates: max candidates for post-processing from the modeldisable_preproc_auto_orientation,disable_preproc_contrast,disable_preproc_grayscale,disable_preproc_static_cropto alter server-side pre-processingmask_decode_modetradeoff_factordisable_active_learning,active_learning_target_dataset,source,source_info- as described above
Configuration of the client
output_visualisation_format: one ofVisualisationResponseFormat.BASE64,VisualisationResponseFormat.NUMPY,VisualisationResponseFormat.PILLOW. Given that server-side visualisation is enabled, you may choose which format should be used in the output.client_downsizing_disabled: set toFalseif you want to perform client-side downsizing. DefaultTrue. Client-side scaling is only supposed to down-scale (keeping aspect ratio) the input for inference, to utilise the internet connection more efficiently (at the price of image manipulation / transcoding). Model input size information is used to determine the target size; if not available,default_max_input_sizeis used.max_concurrent_requests: max number of concurrent requests that can be startedmax_batch_size: max number of elements that can be injected into a single requestworkflow_run_retries_enabled: flag that decides if transient errors in Workflows executions should be retried. Defaults totrueand the default can be altered with the environment variableWORKFLOW_RUN_RETRIES_ENABLED.
Configuration of Workflows execution
profiling_directory: specifies the location where Workflows profiler traces are saved. By default, it is the./inference_profilingdirectory.