---
format: "aidr-story-markdown/v1"
id: "66b71cdbe19cfabcafd62ebcf16cc7e02846eb09a3bcb6b0b1e632b73726349d"
canonical_url: "https://aidr.today/66b71cdb?lang=en"
title: "Right-size generative AI endpoints with concurrency sweeps on Amazon SageMaker AI"
lang: "en"
requested_lang: "en"
available_langs: ["en","vi"]
translation_fallback: null
fallback_fields: []
published_at: "2026-09-22T15:35:53.000Z"
category: "Infra"
topics: ["amazon","sagemaker","inference","benchmark","infra","chips"]
source_urls: ["https://aws.amazon.com/blogs/machine-learning/right-size-generative-ai-endpoints-with-concurrency-sweeps-on-amazon-sagemaker-ai/"]
summary: "Concurrency sweeps help you right-size a generative AI endpoint on Amazon SageMaker AI by systematically benchmarking it at increasing load levels. This post walks through deploying a model, running automated concurrency sweeps with the CreateAIBenchmarkJob API, and using the results to make data-driven capacity decisions about fleet size."
---

# Right\-size generative AI endpoints with concurrency sweeps on Amazon SageMaker AI

> [Open the canonical story](<https://aidr.today/66b71cdb?lang=en>)

**Published:** 2026-09-22T15:35:53.000Z
**Category:** Infra
**Topics:** amazon, sagemaker, inference, benchmark, infra, chips

## Summary

Concurrency sweeps help you right\-size a generative AI endpoint on Amazon SageMaker AI by systematically benchmarking it at increasing load levels\. This post walks through deploying a model, running automated concurrency sweeps with the CreateAIBenchmarkJob API, and using the results to make data\-driven capacity decisions about fleet size\.

## Sources

- [Story source](<https://aws.amazon.com/blogs/machine-learning/right-size-generative-ai-endpoints-with-concurrency-sweeps-on-amazon-sagemaker-ai/>)

