Microsoft Foundry Agent Service を使用して構築されたカスタマーサポートエージェントがあります。このエージェントは、Azure OpenAL モデルのデプロイメントを呼び出します。 負荷テスト中に、呼び出しが断続的に失敗し、HTTP 429 レート制限超過エラーが返されます。 負荷がかかった状態での通話失敗を減らし、信頼性を向上させるためには、スロットリングを適切に処理する必要があります。このソリューションは、サービスおよびモデルの制限内に収まるものでなければなりません。 あなたはどうすべきでしょうか?
正解:A
The correct answer is A. Implement a retry policy that uses exponential backoff and jitter . HTTP 429 indicates that the request rate or token rate has exceeded the configured service limits for the model deployment. Microsoft Foundry Agent Service limits guidance specifically recommends implementing exponential backoff with jitter in application retry logic when agents receive rate-limit 429 errors. It also recommends reviewing Azure OpenAI quotas and token-per-minute and request-per-minute limits for the deployment. This approach improves reliability while remaining within service limits because retries are delayed progressively instead of immediately adding more pressure to an already throttled deployment. Microsoft Foundry Models quota guidance also states that unsuccessful requests still count toward per-minute rate limits and that continuously resending requests without backing off makes throttling worse. It recommends retry logic with exponential backoff and using the Retry-After header when available. Creating a new thread and retrying immediately does not change the deployment's rate limits and can worsen throttling. Reducing registered tools may simplify orchestration but does not directly address model RPM or TPM limits. Splitting uploaded content into smaller files may help ingestion scenarios, but it is not the correct throttling control for intermittent HTTP 429 model calls. Reference topics: Foundry Agent Service limits, Azure OpenAI quota management, throttling, retry policies, and production reliability.