The first step is to create voice signatures for the conversation participants. Creating voice signatures is required for efficient speaker identification. Note: In addition to the standard baseline model used by the Speech Services, you can customize models to your needs with available data, to overcome speech recognition barriers such as speaking style, vocabulary and background noise. References: https://docs.microsoft.com/bs-latn-ba/azure/cognitive-services/speech-service/how-to-use-conversationtranscription-service