thread simplification + handling interuption
This commit is contained in:
@@ -1,9 +1,6 @@
|
||||
- Added cross-sentence stitching support to NvidiaTTSService. Multiple sentences within an LLM turn are fed into a single `SynthesizeOnline` gRPC stream, enabling Magpie's seamless audio across sentence boundaries (requires Magpie TTS model v1.7.0+).
|
||||
- Added new parameters to NvidiaTTSService:
|
||||
- `custom_dictionary` for IPA-based custom pronunciation
|
||||
- `encoding` for output audio encoding format
|
||||
- `zero_shot_audio_prompt_file` for zero-shot voice cloning with file caching
|
||||
- `audio_prompt_encoding` for zero-shot audio prompt format
|
||||
- Added `can_generate_metrics` returning True for NvidiaTTSService.
|
||||
- Added `stop_all_metrics()` call on audio context interruption.
|
||||
- Added gRPC error handling for synthesis config retrieval.
|
||||
- Added enhancements to `NvidiaTTSService`:
|
||||
|
||||
- Cross-sentence stitching: multiple sentences within an LLM turn are fed into a single `SynthesizeOnline` gRPC stream for seamless audio across sentence boundaries (requires Magpie TTS model v1.7.0+).
|
||||
- `custom_dictionary` and `encoding` parameters for IPA-based custom pronunciation and output audio encoding.
|
||||
- Metrics generation (`can_generate_metrics` returns true) and `stop_all_metrics()` when an audio context is interrupted.
|
||||
- gRPC error handling around synthesis config retrieval (`GetRivaSynthesisConfig`).
|
||||
@@ -1,5 +1,8 @@
|
||||
- Updated NvidiaTTSService:
|
||||
- Made `api_key` optional for local NIM deployments
|
||||
- Fixed `_update_settings` to correctly indicate that runtime settings updates (voice, language, quality) take effect on the next turn
|
||||
- Replaced per-sentence synchronous `synthesize_online` calls with async queue-backed gRPC streaming
|
||||
- Renamed Riva references to Nemotron Speech
|
||||
- Updated `NvidiaTTSService`:
|
||||
|
||||
- Made `api_key` optional for local NIM deployments.
|
||||
- Voice, language, and quality can be updated without reconnecting the gRPC client; new values take effect on the next synthesis turn, not for the current turn's in-flight requests.
|
||||
- Replaced per-sentence synchronous `synthesize_online` calls with async queue-backed gRPC streaming.
|
||||
- Streaming now uses asyncio tasks with explicit gRPC cancellation on interruption and stale-response filtering when a stream is aborted or replaced.
|
||||
- Renamed Riva references to Nemotron Speech in docs and messages.
|
||||
- Disabled automatic TTS start frames at the service level (`push_start_frame=False`) and emit `TTSStartedFrame` when a stitched synthesis stream is started for a context.
|
||||
Reference in New Issue
Block a user