Add vision model enhancements and image processing capabilities
- Update `requirements.txt` to include Pillow for image handling. - Refactor vision model validation logic in `voice_webrtc.py` to improve error handling for unsupported image input. - Introduce new functions in `pipeline.py` for image data processing and analysis using vision models. - Implement `VisionCaptureProcessor` to manage video frame requests for auxiliary vision model analysis. - Enhance the pipeline to support image input requests and integrate vision model responses into the processing flow.
This commit is contained in:
@@ -3,6 +3,7 @@
|
||||
# silero -> 本地 VAD(判断用户说话起止),语音必备
|
||||
# openai -> OpenAI 兼容的 LLM/STT/TTS 客户端(DeepSeek、SenseVoice、CosyVoice 都走它)
|
||||
pipecat-ai[webrtc,websocket,silero,openai]==1.3.0
|
||||
Pillow>=11.1.0,<13
|
||||
|
||||
# FastGPT 类型助手:本地 SDK(包 /api/v1/chat/completions 流式 + chatId 会话)
|
||||
fastgpt-client @ file:///Users/wangx/Code/AI-VideoAssistant-Project/fastgpt-python-sdk
|
||||
|
||||
Reference in New Issue
Block a user