Merge pull request #2993 from pipecat-ai/mb/fix-gemini-token-counting
fix: correct GoogleLLMService token counting
This commit is contained in:
@@ -87,6 +87,9 @@ reason")`.
|
|||||||
- `GeminiLiveLLMService` now properly supports context-provided system
|
- `GeminiLiveLLMService` now properly supports context-provided system
|
||||||
instruction and tools.
|
instruction and tools.
|
||||||
|
|
||||||
|
- Fixed `GoogleLLMService` token counting to avoid double-counting tokens when
|
||||||
|
Gemini sends usage metadata across multiple streaming chunks.
|
||||||
|
|
||||||
### Removed
|
### Removed
|
||||||
|
|
||||||
- Removed `needs_mcp_alternate_schema()` from `LLMService`. The mechanism that
|
- Removed `needs_mcp_alternate_schema()` from `LLMService`. The mechanism that
|
||||||
|
|||||||
@@ -899,12 +899,18 @@ class GoogleLLMService(LLMService):
|
|||||||
async for chunk in response:
|
async for chunk in response:
|
||||||
# Stop TTFB metrics after the first chunk
|
# Stop TTFB metrics after the first chunk
|
||||||
await self.stop_ttfb_metrics()
|
await self.stop_ttfb_metrics()
|
||||||
|
# Gemini may send usage_metadata in multiple chunks with varying behavior:
|
||||||
|
# - Sometimes a single chunk, sometimes multiple chunks
|
||||||
|
# - Token counts may be cumulative (growing) or may change between chunks
|
||||||
|
# - Early chunks may include estimates/overhead that gets refined
|
||||||
|
# We use assignment (not accumulation) because the final chunk always contains
|
||||||
|
# the authoritative, billable token usage for the entire response.
|
||||||
if chunk.usage_metadata:
|
if chunk.usage_metadata:
|
||||||
prompt_tokens += chunk.usage_metadata.prompt_token_count or 0
|
prompt_tokens = chunk.usage_metadata.prompt_token_count or 0
|
||||||
completion_tokens += chunk.usage_metadata.candidates_token_count or 0
|
completion_tokens = chunk.usage_metadata.candidates_token_count or 0
|
||||||
total_tokens += chunk.usage_metadata.total_token_count or 0
|
total_tokens = chunk.usage_metadata.total_token_count or 0
|
||||||
cache_read_input_tokens += chunk.usage_metadata.cached_content_token_count or 0
|
cache_read_input_tokens = chunk.usage_metadata.cached_content_token_count or 0
|
||||||
reasoning_tokens += chunk.usage_metadata.thoughts_token_count or 0
|
reasoning_tokens = chunk.usage_metadata.thoughts_token_count or 0
|
||||||
|
|
||||||
if not chunk.candidates:
|
if not chunk.candidates:
|
||||||
continue
|
continue
|
||||||
|
|||||||
Reference in New Issue
Block a user