- Updated the evaluation runner to support live dialogue printing and transcript saving in both Markdown and JSON formats. - Introduced new data classes for managing transcripts and their components, improving structure and readability. - Modified existing YAML case files to align with the new direct client mode and updated stage codes. - Added new test cases for the direct client mode, ensuring comprehensive coverage of user interactions.
Workflow evaluations
The evaluation runner calls a live FastAPI service backed by FastGPT. For now,
each conversation step checks only the SSE stage_code.nextStageCode value.
During the run it also prints the live dialogue and saves transcripts.
Case format
Add YAML files anywhere under evals/cases/:
name: direct_initial_stage
clientMode: direct
steps:
- input: 新对话
expect:
stageCode: "1001"
- input: 继续办理
expect:
stageCode: "1002"
clientMode in the case file is the API request setting (direct or
browser_addon). Steps in one file share the same generated sessionId.
Different cases always use different session IDs.
Run
Start the API, then execute:
python evals/run.py
Use another API address:
python evals/run.py --base-url http://192.168.1.10:8000
Run selected cases:
python evals/run.py --case direct
python evals/run.py --case browser_addon
Dialogue and transcripts
While a case runs, each step prints the user input, bot reply (text_delta),
and stage-code check result.
By default, after the run finishes, transcripts are written under
evals/results/<timestamp>/:
*.md— readable conversation log*.json— structured step results
Save to a specific directory:
python evals/run.py --output-dir evals/results/manual-run --case direct
Print live dialogue only, without writing files:
python evals/run.py --no-save --case direct
The process exits with code 1 when an expectation fails and code 2 when
case configuration is invalid.