AgentDiagnose: An Open Toolkit for Diagnosing LLM Agent Trajectories
July 16, 2025 ยท View on GitHub
Usage
AgentDiagnose provides two main ways to analyze agent trajectories:
Option 1: All-In-One Dashboard
The easiest way to get started is to use the launch dashboard script, which runs the complete pipeline automatically:
./launch_dashboard.sh
This script will:
- Extract verb-noun pairs from trajectories
- Generate tag clouds for reasoning and action phrases
- Generate embeddings on action phrases
- Evaluate trajectories with quality scorers
- Launch the interactive web dashboard
Option 2: Run Evaluator With Custom Options
For more control over the evaluation of specific agent qualities, use the evaluator directly:
# Evaluate with specific scorers
python evaluate_trajectories.py --input examples/sample_trajectories --scorers reasoning_quality objective_quality --output-json results.json
# Dry run to estimate token usage and costs
python evaluate_trajectories.py --input examples/sample_trajectories --scorers reasoning_quality objective_quality --output-json results.json --dry-run
Available Scorers
reasoning_quality: Evaluates the quality of the agent's reasoningobjective_quality: Assesses the task's objective qualitynavigation_path: Analyzes navigation paths within trajectories
Configuration
Set your LLM API key in the environment:
export LLM_API_KEY="your-api-key-here"
Web Dashboard
The web dashboard provides an interactive interface for:
- Visualizing trajectory evaluator's results
- Exploring reasoning patterns through tag clouds
- Analyzing action phrase distributions
- Examining embedding-based trajectory clustering
Access the dashboard at the URL displayed when launching (defaulthttp://localhost:8080).