A synthetic data generation system that produces multi-turn conversations containing multi-step/multi-tool tool-use traces, grounded in ToolBench tool schemas.
Try the interactive demo — generate and inspect a synthetic tool-use trace with no API key.
pip install -e ".[dev]"Ingest ToolBench data and build the tool graph:
# Use bundled sample data (16 tools, 40 endpoints, 8 categories)
toolgen build
# Or point to a ToolBench data directory
toolgen build --data-dir /path/to/toolbench/data/toolenv/toolsThis creates artifacts/registry.json and artifacts/graph.json.
# Generate 10 conversations (default)
toolgen generate --seed 42
# Generate 100 conversations with steering enabled
toolgen generate --count 100 --seed 42 -o dataset_run_b.jsonl
# Run A: steering disabled (for diversity experiment)
toolgen generate --count 100 --seed 42 --no-cross-conversation-steering -o dataset_run_a.jsonl
# Run B: steering enabled (default)
toolgen generate --count 100 --seed 42 -o dataset_run_b.jsonlRequires OPENAI_API_KEY environment variable.
toolgen evaluate -d dataset_run_a.jsonl
toolgen evaluate -d dataset_run_b.jsonl
# Re-score all conversations (ignoring existing scores)
toolgen evaluate -d dataset.jsonl --rescore| Flag | Default | Description |
|---|---|---|
--data-dir |
bundled samples | Path to ToolBench-formatted data directory |
--output-dir |
artifacts |
Where to write registry.json and graph.json |
| Flag | Default | Description |
|---|---|---|
--artifacts-dir |
artifacts |
Path to built artifacts |
-o, --output |
dataset.jsonl |
Output JSONL path |
-n, --count |
10 |
Number of conversations to generate |
--seed |
42 |
Random seed |
--model |
gpt-4o-mini |
LLM model to use |
--no-cross-conversation-steering |
off | Disable diversity steering (for Run A) |
--quality-threshold |
3.0 |
Min mean judge score to accept |
--max-retries |
2 |
Max repair attempts per conversation |
| Flag | Default | Description |
|---|---|---|
-d, --dataset |
(required) | Path to JSONL dataset |
--artifacts-dir |
artifacts |
Path to built artifacts |
--model |
gpt-4o-mini |
LLM model for judging |
--threshold |
3.0 |
Quality threshold for pass rate |
--rescore |
off | Re-score even if scores exist |
Each line in the JSONL output is a conversation record:
{
"conversation_id": "conv_0042",
"messages": [
{"role": "user", "content": "Find me a hotel in Paris"},
{"role": "assistant", "content": "What's your budget?"},
{"role": "user", "content": "Under 200 EUR"},
{"role": "assistant", "tool_calls": [
{"endpoint": "hotel_service/search_hotels", "arguments": {"city": "Paris", "max_price": 200}}
]},
{"role": "tool", "content": {"results": [{"hotel_id": "htl_001", "name": "Hotel du Marais", "price_per_night": 175}]}},
{"role": "assistant", "tool_calls": [
{"endpoint": "hotel_service/book_hotel", "arguments": {"hotel_id": "htl_001", "check_in": "2026-06-01"}}
]},
{"role": "tool", "content": {"booking_id": "boo_002", "status": "confirmed"}},
{"role": "assistant", "content": "Booked Hotel du Marais. Confirmation: boo_002."}
],
"judge_scores": {"naturalness": 4.2, "tool_correctness": 4.8, "task_completion": 5.0},
"metadata": {
"seed": 42,
"tools_used": ["hotel_service/search_hotels", "hotel_service/book_hotel"],
"categories": ["Travel"],
"num_turns": 8,
"num_tool_calls": 2,
"has_disambiguation": true,
"scenario": "Book a hotel in Paris within budget",
"repair_attempts": 0
}
}# Unit + integration tests (no API key needed)
pytest tests/ --ignore=tests/test_e2e.py -v
# Full end-to-end test (requires OPENAI_API_KEY)
pytest tests/test_e2e.py -vsrc/toolgen/
models.py # Pydantic data models
registry.py # ToolBench ingestion + tool registry
graph.py # Tool graph (NetworkX)
sampler.py # Constrained tool-chain sampling
executor.py # Offline mock execution with session state
agents.py # Multi-agent conversation generator
prompts.py # LLM prompt templates
judge.py # LLM-as-judge scoring
repair.py # Automatic retry/repair
context.py # Diversity steering + metrics
cli.py # CLI entry point
data/sample_tools/ # Bundled sample ToolBench data (8 categories, 16 tools)
tests/ # Unit, integration, and e2e tests
See DESIGN.md for architecture details, design decisions, and analysis.