Skip to content
gauravch-codePublic

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

2 Commits

Folders and files

Repository files navigation

toolgen — Multi-Agent Tool-Use Conversation Generator

A synthetic data generation system that produces multi-turn conversations containing multi-step/multi-tool tool-use traces, grounded in ToolBench tool schemas.

Try the interactive demo — generate and inspect a synthetic tool-use trace with no API key.

Quick Start

1. Install

pip install -e ".[dev]"

2. Build artifacts

Ingest ToolBench data and build the tool graph:

# Use bundled sample data (16 tools, 40 endpoints, 8 categories)
toolgen build

# Or point to a ToolBench data directory
toolgen build --data-dir /path/to/toolbench/data/toolenv/tools

This creates artifacts/registry.json and artifacts/graph.json.

3. Generate conversations

# Generate 10 conversations (default)
toolgen generate --seed 42

# Generate 100 conversations with steering enabled
toolgen generate --count 100 --seed 42 -o dataset_run_b.jsonl

# Run A: steering disabled (for diversity experiment)
toolgen generate --count 100 --seed 42 --no-cross-conversation-steering -o dataset_run_a.jsonl

# Run B: steering enabled (default)
toolgen generate --count 100 --seed 42 -o dataset_run_b.jsonl

Requires OPENAI_API_KEY environment variable.

4. Evaluate

toolgen evaluate -d dataset_run_a.jsonl
toolgen evaluate -d dataset_run_b.jsonl

# Re-score all conversations (ignoring existing scores)
toolgen evaluate -d dataset.jsonl --rescore

CLI Reference

toolgen build

Flag Default Description
--data-dir bundled samples Path to ToolBench-formatted data directory
--output-dir artifacts Where to write registry.json and graph.json

toolgen generate

Flag Default Description
--artifacts-dir artifacts Path to built artifacts
-o, --output dataset.jsonl Output JSONL path
-n, --count 10 Number of conversations to generate
--seed 42 Random seed
--model gpt-4o-mini LLM model to use
--no-cross-conversation-steering off Disable diversity steering (for Run A)
--quality-threshold 3.0 Min mean judge score to accept
--max-retries 2 Max repair attempts per conversation

toolgen evaluate

Flag Default Description
-d, --dataset (required) Path to JSONL dataset
--artifacts-dir artifacts Path to built artifacts
--model gpt-4o-mini LLM model for judging
--threshold 3.0 Quality threshold for pass rate
--rescore off Re-score even if scores exist

Output Format

Each line in the JSONL output is a conversation record:

{
  "conversation_id": "conv_0042",
  "messages": [
    {"role": "user", "content": "Find me a hotel in Paris"},
    {"role": "assistant", "content": "What's your budget?"},
    {"role": "user", "content": "Under 200 EUR"},
    {"role": "assistant", "tool_calls": [
      {"endpoint": "hotel_service/search_hotels", "arguments": {"city": "Paris", "max_price": 200}}
    ]},
    {"role": "tool", "content": {"results": [{"hotel_id": "htl_001", "name": "Hotel du Marais", "price_per_night": 175}]}},
    {"role": "assistant", "tool_calls": [
      {"endpoint": "hotel_service/book_hotel", "arguments": {"hotel_id": "htl_001", "check_in": "2026-06-01"}}
    ]},
    {"role": "tool", "content": {"booking_id": "boo_002", "status": "confirmed"}},
    {"role": "assistant", "content": "Booked Hotel du Marais. Confirmation: boo_002."}
  ],
  "judge_scores": {"naturalness": 4.2, "tool_correctness": 4.8, "task_completion": 5.0},
  "metadata": {
    "seed": 42,
    "tools_used": ["hotel_service/search_hotels", "hotel_service/book_hotel"],
    "categories": ["Travel"],
    "num_turns": 8,
    "num_tool_calls": 2,
    "has_disambiguation": true,
    "scenario": "Book a hotel in Paris within budget",
    "repair_attempts": 0
  }
}

Running Tests

# Unit + integration tests (no API key needed)
pytest tests/ --ignore=tests/test_e2e.py -v

# Full end-to-end test (requires OPENAI_API_KEY)
pytest tests/test_e2e.py -v

Project Structure

src/toolgen/
  models.py       # Pydantic data models
  registry.py     # ToolBench ingestion + tool registry
  graph.py        # Tool graph (NetworkX)
  sampler.py      # Constrained tool-chain sampling
  executor.py     # Offline mock execution with session state
  agents.py       # Multi-agent conversation generator
  prompts.py      # LLM prompt templates
  judge.py        # LLM-as-judge scoring
  repair.py       # Automatic retry/repair
  context.py      # Diversity steering + metrics
  cli.py          # CLI entry point

data/sample_tools/  # Bundled sample ToolBench data (8 categories, 16 tools)
tests/              # Unit, integration, and e2e tests

See DESIGN.md for architecture details, design decisions, and analysis.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages