Agent Eval & Testing (Cage Eval)

Cage Eval lets you define reproducible test scenarios for AI agents. Run agents against assertions, detect flaky behavior, and catch regressions before shipping.

Concepts

  • Scenario — A single test case: a prompt, a sandbox template, and assertions to verify the output
  • Suite — A group of related scenarios for batch execution
  • Run — A single execution of a scenario, producing assertion results and a score
  • Assertion — A verification check (file exists, output contains, cost below, etc.)

YAML File Format

Suite File (.caged-eval.yaml)

name: "API Agent Tests"
description: "Verify the agent can build REST APIs"
tags: [api, backend]
defaults:
  template: node-20
  budget: 0.50
  timeout: 5m
  model: claude-sonnet-4-20250514
scenarios:
  - name: "Create Express server"
    prompt: "Create a basic Express.js server with a /health endpoint"
    assertions:
      - { id: server, type: file_exists, path: "server.js" }
      - { id: health, type: file_contains, path: "server.js", contains: "/health" }
      - { id: cost, type: cost_below, max_cost: 0.25 }

  - name: "Add unit tests"
    prompt: "Add Jest tests for the server"
    assertions:
      - { id: tests, type: file_exists, path: "test/server.test.js" }
      - { id: passes, type: exit_code, command: "npm test", exit_code: 0 }

Single Scenario File (*.scenario.yaml)

name: "Hello World Python"
template: python-3.12
prompt: "Create a hello world Python script"
budget: 0.10
assertions:
  - { id: main, type: file_exists, path: "main.py" }
  - { id: output, type: output_contains, command: "python main.py", contains: "Hello" }

Assertion Types

Type Description Key Fields
file_exists File exists at path path
file_contains File contains substring path, contains
file_matches File content matches regex path, pattern
exit_code Command exits with code command, exit_code
output_contains Command output contains string command, contains
output_matches Command output matches regex command, pattern
no_errors No error events in session
cost_below Total cost below threshold max_cost
trust_above Trust score above threshold min_score
custom Custom script returns 0 command

CLI Commands

# Discover scenario files in current directory
caged eval discover

# Load a suite file (creates suite + scenarios via API)
caged eval load .caged-eval.yaml

# List all scenarios
caged eval list

# Run a specific scenario
caged eval run <scenario-id>

# Run flakiness test (10 iterations)
caged eval flakiness <scenario-id> --iterations 10

# Detect regression against last 5 runs
caged eval regression <scenario-id>

# View run history
caged eval runs <scenario-id>

Flakiness Detection

Run a scenario N times and measure consistency:

caged eval flakiness <id> --iterations 10

Returns:

  • Pass rate (e.g. 8/10 = 80%)
  • Per-assertion flakiness breakdown
  • Verdict: stable, flaky, or broken

Regression Detection

Compare current run against baseline (last N runs):

caged eval regression <id>

Detects:

  • Score drops (>10% threshold)
  • New failing assertions
  • Cost increases (>20% threshold)

CI/CD Integration

Add to your CI pipeline:

# GitHub Actions
- name: Run agent evals
  run: |
    caged eval load .caged-eval.yaml
    caged eval run --suite $SUITE_ID --fail-on-regression

API

# Create scenario
curl -X POST -H "Authorization: Bearer $TOKEN" \
  -d '{"name":"test","template":"node-20","prompt":"...","assertions":[...]}' \
  https://api.caged.dev/v1/eval/scenarios

# Run scenario
curl -X POST -H "Authorization: Bearer $TOKEN" \
  https://api.caged.dev/v1/eval/scenarios/{id}/run

# Get run results
curl -H "Authorization: Bearer $TOKEN" \
  https://api.caged.dev/v1/eval/runs/{run-id}
Was this page helpful?