Eval-Gated Pipeline Promotion
Eval-gated pipelines let you enforce quality gates before code progresses through your CI/CD workflow. Run Cage Eval scenarios as pipeline stages and automatically block promotion when tests fail or regress.
Overview
An eval gate is a pipeline stage that:
- Runs one or more eval scenarios against your agent
- Compares results against configurable thresholds
- Detects regressions vs. historical baselines
- Blocks or allows pipeline progression based on results
Stage Types
eval Stage
Runs a Cage Eval scenario or suite within the pipeline.
stages:
- name: run-tests
type: eval
config:
scenario_name: "fix-bug-scenario" # OR scenario_id: uuid
min_pass_rate: 0.8 # 80% assertions must pass
max_cost_usd: 0.50 # Budget cap
fail_on_regression: true
baseline_count: 5 # Compare against last 5 runs
regression_threshold: 0.1 # 10% max score drop
Configuration options:
| Field | Type | Description |
|---|---|---|
scenario_id |
UUID | ID of the scenario to run |
scenario_name |
string | Name of the scenario (alternative to ID) |
suite_id |
UUID | Run all scenarios in a suite |
suite_name |
string | Suite name (alternative to ID) |
min_pass_rate |
float | Minimum pass rate (0-1) to pass gate |
max_cost_usd |
float | Maximum allowed cost for the eval run |
fail_on_regression |
bool | Block if score regresses vs baseline |
baseline_count |
int | Number of previous runs for baseline (default: 5) |
regression_threshold |
float | Max allowed score drop (default: 0.1 = 10%) |
store_result_in_state |
bool | Save eval result to pipeline state |
result_state_key |
string | State key for result (default: "eval_result") |
gate Stage
Evaluates trust score and cost thresholds.
stages:
- name: quality-gate
type: gate
config:
trust_above: 70 # Session trust score >= 70
cost_below: 1.00 # Total run cost < $1.00
Configuration options:
| Field | Type | Description |
|---|---|---|
trust_above |
int | Minimum trust score (0-100) to pass |
cost_below |
float | Maximum cost in USD to pass |
expression |
string | Custom CEL expression (coming soon) |
state_key |
string | State key for dynamic evaluation |
fail_fast |
bool | Stop pipeline immediately on failure |
Regression Detection
When fail_on_regression: true is set, the eval gate:
- Fetches the last N passing runs (configurable via
baseline_count) - Calculates the baseline average score
- Compares the current run's score against the baseline
- Fails if the score dropped more than
regression_threshold
Example:
- Baseline (last 5 runs): 95%, 92%, 98%, 94%, 96% → average: 95%
- Current run: 82%
- Delta: -13% (exceeds 10% threshold)
- Result: Gate fails with regression error
Pipeline Example
name: agent-ci
version: 1
stages:
- name: build
type: command
command: "npm run build"
- name: unit-tests
type: command
command: "npm test"
depends_on: [build]
- name: eval-agent
type: eval
depends_on: [unit-tests]
config:
scenario_name: "code-review-task"
min_pass_rate: 0.9
fail_on_regression: true
baseline_count: 10
regression_threshold: 0.05 # 5% max regression
- name: quality-check
type: gate
depends_on: [eval-agent]
config:
trust_above: 80
cost_below: 2.00
- name: deploy
type: command
command: "npm run deploy"
depends_on: [quality-check]
Conditional Branching
Use on_success and on_failure conditions to branch based on gate results:
stages:
- name: eval-agent
type: eval
config:
scenario_name: "main-workflow"
min_pass_rate: 0.8
- name: deploy-prod
type: command
command: "deploy --env=production"
depends_on: [eval-agent]
condition:
on_success: true # Only runs if eval passed
- name: notify-failure
type: command
command: "slack-notify --channel=alerts"
depends_on: [eval-agent]
condition:
on_failure: true # Only runs if eval failed
Gate Results
When a gate blocks a pipeline, the gate_result field on the stage provides details:
{
"passed": false,
"failed_checks": [
{
"type": "min_pass_rate",
"expected": "0.80",
"actual": "0.65",
"passed": false,
"message": "Pass rate 65.0% below required 80.0%"
},
{
"type": "regression",
"expected": ">= 85.0% (baseline)",
"actual": "65.0%",
"passed": false,
"message": "Score regressed 23.5% vs baseline (85.0% → 65.0%)"
}
],
"passed_checks": [],
"message": "Eval gate failed: 2 check(s) failed",
"score": 0.65,
"baseline_score": 0.85,
"is_regression": true,
"regression_pct": 23.5
}
Best Practices
1. Start with Lenient Thresholds
Begin with conservative thresholds and tighten as your agent improves:
# Week 1: Establish baseline
min_pass_rate: 0.5
regression_threshold: 0.2 # 20%
# Week 4: Tighten as agent stabilizes
min_pass_rate: 0.8
regression_threshold: 0.1 # 10%
# Production: Strict enforcement
min_pass_rate: 0.95
regression_threshold: 0.05 # 5%
2. Use Suites for Comprehensive Testing
Run multiple scenarios as a suite for broader coverage:
- name: full-eval
type: eval
config:
suite_name: "agent-capabilities"
min_pass_rate: 0.85
3. Set Cost Budgets
Prevent runaway costs with budget gates:
- name: cost-check
type: gate
config:
cost_below: 5.00 # $5 max per pipeline run
4. Combine with Trust Scoring
Use trust gates to ensure agent behavior stays safe:
- name: trust-gate
type: gate
depends_on: [eval-agent]
config:
trust_above: 75 # Block if risky behavior detected
Recipe: "Don't Merge Until Agent Tests Pass"
Block PRs from merging until the agent passes its own test suite:
# .caged/pipelines/pr-check.yaml
name: pr-agent-check
version: 1
stages:
- name: checkout
type: command
command: "git checkout $CAGED_PR_BRANCH"
- name: run-agent-tests
type: eval
depends_on: [checkout]
config:
suite_name: "pr-validation-suite"
min_pass_rate: 1.0 # All tests must pass
fail_on_regression: true
regression_threshold: 0 # No regression allowed
max_cost_usd: 1.00
- name: update-pr-status
type: command
depends_on: [run-agent-tests]
command: |
gh pr review $CAGED_PR_NUMBER --approve \
--body "✅ Agent tests passed (score: $(cat eval_result.json | jq .score))"
condition:
on_success: true
- name: block-pr
type: command
depends_on: [run-agent-tests]
command: |
gh pr review $CAGED_PR_NUMBER --request-changes \
--body "❌ Agent tests failed. See pipeline for details."
condition:
on_failure: true
API Reference
Get Stage Gate Result
GET /v1/pipelines/runs/{run_id}/stages/{stage_id}
Response includes gate_result for eval and gate stages.
List Pipeline Runs with Gate Failures
GET /v1/pipelines/{pipeline_id}/runs?status=failed&has_gate_failure=true
Dashboard
The pipeline dashboard shows:
- Which gate blocked each run
- Pass/fail checks with details
- Regression trends over time
- Cost accumulation per stage
Navigate to Pipelines → [Your Pipeline] → Runs to see gate results inline with each stage.