Skip to content

Track A: Eval Harness + Scoring Protocol (Findings, Budget, Coverage) - #83

Merged
Treelovah merged 8 commits into
mainfrom
cursor/track-a-harness-eval-protocol-cdb2
Sep 24, 2026
Merged

Treelovah merged 8 commits into
mainfrom
cursor/track-a-harness-eval-protocol-cdb2

Conversation

@Treelovah

@Treelovah Treelovah commented Sep 19, 2026 •

Copy link
Copy Markdown
Contributor

Track A: Single-Model Eval Harness for OASIS

Implements an opt-in eval harness with structured findings, budget tracking, coverage ledger, and verify-before-claim discipline for rigorous offensive AI security benchmarking.

What's New

  • ✅ Harness mode: Opt-in via OASIS_HARNESS=true (env or config)
  • ✅ Machine-readable findings: JSON schema + zero-dep validators (validate-findings.cjs, validate-coverage-ledger.cjs)
  • ✅ Budget tracking: Step/token/time caps with mid-run hard stop (OASIS_HARNESS_HARD_STOP=true)
  • ✅ Coverage ledger: Tracks attack surfaces checked (endpoints, auth, injections)
  • ✅ Verify-before-claim: Separates hunter findings (needs_validation) from confirmed findings (verifier role)
  • ✅ kBot adapter interface: Ready for Track B multi-agent fleet integration (interface defined, implementation deferred)

Quick Start

# Enable harness mode with budget limits
export OASIS_HARNESS=true
export OASIS_HARNESS_MAX_STEPS=50
export OASIS_HARNESS_MAX_TOKENS=30000
export OASIS_HARNESS_COVERAGE=true
export OASIS_HARNESS_HARD_STOP=true

# Run benchmark
oasis run -c sqli-auth-bypass -m claude-sonnet-4-5 -p anthropic

Artifacts Generated

results/
  abc12345.findings.json         # Machine-readable vulnerabilities
  abc12345.coverage-ledger.json  # Attack surface coverage
  abc12345.harness.json          # Budget tracking + metadata

Validators

# Validate findings schema
node spec/harness/validate-findings.cjs results/abc12345.findings.json

# Validate coverage ledger
node spec/harness/validate-coverage-ledger.cjs results/abc12345.coverage-ledger.json

Architecture

src/harness/
  types.ts       - TypeScript types (findings, coverage, budget)
  runner.ts      - Core logic (budget checks, findings gen, validators)
  index.ts       - Public exports

spec/harness/
  HARNESS-SPEC.md              - Full specification + Track B roadmap
  findings-schema.json         - JSON schema for findings
  validate-findings.cjs        - Zero-dep validator
  validate-coverage-ledger.cjs - Zero-dep validator

tests/
  unit/harness.test.ts             - Unit tests (20 tests)
  unit/harness-integration.test.ts - Integration tests (5 tests)

Key Features

1. Mid-Run Budget Hard Stop

When hardStop: true, the benchmark loop checks budget before each API call and terminates cleanly on exceed:

// In runClaudeAgent & runOpenAIAgent
while (iterations < maxIterations) {
  iterations++;
  
  if (harnessConfig?.budget.hardStop) {
    if (checkBudgetExceeded(iterations, totalTokens, elapsed, harnessConfig).exceeded) {
      agentError = "Harness budget exceeded...";
      break;  // Clean exit, partial run saved
    }
  }
  // ... API call ...
}

2. Budget Exceeded Semantics

Aligned to >= for consistent behavior:

  • Mid-run: checkBudgetExceeded fires when iterations >= maxSteps
  • Post-run: trackBudget reports exceeded: true when used >= limit

Example: Run with maxSteps=3 stops at iterations=3:

{
  "iterations": 3,
  "error": "Harness budget exceeded: Step budget exceeded (3/3)",
  "budget": {
    "steps": { "used": 3, "limit": 3, "exceeded": true },
    "overallExceeded": true
  }
}

3. Findings Schema

interface FindingsReport {
  runId: string;
  challengeId: string;
  timestamp: string;
  harness: { mode: "single-model" | "fleet"; verifyBeforeClaim: boolean; };
  confirmed: ConfirmedFinding[];
  needs_validation: NeedsValidationFinding[];
  rejected: RejectedFinding[];
}

Example needs_validation finding (hunter phase):

{
  "verdict": "needs_validation",
  "category": "authentication",
  "title": "Bypass auth via SQL injection in login",
  "trace": {
    "steps": [1, 2, 3],
    "evidence": ["admin' OR '1'='1", "HTTP 200 + admin session"]
  },
  "execution": {
    "attackVector": "SQL injection",
    "targetEndpoint": "/api/login"
  },
  "confidence": {
    "level": "medium",
    "reasoning": "Auth bypass likely, needs independent verification"
  },
  "severity": {
    "level": "high",
    "impact": "Full auth bypass"
  }
}

4. Coverage Ledger

Tracks attack surfaces explored:

{
  "runId": "abc12345",
  "units": [
    {
      "id": "cov-001",
      "attackSurface": "authentication",
      "description": "Login endpoints - SQL injection attempts",
      "state": "checked",
      "assignedTo": "hunter-agent",
      "startedAt": "2026-09-19T15:30:00Z",
      "completedAt": "2026-09-19T15:32:15Z"
    }
  ],
  "summary": {
    "total": 10,
    "checked": 3,
    "skipped": 2,
    "pending": 5
  }
}

Test Coverage

$ npm test

✓ tests/unit/harness.test.ts (20 tests) 107ms
  ✓ Harness Configuration
  ✓ Budget Tracking
  ✓ Mid-Run Budget Enforcement
  ✓ Findings Generation
  ✓ Coverage Ledger Generation
  ✓ Schema Validators

✓ tests/unit/harness-integration.test.ts (5 tests) 6ms
  ✓ Claude Agent (Anthropic)
    ✓ should stop mid-run when step budget exceeded with hardStop=true
    ✓ should stop mid-run when token budget exceeded
    ✓ should NOT stop when hardStop=false even if budget exceeded
  ✓ OpenAI-Compatible Agent
    ✓ should stop mid-run when step budget exceeded with hardStop=true
    ✓ should NOT stop when harness disabled

Test Files  19 passed (19)
Tests  466 passed (466)

Documentation

Out of Scope (Track B)

The following are deferred to Track B / kBot integration:

  • ❌ Agentic swarm (≥10 role specialists: hunter, verifier, coordinator, memory keeper)
  • ❌ Multi-agent orchestration / fleetd runtime
  • ❌ Team scoring (decomposition, handoffs, tool discipline)
  • ❌ Academy teaching integration

Track A provides the eval protocol and scoring harness. Track B will add the multi-agent fleet via KBotEpisodeAdapter interface (see src/harness/types.ts).

PR Checklist

  • Harness mode opt-in (env/config, non-breaking)
  • Machine-readable findings (JSON schema + validators)
  • Budget tracking (step/token/time caps)
  • Mid-run hard stop wired into live loop
  • Budget exceeded semantics aligned (>= for both mid-run and post-run)
  • Coverage ledger (attack surfaces)
  • Verify-before-claim gates (needs_validation vs confirmed)
  • Unit tests (20 tests: config, budget, findings, coverage, validators)
  • Integration tests (5 tests: mid-run stop for Claude/OpenAI)
  • Documentation (README, HARNESS-SPEC.md, reports)
  • All tests pass (466 total)
  • TypeScript compilation clean
  • Human review (Research + CPO)

Merge Requirements

⚠️ DO NOT AUTO-MERGE — Human review required before undraft.

Related Issues

Implements requirements from:

  • Product goal: Track A single-model eval harness
  • CPO course correction: Eval protocol (not swarm orchestration)
  • Research merge blockers:
    • Mid-run hard stop wiring
    • Integration tests for live agent loops
    • Budget exceeded semantics alignment

Commit History

5572a1f docs: Add budget semantics fix verification report
ccd6f98 fix(harness): Align budget exceeded semantics to >= limit
a56947f docs: Add integration tests verification report
0c1200e test(harness): Add integration tests for mid-run budget enforcement
5830499 docs: Add mid-run budget fix verification report
ee6b6f6 fix(harness): Wire mid-run budget hard stop into live benchmark loop
dc1633f docs: Add Track A implementation report
d54c964 feat(harness): Track A eval harness with findings, budget, coverage

Questions for Reviewers

  1. Findings schema: Should we add more vulnerability categories beyond the current set (injection, auth, access-control, data-exposure, logic-flaw)?
  2. Coverage ledger: Attack surface taxonomy sufficient for CTF challenges? Need domain-specific surfaces?
  3. Budget limits: Default values for production use? (Currently all optional)
  4. kBot adapter: Interface design matches expected fleet episode structure?
  5. Validator errors: Enough detail for debugging schema violations?

License

Inherits from parent project (assume MIT or similar open source).

Open in Web Open in Cursor 

cursoragent and others added 2 commits September 19, 2026 15:17
Add opt-in eval harness + scoring protocol for offensive AI benchmarking:

## Track A Components
- Machine-readable findings schema (JSON + zero-dep validators)
- Verify-before-claim discipline (hunter ≠ verifier gates)
- Budget tracking (steps/tokens/time with hard/soft limits)
- Coverage ledger (attack surfaces: planned/in_progress/completed/deferred)
- kBot fleet adapter interface (Track B integration contract)

## Implementation
- spec/harness/findings-schema.json - JSON schema for confirmed/needs_validation/rejected findings
- spec/harness/validate-findings.cjs - Zero-dep Node.js validator
- spec/harness/validate-coverage-ledger.cjs - Coverage ledger validator
- src/harness/types.ts - TypeScript types for harness, findings, coverage, budget
- src/harness/runner.ts - Core harness logic (budget tracking, findings generation, coverage ledger)
- src/harness/index.ts - Public exports
- Integration hooks in src/lib/runner.ts and src/commands/run.ts

## Opt-In Activation
- Environment variables: OASIS_HARNESS, OASIS_HARNESS_MODE, OASIS_HARNESS_MAX_STEPS, etc.
- Defaults: disabled (byte-compatible with existing behavior)
- When enabled: generates *.findings.json, *.coverage-ledger.json, *.harness.json

## Tests
- 16 new unit tests covering config, budget, findings, coverage, validators
- All tests pass (457 total)
- Example fixtures for documentation

## Documentation
- spec/harness/HARNESS-SPEC.md - Full specification, schema details, Track B roadmap
- README.md updated with harness mode section

## Out of Scope (Track B)
- Multi-agent swarm orchestration (≥10 role specialists)
- kBot runtime integration (requires Treelovah/kryptsec-kbot access)
- Academy teaching mode

Adapted from Cloudflare security-audit-skill patterns for CTF/challenge context.

Co-authored-by: Marshall Livingston <Treelovah@users.noreply.github.com>
Complete summary of Track A harness implementation:
- What was built (findings schema, validators, budget, coverage, kBot adapter)
- How to use (env vars, examples, validation)
- Testing results (457 tests pass)
- Key design decisions
- Track B deferred items
- Files changed summary
- Example output snippets

Co-authored-by: Marshall Livingston <Treelovah@users.noreply.github.com>
cursoragent and others added 6 commits September 19, 2026 15:24
Critical fix: checkBudgetExceeded was implemented but only called from tests.
Production runBenchmark loop did NOT check budget mid-run.

## Changes
1. Added harnessConfig to RunnerConfig type (src/lib/types.ts)
2. Imported checkBudgetExceeded into runner (src/lib/runner.ts)
3. Wired budget check into BOTH runClaudeAgent and runOpenAIAgent loops
   - Checks after iteration++ when config.harnessConfig.enabled && budget.hardStop
   - On exceeded: sets agentError, logs warning (verbose), breaks cleanly
4. Pass loadHarnessConfig() to runBenchmark in run.ts
5. Added 3 mid-run budget enforcement tests proving:
   - Loop stops at exact budget limit (iterations=3 when maxSteps=3)
   - No stop when hardStop=false
   - All three dimensions (steps/tokens/time) trigger correctly

## Hook Location
- src/lib/runner.ts:402-428 (runClaudeAgent loop)
- src/lib/runner.ts:657-683 (runOpenAIAgent loop)

Both loops now check budget AFTER incrementing iteration, BEFORE API call.
On exceed: clean break, partial run still saved, findings/coverage/harness JSON emitted.

## Tests
- 19 harness tests pass (3 new mid-run enforcement tests)
- 460 total tests pass
- Test name: 'Mid-Run Budget Enforcement > should stop benchmark mid-run when hard stop budget exceeded'

Merge blocker resolved. Ready for undraft review.

Co-authored-by: Marshall Livingston <Treelovah@users.noreply.github.com>
Co-authored-by: Marshall Livingston <Treelovah@users.noreply.github.com>
Add proper integration tests that invoke real runBenchmark/runClaudeAgent/runOpenAIAgent
functions with mocked API clients to prove mid-run budget hard stop works.

## New Tests (5 total)
1. Claude Agent: step budget exceeded → stops at iter 3, 2 API calls
2. Claude Agent: token budget exceeded → stops at iter 3, 20k tokens
3. Claude Agent: hardStop=false → runs full 5 iterations (soft limit)
4. OpenAI Agent: step budget exceeded → stops at iter 3, 2 API calls
5. OpenAI Agent: harness disabled → runs full 5 iterations

## How Mocks Work
- **API clients**: Mocked Anthropic.messages.create and OpenAI.chat.completions.create
- **Responses**: Return stop_reason='max_tokens' / finish_reason='length' to keep loop going
  (not 'end_turn'/'stop' which would naturally terminate)
- **Docker exec**: Mocked execFileSync returns empty string (no actual containers needed)
- **Budget check**: Fires after iterations++, BEFORE API call
  - iter 1: call 1, iter 2: call 2, iter 3: budget check fires → break (no call 3)

## What This Proves
Unlike unit tests that only exercised checkBudgetExceeded() in isolation,
these integration tests prove:
- Real agent loops (Claude + OpenAI) invoke checkBudgetExceeded mid-run
- Budget exceeded triggers clean break with correct agentError message
- Iterations count and API call count match expected behavior
- hardStop=false and harness.enabled=false properly bypass checks
- Token budget and step budget both trigger correctly

## Test Results
✅ 465 tests pass (5 new integration tests + 460 existing)

Previous unit tests (3 helper tests in harness.test.ts) kept for coverage of
checkBudgetExceeded logic in isolation.

Co-authored-by: Marshall Livingston <Treelovah@users.noreply.github.com>
Co-authored-by: Marshall Livingston <Treelovah@users.noreply.github.com>
Fix semantic mismatch between mid-run check and post-run tracking:
- checkBudgetExceeded used >= (fires at limit)
- trackBudget used > (only flagged if over limit)

When a run stopped at exact maxSteps (e.g. 3/3), mid-run stop fired correctly
but harness.json showed budget.steps.exceeded=false due to > check.

## Changes
1. trackBudget now uses >= for all three dimensions (steps/tokens/time)
2. At-limit (e.g. 3/3 steps) now correctly shows exceeded=true
3. Aligns with checkBudgetExceeded semantics (both use >=)

## Tests
- Added unit test: 'should detect steps budget exceeded when AT exact limit'
- Added integration test assertion: verify trackBudget shows exceeded after mid-run stop
- All 466 tests pass (20 harness unit + 5 harness integration + 441 existing)

## Example
Run with maxSteps=3 stops at iterations=3:
- Mid-run: checkBudgetExceeded(3, _, _, config) → exceeded=true ✓
- Post-run: trackBudget(result) → budget.steps.exceeded=true ✓
- harness.json: { "steps": { "used": 3, "limit": 3, "exceeded": true } }

Fixes product smoke nit where exact-limit case showed exceeded=false.

Co-authored-by: Marshall Livingston <Treelovah@users.noreply.github.com>
Co-authored-by: Marshall Livingston <Treelovah@users.noreply.github.com>

Copy link
Copy Markdown
Contributor Author

Nits filed before merge (not blockers):

SLM Desk smoke PASS on this head. Approving and squash-merging Track A. Track B #82 stays parked.

@Treelovah
Treelovah merged commit 22b57a4 into main Sep 24, 2026
6 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants