Track A: Eval Harness + Scoring Protocol (Findings, Budget, Coverage) - #83
Merged
Merged
Conversation
Add opt-in eval harness + scoring protocol for offensive AI benchmarking: ## Track A Components - Machine-readable findings schema (JSON + zero-dep validators) - Verify-before-claim discipline (hunter ≠ verifier gates) - Budget tracking (steps/tokens/time with hard/soft limits) - Coverage ledger (attack surfaces: planned/in_progress/completed/deferred) - kBot fleet adapter interface (Track B integration contract) ## Implementation - spec/harness/findings-schema.json - JSON schema for confirmed/needs_validation/rejected findings - spec/harness/validate-findings.cjs - Zero-dep Node.js validator - spec/harness/validate-coverage-ledger.cjs - Coverage ledger validator - src/harness/types.ts - TypeScript types for harness, findings, coverage, budget - src/harness/runner.ts - Core harness logic (budget tracking, findings generation, coverage ledger) - src/harness/index.ts - Public exports - Integration hooks in src/lib/runner.ts and src/commands/run.ts ## Opt-In Activation - Environment variables: OASIS_HARNESS, OASIS_HARNESS_MODE, OASIS_HARNESS_MAX_STEPS, etc. - Defaults: disabled (byte-compatible with existing behavior) - When enabled: generates *.findings.json, *.coverage-ledger.json, *.harness.json ## Tests - 16 new unit tests covering config, budget, findings, coverage, validators - All tests pass (457 total) - Example fixtures for documentation ## Documentation - spec/harness/HARNESS-SPEC.md - Full specification, schema details, Track B roadmap - README.md updated with harness mode section ## Out of Scope (Track B) - Multi-agent swarm orchestration (≥10 role specialists) - kBot runtime integration (requires Treelovah/kryptsec-kbot access) - Academy teaching mode Adapted from Cloudflare security-audit-skill patterns for CTF/challenge context. Co-authored-by: Marshall Livingston <Treelovah@users.noreply.github.com>
Complete summary of Track A harness implementation: - What was built (findings schema, validators, budget, coverage, kBot adapter) - How to use (env vars, examples, validation) - Testing results (457 tests pass) - Key design decisions - Track B deferred items - Files changed summary - Example output snippets Co-authored-by: Marshall Livingston <Treelovah@users.noreply.github.com>
Critical fix: checkBudgetExceeded was implemented but only called from tests. Production runBenchmark loop did NOT check budget mid-run. ## Changes 1. Added harnessConfig to RunnerConfig type (src/lib/types.ts) 2. Imported checkBudgetExceeded into runner (src/lib/runner.ts) 3. Wired budget check into BOTH runClaudeAgent and runOpenAIAgent loops - Checks after iteration++ when config.harnessConfig.enabled && budget.hardStop - On exceeded: sets agentError, logs warning (verbose), breaks cleanly 4. Pass loadHarnessConfig() to runBenchmark in run.ts 5. Added 3 mid-run budget enforcement tests proving: - Loop stops at exact budget limit (iterations=3 when maxSteps=3) - No stop when hardStop=false - All three dimensions (steps/tokens/time) trigger correctly ## Hook Location - src/lib/runner.ts:402-428 (runClaudeAgent loop) - src/lib/runner.ts:657-683 (runOpenAIAgent loop) Both loops now check budget AFTER incrementing iteration, BEFORE API call. On exceed: clean break, partial run still saved, findings/coverage/harness JSON emitted. ## Tests - 19 harness tests pass (3 new mid-run enforcement tests) - 460 total tests pass - Test name: 'Mid-Run Budget Enforcement > should stop benchmark mid-run when hard stop budget exceeded' Merge blocker resolved. Ready for undraft review. Co-authored-by: Marshall Livingston <Treelovah@users.noreply.github.com>
Co-authored-by: Marshall Livingston <Treelovah@users.noreply.github.com>
Add proper integration tests that invoke real runBenchmark/runClaudeAgent/runOpenAIAgent functions with mocked API clients to prove mid-run budget hard stop works. ## New Tests (5 total) 1. Claude Agent: step budget exceeded → stops at iter 3, 2 API calls 2. Claude Agent: token budget exceeded → stops at iter 3, 20k tokens 3. Claude Agent: hardStop=false → runs full 5 iterations (soft limit) 4. OpenAI Agent: step budget exceeded → stops at iter 3, 2 API calls 5. OpenAI Agent: harness disabled → runs full 5 iterations ## How Mocks Work - **API clients**: Mocked Anthropic.messages.create and OpenAI.chat.completions.create - **Responses**: Return stop_reason='max_tokens' / finish_reason='length' to keep loop going (not 'end_turn'/'stop' which would naturally terminate) - **Docker exec**: Mocked execFileSync returns empty string (no actual containers needed) - **Budget check**: Fires after iterations++, BEFORE API call - iter 1: call 1, iter 2: call 2, iter 3: budget check fires → break (no call 3) ## What This Proves Unlike unit tests that only exercised checkBudgetExceeded() in isolation, these integration tests prove: - Real agent loops (Claude + OpenAI) invoke checkBudgetExceeded mid-run - Budget exceeded triggers clean break with correct agentError message - Iterations count and API call count match expected behavior - hardStop=false and harness.enabled=false properly bypass checks - Token budget and step budget both trigger correctly ## Test Results ✅ 465 tests pass (5 new integration tests + 460 existing) Previous unit tests (3 helper tests in harness.test.ts) kept for coverage of checkBudgetExceeded logic in isolation. Co-authored-by: Marshall Livingston <Treelovah@users.noreply.github.com>
Co-authored-by: Marshall Livingston <Treelovah@users.noreply.github.com>
Fix semantic mismatch between mid-run check and post-run tracking:
- checkBudgetExceeded used >= (fires at limit)
- trackBudget used > (only flagged if over limit)
When a run stopped at exact maxSteps (e.g. 3/3), mid-run stop fired correctly
but harness.json showed budget.steps.exceeded=false due to > check.
## Changes
1. trackBudget now uses >= for all three dimensions (steps/tokens/time)
2. At-limit (e.g. 3/3 steps) now correctly shows exceeded=true
3. Aligns with checkBudgetExceeded semantics (both use >=)
## Tests
- Added unit test: 'should detect steps budget exceeded when AT exact limit'
- Added integration test assertion: verify trackBudget shows exceeded after mid-run stop
- All 466 tests pass (20 harness unit + 5 harness integration + 441 existing)
## Example
Run with maxSteps=3 stops at iterations=3:
- Mid-run: checkBudgetExceeded(3, _, _, config) → exceeded=true ✓
- Post-run: trackBudget(result) → budget.steps.exceeded=true ✓
- harness.json: { "steps": { "used": 3, "limit": 3, "exceeded": true } }
Fixes product smoke nit where exact-limit case showed exceeded=false.
Co-authored-by: Marshall Livingston <Treelovah@users.noreply.github.com>
Co-authored-by: Marshall Livingston <Treelovah@users.noreply.github.com>
Treelovah
marked this pull request as ready for review
September 19, 2026 15:38
This was referenced Sep 24, 2026
Contributor
Author
|
Nits filed before merge (not blockers):
SLM Desk smoke PASS on this head. Approving and squash-merging Track A. Track B #82 stays parked. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Track A: Single-Model Eval Harness for OASIS
Implements an opt-in eval harness with structured findings, budget tracking, coverage ledger, and verify-before-claim discipline for rigorous offensive AI security benchmarking.
What's New
OASIS_HARNESS=true(env or config)validate-findings.cjs,validate-coverage-ledger.cjs)OASIS_HARNESS_HARD_STOP=true)needs_validation) from confirmed findings (verifier role)Quick Start
Artifacts Generated
Validators
Architecture
Key Features
1. Mid-Run Budget Hard Stop
When
hardStop: true, the benchmark loop checks budget before each API call and terminates cleanly on exceed:2. Budget Exceeded Semantics
Aligned to
>=for consistent behavior:checkBudgetExceededfires wheniterations >= maxStepstrackBudgetreportsexceeded: truewhenused >= limitExample: Run with
maxSteps=3stops atiterations=3:{ "iterations": 3, "error": "Harness budget exceeded: Step budget exceeded (3/3)", "budget": { "steps": { "used": 3, "limit": 3, "exceeded": true }, "overallExceeded": true } }3. Findings Schema
Example
needs_validationfinding (hunter phase):{ "verdict": "needs_validation", "category": "authentication", "title": "Bypass auth via SQL injection in login", "trace": { "steps": [1, 2, 3], "evidence": ["admin' OR '1'='1", "HTTP 200 + admin session"] }, "execution": { "attackVector": "SQL injection", "targetEndpoint": "/api/login" }, "confidence": { "level": "medium", "reasoning": "Auth bypass likely, needs independent verification" }, "severity": { "level": "high", "impact": "Full auth bypass" } }4. Coverage Ledger
Tracks attack surfaces explored:
{ "runId": "abc12345", "units": [ { "id": "cov-001", "attackSurface": "authentication", "description": "Login endpoints - SQL injection attempts", "state": "checked", "assignedTo": "hunter-agent", "startedAt": "2026-09-19T15:30:00Z", "completedAt": "2026-09-19T15:32:15Z" } ], "summary": { "total": 10, "checked": 3, "skipped": 2, "pending": 5 } }Test Coverage
Documentation
Out of Scope (Track B)
The following are deferred to Track B / kBot integration:
Track A provides the eval protocol and scoring harness. Track B will add the multi-agent fleet via
KBotEpisodeAdapterinterface (seesrc/harness/types.ts).PR Checklist
needs_validationvsconfirmed)Merge Requirements
Related Issues
Implements requirements from:
Commit History
Questions for Reviewers
License
Inherits from parent project (assume MIT or similar open source).