Conversation
Seeded with the one thing 1.5.0's own release turned up: the release workflow's last job could not run, and its dry-run mode could not have caught that. `github-release` is gated on `refs/tags/v*`, so a `workflow_dispatch` from a branch skips it -- the rehearsal path structurally excludes the only job that had never run. That is the item: a dispatch that exercises `gh release create` against real artifacts and the real token, and a deliberately broken step that fails the rehearsal rather than the next tag. Also records four things noticed during 1.5.0 and deliberately left alone, so they are not lost: the tracked judge usage database that churns on every local run, Playwright adopting a stray server as the system under test, Carbon's always-mounted modal, and the docs directory being discovered rather than configured. None investigated beyond the observation. The attic README points at the open plan again. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Signed-off-by: Oktie Hassanzadeh <oktie@oktie.com>
…boxed The question and db_id were set at 0.875rem in secondary text, smaller and dimmer than the body copy around them, with the db_id run onto the end of the same line. They are the subject of the whole view and were the hardest thing on it to locate. The question now has its own block: a QUESTION label, the text at 1.25rem in primary colour, and the database as a labelled tag rather than an inline fragment. Tinted with layer-02 and not layer-01 -- inside this content layer the latter resolves to the same white as the page, so the first attempt had a background in the code and none on screen. Ground truth and predicted SQL each get a box. A bare Carbon text area carries only an underline, so two side by side read as one undivided region instead of as the two queries being compared. The border matches the panels below them. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Signed-off-by: Oktie Hassanzadeh <oktie@oktie.com>
The detail panels show a record read-only -- the question, both queries, the results, the metrics. The playground is where that same record can be edited and re-run. Getting between them meant reading the record id off the address bar and assembling a /run/... URL by hand, which is a thing nobody does twice. One button, in both places that open a record detail: Error Analysis and a pipeline's detail view. It carries the benchmark, the record and the pipeline being read, so the playground opens on the prediction you were just looking at rather than picking its own default. An anchor, not a button with a click handler. The playground address is worth copying and worth opening in a new tab, and only a plain left click is intercepted for single-page navigation -- every other gesture goes back to the browser. Without an onNavigate the control stays a working link and the browser navigates normally, so a caller that forgets to wire it degrades rather than breaks. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Signed-off-by: Oktie Hassanzadeh <oktie@oktie.com>
Picking a judge config by name said nothing about what it would ask. The model it runs and the prompt it sends are the whole of what a judge is, and finding either meant leaving for the config editor and coming back. The box names the config, tags its model, and puts the YAML behind a Show prompt toggle. Collapsed by default: the prompt template is the bulk of every config and runs to forty-odd lines, which would push the run controls and the verdict off the screen for a reader who only wanted to know which model was judging. A `pre`, not the CodeMirror editor the config view uses. That editor's chunk is 416 KB against this view's 44 KB -- a great deal of syntax highlighting for a read-only prompt template. This adds 4 KB. Also fixes a sentence I broke adding Judge again: JSX drops a newline between an element and the text after it, so "Judge again ignores the cache" rendered as "Judge againignores the cache". And the model tag carries its full value as a title, since Carbon truncates a long provider:model id. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Signed-off-by: Oktie Hassanzadeh <oktie@oktie.com>
…never closed is_stale() and every JudgeStore method opened SQLite with `with sqlite3.connect(...)`, which commits but does not close. A connection refers to itself through its statement cache, so each stayed open until the cyclic collector ran: six descriptors per landing-page load, one per /api/me from a signed-in caller. On the public deployment that reached Docker's default limit of 1024 and the app could no longer open the results it serves. Both now close what they open. An index also reaps the connections of worker threads that have exited, a smaller leak found on the way. Compose raises the app's nofile limit to 65536, with a CI check, and the deployment guide gains a troubleshooting row. Recorded as item 2 of the 1.6.0 plan and in the project log. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Signed-off-by: Oktie Hassanzadeh <oktie@oktie.com>
Items 2 to 5 and the items carried over from 1.5.0, as they stood before work on item 5 began: the file-descriptor exhaustion on the deployment, Archer's LLM-judge results, gpt-oss-120b baselines and judge configs, and keeping Beaver behind a sign-in. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Signed-off-by: Oktie Hassanzadeh <oktie@oktie.com>
Plan item 5, in the dashboard. Every read route was public tier, and a tier cannot tell an anonymous visitor from a signed-in read_only user, so this is a check on identity, enforced in the tier middleware ahead of the tier gate. - A registry entry with "requires_sign_in": true (beaver, beaver_test_10) is hidden from anonymous callers on shared deployments. The flag is read from every registry copy, packaged ones included, because provision.sh never overwrites a data root's seeded benchmarks.json. TEXT2SQL_SIGN_IN_BENCHMARKS adds ids and is passed through compose. - Every route whose declared parameters name such a benchmark answers 401 with sign_in_required, the logo included; /api/benchmarks leaves it out. Ids are compared casefolded, and an id outside [A-Za-z0-9_-] is a 400 for a caller the wall applies to: /api/compare?left_id=charts/../beaver read Beaver's summary. - X-Robots-Tag: noindex, nofollow on any address that mentions it, API and app shell alike. - A Beaver link opened anonymously asks the reader to sign in and returns them to the same address; a signed-in listing tags it "Signed-in only". Editing a benchmark keeps its flag. tests/test_benchmark_access.py parametrizes over the live route table and fails on any route parameter not classified as naming a benchmark or not, so a new route cannot serve one unnoticed. The GitHub and Hugging Face copies of Beaver's data are untouched; how to handle them is still an open question in the plan. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Signed-off-by: Oktie Hassanzadeh <oktie@oktie.com>
archer_en_dev was the only benchmark with no judge scores at all, which is why its Metric Insights page showed no evidence under LLM judge comparison. Run on 2026-09-01 with llm_judge_default_config (wxai:meta-llama/llama-3-3-70b-instruct), the same model and prompt the existing summary records. 1,093 of 1,144 predictions scored: 543 Yes, 550 No. The other 51 are agentic predictions with no SQL (20 in baseline4, 31 in baseline5). The eval file keeps HEAD's formatting and key order, so its diff is the llm_score and llm_explanation lines and the trailing commas they need; the summary, errors report and charts are regenerated from it. archer_en_dev-predictions.json is unchanged. Plan item 3, step 1. Not yet on the Hub or the deployment. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Signed-off-by: Oktie Hassanzadeh <oktie@oktie.com>
…ign-in Beaver's questions, SQL and schema are distributed under gated access, and only what its own leaderboard shows -- overall scores -- may be published. The wall added the day before hid the whole benchmark from anonymous callers; it now inverts to what may be shown. - Anonymous callers get Beaver's tile, logo, overall summary, alias table and /api/compare. SUMMARY_ROUTES is an allowlist: every other route that names Beaver -- the breakdown by query category, error analysis, record detail, playground, insights, config, the judge -- answers 401 sign_in_required, and a route added later is refused until it is added on purpose. Ids outside [A-Za-z0-9_-] are still refused first. - /api/benchmarks lists Beaver for everyone with details_locked for anonymous callers. The summary page falls back to the overall table; every other Beaver view asks the reader to sign in, with a way back to the overall scores. - The repository follows the history purge: beaver_test_10 unregistered, Beaver's summary report cut to overall scores, its per-category charts untracked, local copies of its files git-ignored. The MySQL read-only grant takes MYSQL_READONLY_DATABASES instead of naming Beaver's databases, and the Beaver database docs point at its gated distribution. - Docs, changelog, plan item 5 and the project log record the decisions, the GitHub and Hugging Face purge, and what is still open: GitHub Support for PR refs and caches, forks and old clones, and a gated source for the deployment's Beaver details. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Signed-off-by: Oktie Hassanzadeh <oktie@oktie.com>
1.5.0's release published to PyPI and then failed on its last job, which had never run: the dispatch path that looked like a rehearsal skipped it with `if: startsWith(github.ref, 'refs/tags/v')`. Plan item 1. - The Release job now runs on workflow_dispatch. It uses the same `gh release create` step a tag uses, with --draft under a throwaway name, and a following step deletes the draft even if a later check fails, and fails the job if the delete does. - The build job extracts release notes on every run, for the version in pyproject.toml when there is no tag, and always uploads them. - A rehearsal's TestPyPI upload skips a version already there, so it can be re-run; a real publish still refuses one. - CONTRIBUTING.md says how to rehearse and what a rehearsal does not cover: the tag/version check, the pypi reviewer gate, PyPI itself. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Signed-off-by: Oktie Hassanzadeh <oktie@oktie.com>
…ark's details The upload script sent everything under results/ except logs. Run from a maintainer's checkout that would have published the derived query indices in results/.index/ -- which carry every record's raw bytes -- backups, and the local copies of Beaver's predictions, evaluation file and errors report, which are distributed under gated access. It now uploads an exact allow list. Hidden directories, bak/ and logs are never included, and a benchmark flagged requires_sign_in in any registry copy publishes only its overall summary (json, csv, md) and overall chart. The manifest lists only what is uploaded, and an upload whose restricted summary report still breaks results down by query category is refused. tests/test_upload_results_script.py pins the file list and the dry run's output. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Signed-off-by: Oktie Hassanzadeh <oktie@oktie.com>
The summaries for bird_mini_dev_postgres, bird_mini_dev_sqlite, spider_dev and spider_realistic listed gemini:gemini-3-flash-preview-greedy-zero-shot-chatapi, which none of those benchmarks' predictions or eval files contain. Plan item 3, step 3. The row is dropped from each summary JSON, with every other pipeline's values unchanged, and the CSV and overall chart are regenerated from it with the toolkit's own writers. The Markdown reports, rebuilt from the eval records, came out identical, and no errors report mentioned the pipeline. make_summary_report.py was not used: it rewrites the hand-written data/results/README.md. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Signed-off-by: Oktie Hassanzadeh <oktie@oktie.com>
The first rehearsal (run 34702494907) built, published to TestPyPI and created the draft Release with both artifacts, then failed deleting it: --cleanup-tag asked for refs/tags/rehearsal-<id>, and a draft creates no tag, so the API answered 422 after the draft itself was already gone. Exactly the kind of failure the rehearsal exists to find before a tag. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Signed-off-by: Oktie Hassanzadeh <oktie@oktie.com>
Run 34702494907 found the draft-cleanup bug, run 34702714569 passed all three jobs, and run 34702716866 -- GH_REPO removed on a throwaway branch, as on 1.5.0 -- failed the Release job with 1.5.0's own error. No draft, rehearsal tag or throwaway branch remains. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Signed-off-by: Oktie Hassanzadeh <oktie@oktie.com>
CHANGELOG: rehearsable releases and Archer's judge scores under Added; the upload script's explicit file list and the stale Gemini summary rows under Fixed. Plan item 3 records steps 1-4: the committed archer files, the new v1.6.0 Hub tag, the dropped Gemini rows, and the verified upload. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Signed-off-by: Oktie Hassanzadeh <oktie@oktie.com>
The 1.6.0 branch is on the deployment: rebuilt image, re-provisioned onto results v1.6.0, recreated with up -d. ulimit -n is 65536, every index is current, archer's judge comparison has evidence for 11 of 11 pipelines, and the documented health checks pass on the public domain. Item 3 is done; item 2 keeps its restart cron until the log shows the descriptor count flat across several twelve-hour periods, as its done-when asks. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Signed-off-by: Oktie Hassanzadeh <oktie@oktie.com>
Two packaged judge configs replace the four Llama-based ones, both wxai:openai/gpt-oss-120b at max_new_tokens 8192: llm_judge_default_config (with ground truth) and llm_judge_no_gt. The retired names resolve to their replacement in the loader and in the dashboard, unless a user config has the name. They were chosen against data/judge_calibration/: 202 labelled execution mismatches from BIRD, Spider, Spider Realistic and Archer, split into tune, test and a holdout drawn and labelled before any judge saw it. On the holdout the new default accepts 2 of 28 wrong predictions (Llama: 7), and the no-ground-truth judge 7 (Llama: 15). scripts/analysis/judge_calibration.py builds, runs and scores the set. Two defects found on the way: - An agentic prediction's trace was cut to 500 characters per message, task included, so the judge never saw its schema or hints. The task is now kept whole (evaluation/judge_inputs.py, shared with the calibration script). - The watsonx client salvaged a reasoning model's thinking as its answer when it ran out of tokens, which the judge scored N/A. Text generation now raises; SQL generation still falls back. Published scores are unchanged and still come from the Llama judge. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Signed-off-by: Oktie Hassanzadeh <oktie@oktie.com>
The batch judge reused any stored llm_score, whichever config had given it. Evaluating Llama-judged results with the gpt-oss-120b default would have kept every Llama score and recorded the new config in the summary. - Each verdict the judge gives stores llm_judge_config_digest, the same digest the dashboard's verdict cache keys on (judge_config_digest). A stored verdict is reused only under that digest; verdicts from before 1.6.0 carry none and are judged again. - llm_judge_reuse="any" keeps stored verdicts from any config, each with its own digest, and the run warns that the summary records the current config. rerun_metrics.py --preserve-llm-judge uses it, so it still calls no judge for a stored verdict. - Scores decided without the judge (subset match, no result) carry no digest and are worked out afresh; a stored "did not use LLM" note is never reused as a verdict. - The judge is asked once per prediction, about the ground truth that decided the result, instead of once per ground truth with every answer but the last discarded. - The digest is declared in metric_definitions and kept out of summary aggregation and the errors report's metric table. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Signed-off-by: Oktie Hassanzadeh <oktie@oktie.com>
All 710 Archer predictions the judge is asked about were judged again with llm_judge_default_config (wxai:openai/gpt-oss-120b), with no judge errors. The eval file's diff is exactly the 710 explanations, the 710 llm_judge_config_digest lines added, and 166 changed scores; every other field is unchanged against a backup taken before the run, and every column of the summary report other than LLM Score is identical. The Llama judge accepted 217 of the 710; gpt-oss-120b accepts 85. Every pipeline's LLM score falls, e.g. gpt-oss-120b zero-shot 0.558 -> 0.423. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Signed-off-by: Oktie Hassanzadeh <oktie@oktie.com>
Every LLM-judge call built a new ModelInference, which requests the project's details and an IAM token; watsonx rate-limits both. Re-judging the published results had most calls refused with "Exceeded limit of calls to endpoint" or a /token rate limit before any inference, and repeating the run could not get past it. WXAIClientChatAPI now takes its handle from a bounded cache keyed on a digest of the model, parameters and all credentials, built under a lock so concurrent evaluation threads do not all miss at once. A failed build is not cached, and clients with different keys never share a handle. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Signed-off-by: Oktie Hassanzadeh <oktie@oktie.com>
…udge Spider Realistic, Spider Dev, BIRD SQLite, BIRD PostgreSQL and Beaver were judged again with llm_judge_default_config (wxai:openai/gpt-oss-120b): 1,120, 2,051, 2,120, 2,190 and 941 judge calls, none failing. Archer was committed earlier (710). Summaries, CSVs, summary reports, charts and errors reports are regenerated from the new eval files; Beaver's summary report stays overall-only and its details are not in the repository. Every eval file was compared with a backup taken before the run. The judge accepts far fewer predictions than the Llama judge did (on BIRD PostgreSQL, 753 of its Yes verdicts become No), so every LLM score falls. Evaluating again also recomputed the other metrics under current code and dependencies, for the five benchmarks last evaluated in August: syntactic-equivalence scores move (Spider Dev: 1,915 of 10,340 predictions change sqlglot_equivalence), and 30 predictions that did not match their reference now subset-match (BIRD PostgreSQL 24, BIRD SQLite 5, Beaver 1), moving a pipeline's subset execution accuracy by at most 0.008. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Signed-off-by: Oktie Hassanzadeh <oktie@oktie.com>
The changelog replaces "published scores are unchanged" with what the re-judge did: 9,132 judge calls across the six benchmarks, none failing, every LLM score lower, and the other metrics recomputed under current code for the five benchmarks last evaluated in August (syntactic equivalence moved; 30 mismatches became subset matches, at most 0.008 on a pipeline). Plan item 4 records the re-judge, the re-scoring change, the v1.6.0 results tag moved to dataset commit f52b8d2 and checked file by file, the watsonx rate limit it surfaced, and answers "replace or add?" with replace. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Signed-off-by: Oktie Hassanzadeh <oktie@oktie.com>
…2xmp) Dependabot opened two alerts against anyio 4.12.1 in uv.lock: a critical one, IDNA 2003 host-name encoding in TLSStream that can let a certificate for another host be accepted, and a medium one, process-pool workers blocking on undrained stderr. Both are fixed in 4.14.2. anyio arrives through httpx, openai, google-genai and starlette, whose own floors allow the vulnerable versions, so it joins the security floors in pyproject.toml; the lockfile moves to 4.15.1 (typing-extensions 4.16.0 comes with it). The hermetic suite passes. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Signed-off-by: Oktie Hassanzadeh <oktie@oktie.com>
- Item 4: no new baselines; the existing gpt-oss-120b zero-shot and agentic baseline0-5 predictions are the gpt-oss-120b baselines. - Item 5: Beaver's details stay where they are for now -- served to signed-in users by the deployment from a copy placed on the host, with nowhere to download them from; distribution is to be worked out later. The GitHub Support purge request is dropped. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Signed-off-by: Oktie Hassanzadeh <oktie@oktie.com>
- pyproject.toml and uv.lock at 1.6.0; the changelog's Unreleased section becomes 1.6.0, dated 2026-09-19, with its link reference. The release notes extract cleanly for v1.6.0. - Plan item 2: the descriptor count stayed between 15 and 27 across thirteen twelve-hour periods since the fix, so the leak is fixed; the twice-daily restart cron stays on the host by decision. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Signed-off-by: Oktie Hassanzadeh <oktie@oktie.com>
The plan is deleted, as the attic's rule says: a plan stops being true the moment it is carried out. What survived it moves to where it will be met. The log entry records what the release turned on rather than what it shipped: that a split whose errors you read is a development set from then on, which is why the holdout was drawn and why the prompt written from the test split lost on it; that the judge's inputs were wrong before its prompt was, so every agentic verdict was reached without the schema and a model that ran out of tokens looked like one that said no; that a score has to carry the judge that gave it; that re-judging re-evaluates everything, so August's artifacts moved under September's dependencies; and that a rate limit can belong to the client rather than the workload. The attic's README no longer names an open plan. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Signed-off-by: Oktie Hassanzadeh <oktie@oktie.com>
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Judge inputs and config previews can misrepresent content, MySQL initialization mishandles valid secrets, and Beaver documentation remains contradictory.
Get a fresh assessment by requesting another Copilot review.
Review effort: Balanced
Findings: 1
Open (4)
What changed in this PR
Prepares release 1.6.0 with safer benchmark publishing, improved LLM judging, connection lifecycle fixes, deployment hardening, and dashboard enhancements.
Changes:
- Gates Beaver details behind sign-in and prevents restricted-result publication.
- Updates judge models, provenance tracking, prompt construction, and watsonx client reuse.
- Fixes SQLite leaks and adds release/deployment safeguards and dashboard navigation improvements.
| File | Description |
|---|---|
tests/test_watsonx_client_reuse.py |
Tests watsonx handle caching. |
tests/test_upload_results_script.py |
Tests restricted upload filtering. |
tests/test_security_hardening.py |
Updates packaged-config checks. |
tests/test_judge_reasoning_models.py |
Tests reasoning replies and judge inputs. |
tests/test_judge_endpoint.py |
Uses the new no-GT config. |
tests/test_judge_config_storage.py |
Updates packaged config expectations. |
tests/test_judge_config_retired.py |
Tests retired-config compatibility. |
tests/test_index_concurrency.py |
Tests connection reaping. |
tests/test_evaluate_prediction.py |
Tests digest-aware verdict reuse. |
tests/test_connection_leaks.py |
Tests SQLite connection closure. |
src/text2sql_eval_toolkit/ui/routers_judge.py |
Resolves retired judge names. |
src/text2sql_eval_toolkit/ui/routers_benchmarks.py |
Exposes and preserves access flags. |
src/text2sql_eval_toolkit/ui/models.py |
Adds benchmark lock metadata. |
src/text2sql_eval_toolkit/ui/middleware.py |
Enforces restricted access and noindex. |
src/text2sql_eval_toolkit/ui/judge_budget.py |
Closes ledger connections. |
src/text2sql_eval_toolkit/inference/inference_tools.py |
Reuses watsonx handles and rejects empty text. |
src/text2sql_eval_toolkit/indexing/store.py |
Reaps retired-thread connections. |
src/text2sql_eval_toolkit/indexing/builder.py |
Closes staleness-check connections. |
src/text2sql_eval_toolkit/evaluation/metric_definitions.py |
Documents judge-config digests. |
src/text2sql_eval_toolkit/evaluation/llm_judge_config/llm_judge_no_gt.yaml |
Adds the new no-GT judge. |
src/text2sql_eval_toolkit/evaluation/llm_judge_config/llm_judge_no_gt_v2.yaml |
Removes the retired config. |
src/text2sql_eval_toolkit/evaluation/llm_judge_config/llm_judge_default_config.yaml |
Replaces the default judge. |
src/text2sql_eval_toolkit/evaluation/llm_as_judge.py |
Adds digests and retired-name mapping. |
src/text2sql_eval_toolkit/evaluation/judge_inputs.py |
Centralizes judge input rendering. |
src/text2sql_eval_toolkit/data/test-benchmarks.json |
Removes the Beaver test subset. |
src/text2sql_eval_toolkit/data/benchmarks.json |
Marks Beaver as restricted. |
src/text2sql_eval_toolkit/analysis/error_analysis.py |
Hides digest metadata from reports. |
scripts/evaluation/rerun_metrics.py |
Adds configurable verdict reuse. |
scripts/analysis/llm_judge_comparison.py |
Uses current judge configs. |
pyproject.toml |
Bumps version and anyio floor. |
docs/guide/llm-judge.md |
Documents new judges and provenance. |
docs/dashboard/deployment.md |
Documents access and leak behavior. |
docs/dashboard/capability-tiers.md |
Documents benchmark restrictions. |
docs/attic/README.md |
Retires the 1.6.0 plan. |
deploy/sql/mysql-init/10-readonly-role.sh |
Creates configurable read-only grants. |
deploy/load-bird-postgres.sh |
Clarifies dump configuration. |
deploy/env.deploy.example |
Adds restriction and database settings. |
deploy/docker-compose.yml |
Raises file limits and passes settings. |
data/test-benchmarks.json |
Removes the Beaver test subset. |
data/results/spider_realistic-predictions_eval_summary.md |
Refreshes published metrics. |
data/results/spider_dev-predictions_eval_summary.md |
Refreshes published metrics. |
data/results/bird_mini_dev_postgres-predictions_eval_summary.md |
Refreshes published metrics. |
data/judge_calibration/README.md |
Documents judge calibration. |
data/judge_calibration/configs/llama_no_gt_v1.yaml |
Preserves a calibration baseline. |
data/judge_calibration/configs/llama_default.yaml |
Updates the calibration baseline. |
data/benchmarks/test_benchmarks/results/README.md |
Removes Beaver test results. |
data/benchmarks/README.md |
Documents gated Beaver access. |
data/benchmarks/dbs/README.md |
Updates Beaver database setup. |
data/benchmarks.json |
Marks Beaver as restricted. |
dashboard/src/views/SignInRequired.tsx |
Adds a sign-in-required view. |
dashboard/src/views/SignInRequired.test.tsx |
Tests sign-in navigation. |
dashboard/src/views/RunEvaluationView.tsx |
Improves playground layout. |
dashboard/src/views/PipelineDetailView.tsx |
Links records to the playground. |
dashboard/src/views/OpenInPlaygroundButton.tsx |
Adds a reusable playground link. |
dashboard/src/views/JudgePlayground.tsx |
Displays selected judge configuration. |
dashboard/src/views/ErrorAnalysis.tsx |
Links errors to the playground. |
dashboard/src/views/BenchmarkTiles.tsx |
Labels restricted benchmarks. |
dashboard/src/views/BenchmarkDetail.tsx |
Falls back to public overall scores. |
dashboard/src/views/BenchmarkDetail.test.tsx |
Tests summary fallback. |
dashboard/src/types/benchmark.ts |
Adds restriction fields. |
dashboard/src/services/benchmarks.ts |
Recognizes sign-in-required responses. |
dashboard/src/pages/App.tsx |
Routes locked views to sign-in. |
dashboard/src/lib/api.ts |
Adds a typed sign-in error. |
dashboard/src/lib/api.test.ts |
Tests sign-in error handling. |
dashboard/e2e/shareable-urls.spec.ts |
Tests playground links. |
dashboard/dist/index.html |
Points to the rebuilt frontend. |
dashboard/dist/assets/TableRow-Td9PgCqq.js |
Adds rebuilt table chunk. |
dashboard/dist/assets/TableRow-B81J3hyC.js |
Removes obsolete table chunk. |
dashboard/dist/assets/TableHeader-CH8CR35j.js |
Rebuilds table-header chunk. |
dashboard/dist/assets/TableContainer-O9G7CXka.js |
Adds rebuilt container chunk. |
dashboard/dist/assets/TableContainer-B7wsClud.js |
Removes obsolete container chunk. |
dashboard/dist/assets/swimlanesDiagram-VR7AAH4N-JB22Rs9l.js |
Adds rebuilt diagram chunk. |
dashboard/dist/assets/swimlanesDiagram-VR7AAH4N-D3ZXWkB1.js |
Removes obsolete diagram chunk. |
dashboard/dist/assets/stateDiagram-v2-MP3YSRHH-FQIrQuHa.js |
Removes obsolete state chunk. |
dashboard/dist/assets/stateDiagram-v2-MP3YSRHH-DwmnbYBy.js |
Adds rebuilt state chunk. |
dashboard/dist/assets/sizeCapture-INFHLROL-D3y8d2OJ.js |
Refreshes generated imports. |
dashboard/dist/assets/ResultTableView-D7zeyOYh.js |
Removes obsolete result chunk. |
dashboard/dist/assets/railroadDiagram-O6MQD6OU-D3WtqEwj.js |
Refreshes generated imports. |
dashboard/dist/assets/pieDiagram-E7YTZNPT-DF2dAtj3.js |
Refreshes generated imports. |
dashboard/dist/assets/pegDiagram-XKGWAZYB-C43s2JK-.js |
Refreshes generated imports. |
dashboard/dist/assets/OpenInPlaygroundButton-CCS_tmfs.js |
Adds rebuilt playground chunk. |
dashboard/dist/assets/metricInsightsSelect-Ca-YJ6Xf.js |
Refreshes generated imports. |
dashboard/dist/assets/infoDiagram-27XIBGKW-Bi6ru0tL.js |
Refreshes generated imports. |
dashboard/dist/assets/flowDiagram-HODETNUW-BZ_5n_mm.js |
Refreshes generated imports. |
dashboard/dist/assets/ebnfDiagram-PWID7BFC-u2JqIZgy.js |
Refreshes generated imports. |
dashboard/dist/assets/classDiagram-ZZMXUADV-Syw15DOs.js |
Removes obsolete class chunk. |
dashboard/dist/assets/classDiagram-ZZMXUADV-BUlv72DX.js |
Adds rebuilt class chunk. |
dashboard/dist/assets/classDiagram-v2-VYDZK3BY-Syw15DOs.js |
Removes obsolete class-v2 chunk. |
dashboard/dist/assets/classDiagram-v2-VYDZK3BY-BUlv72DX.js |
Adds rebuilt class-v2 chunk. |
dashboard/dist/assets/chunk-XXDRQBXY-B_g5D6EZ.js |
Refreshes generated chunk. |
dashboard/dist/assets/chunk-SVP7TREG-DLSceb9V.js |
Refreshes generated chunk. |
dashboard/dist/assets/chunk-POPQ4Y6H-DPe82o4u.js |
Refreshes generated chunk. |
dashboard/dist/assets/chunk-JWPE2WC7-B1bCvL2R.js |
Refreshes generated chunk. |
dashboard/dist/assets/chunk-F27PBJKO-Slp3VtyJ.js |
Refreshes generated chunk. |
dashboard/dist/assets/chunk-5VM5RSS4-CeSGRynF.js |
Refreshes generated chunk. |
dashboard/dist/assets/chunk-2Q5K7J3B-yhnntLfj.js |
Refreshes generated chunk. |
dashboard/dist/assets/channel-YHNFbJNY.js |
Removes obsolete channel chunk. |
dashboard/dist/assets/channel-luFgzhf3.js |
Adds rebuilt channel chunk. |
dashboard/dist/assets/bucket-0-Bdefr4Uj.js |
Refreshes generated imports. |
dashboard/dist/assets/abnfDiagram-VCTEODGH-D99xhN_E.js |
Refreshes generated imports. |
CONTRIBUTING.md |
Documents release rehearsals. |
.gitignore |
Excludes gated Beaver artifacts. |
.github/workflows/ci.yml |
Checks the app file limit. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
- The MySQL read-only role script interpolated the password and each database name straight into SQL. A password containing an apostrophe produced a malformed IDENTIFIED BY, and a backtick in a database name would have ended the quoted identifier. The password is now escaped for a string literal, and a database name that is not a plain identifier is refused rather than emitted. - The judge playground kept the previous config's YAML and model on screen while the new one loaded, under the new config's name. Both are cleared before the request. - Every later message in an agentic trace was suffixed with "...", whether or not it had reached the 500-character limit, telling the judge that what it had been given was incomplete. The marker is added only where something was removed. data/judge_calibration/README.md notes that the stored inputs predate this. - data/benchmarks/dbs/README.md said Beaver's databases are gated and then described public dumps. The coverage figure stays; the contradiction is gone. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Signed-off-by: Oktie Hassanzadeh <oktie@oktie.com>
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Public uploads can expose environment-restricted benchmarks, and the sign-in page can incorrectly hide authentication.
Get a fresh assessment by requesting another Copilot review.
Review effort: Balanced
Findings: 1
Open (2)
- The upload script decided what to publish from the registry flag alone, so a benchmark marked restricted only through TEXT2SQL_SIGN_IN_BENCHMARKS -- the other way a deployment marks one, and the one that needs no file edit -- would have had its predictions, evaluation, errors report and category charts uploaded from that same host. The environment's ids are now restricted too, and ids are compared without case, since on a case-insensitive filesystem `BEAVER-…` opens Beaver's files. - SignInRequired treated "not asked yet" as "this server has no sign-in": the claim appeared on first render and stayed there if the deployment request failed, leaving no way to authenticate. It now distinguishes asking, answered and failed, claims sign-in is unavailable only when the server said so, and offers sign-in anyway when the request failed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Signed-off-by: Oktie Hassanzadeh <oktie@oktie.com>
- A failed judge call erased the verdict already stored. Re-judging under a new config declines to reuse it, so a call the provider refused left the prediction with no llm_score -- and compute_summary averages a missing score as 0, which is how a run interrupted by a rate limit would have published scores far below the ones the judge gave. The stored verdict is kept, under the digest of the config that gave it, alongside the error, and a run with any judge errors says so. - Keeping an agent's task whole put no bound on the judge's prompt: one Beaver trace runs to 160,000 characters, which a judge config on a smaller-context model cannot take -- and each such failure then scored 0. The task is kept to 40,000 characters, longer than 99% of the traces in data/judge_calibration/, and cut with a marker beyond that. - The playground listed the 64-character judge digest as a metric, because its table is built from the evaluation's own keys. It is shown under the LLM judge section, shortened, with the whole value in the title. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Signed-off-by: Oktie Hassanzadeh <oktie@oktie.com>
|
Did another round of review. Three findings, all fixed in 8050cf1. A failed judge call erased the verdict already stored. Re-judging under a new config declines to reuse the stored verdict, so a call the provider refused left the prediction with no Keeping an agent's task whole left the judge's prompt unbounded. That fix was right — agentic predictions were being judged without their schema — but one Beaver trace runs to 160,000 characters, about 40k tokens. gpt-oss-120b takes it; a judge on a smaller-context model (the retired Llama config, still loadable by its old name, or a user's own) would fail on exactly the records with the richest traces, and each failure then scored 0. The task is now kept to 40,000 characters, longer than 99% of the traces in The playground listed the judge digest as a metric. Its metric table is built from the evaluation's own keys, so every judged record showed a 64-character hex string next to Checks: 1,197 Python tests, 212 dashboard tests, lint clean, bundle rebuilt. Tests were added for the first two. Not changed: the published results keep the verdicts they have. The prompt cap changes what a rebuild of the calibration inputs would produce, which |
Both came out of the last commit rather than the release: - The judge prompt was still unbounded. The task cap applied per message, so a first step carrying three long messages rendered 120,000 characters against a 40,000 limit. The task's messages now share one budget, and each keeps at least TRACE_MESSAGE_CHARS so the question survives a schema that spent the budget before it. - A summary could contradict itself. A verdict kept through a failed call is in the llm_score average, but num_correct_llm required the absence of an error, so a pipeline could read "1.00 average, 1 of 2 correct". The count now follows the score; a failure with nothing stored has no score and is counted by neither. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Signed-off-by: Oktie Hassanzadeh <oktie@oktie.com>
|
Second review pass, before merge. Two findings, both from the previous round of fixes rather than the release itself; fixed in e2fb3b3. The judge prompt was still unbounded. The 40,000-character task cap applied per message, not to the task, so a first step carrying three long messages rendered 120,000 characters — verified by rendering one. The task's messages now share a single budget, and each keeps at least 500 characters so the question still appears after a schema that spent the budget. A summary could contradict itself. A verdict kept through a failed call is in the Also re-checked this pass, nothing found: the calibration script's report and resume paths, the CI compose assertions, the registry's Checks: 1,199 Python tests, 212 dashboard tests, lint clean, bundle rebuilt. |
A stored `llm_judge_error` says the last judge call failed. It does not say there is no verdict: a run that fails keeps the one an earlier run stored, beside the error. `_is_judge_verdict` answered both questions at once, so: - The second refusal in a row dropped the verdict the first one had kept, and the prediction ended with no `llm_score` at all -- which a summary counts as 0. That is the collapse the handler exists to prevent, and it came back on the retry the run itself tells the operator to do. - The "verdicts were kept from a different judge config" warning counted only verdicts with no error beside them, which is never how one of these is stored. Under the default `llm_judge_reuse="matching"` it never fired, so a run the provider refused published old-config scores labelled with the new one without a word. Split the two questions with `past_failure`, defaulting to the old answer so the reuse path still retries a failed prediction rather than settling for what is stored. Separately, `_check_restricted_summaries` built its path from the restricted id's spelling while `_is_publishable` decides what to upload by the file's. With `TEXT2SQL_SIGN_IN_BENCHMARKS=Beaver` on a case-sensitive filesystem the check found nothing, passed, and let the real `beaver-...md` through with its `## Category:` sections -- the one publishable artifact that can still carry gated derived data. Casefold the ids at the source, as `ui/benchmark_access.py` does, and match the reports on disk. The nested-layout manifest check casefolds too, or a directory the upload skips would be named in the manifest and `results fetch` would fail on it. Six tests; five fail without these changes. The sixth pins that the first fix did not over-correct into reusing a stale verdict instead of asking again. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Signed-off-by: Oktie Hassanzadeh <oktie@oktie.com>
Review round 3 — three findings, fixed in 67642bfA full pass over the ~4,400 lines of actual code in the diff (Python source, scripts, dashboard sources, Findings1. A second consecutive judge failure erased the verdict the first one kept —
So: re-judge under a new config, watsonx refuses N calls. Run 1 stores That is precisely the published-scores collapse the handler was written to prevent, reappearing on the retry the run recommends. Reproduced directly: run 1 → 2. The gated-category guard missed when the restricted id's case differed —
On macOS the path lookup is case-insensitive, so the guard fires by accident and the bug is invisible locally. 3. The "kept from a different judge config" warning never fired under the default mode — Same root cause as 1: the The fixesFindings 1 and 3 share a cause, so they share a fix: Finding 2 is fixed at the source: VerificationSix new tests. Five fail against the unfixed source (verified by stashing the two source files and re-running). The sixth, Because macOS resolves paths case-insensitively, the case test asserts on the filename in the refusal message rather than on the refusal itself, which distinguishes old from new on any platform. Full suite 1205 passed (was 1199); Checked and found sound
🤖 Generated with Claude Code |
`data/judge/usage.sqlite` holds the dashboard's LLM-judge spend counters and verdict cache, both written at runtime. `JudgeStore` creates the file, its parent directory and its schema on first use, so the tracked copy seeded nothing -- and it had gone stale, missing the `user_caps` table a fresh database now creates. Tracking it meant every local use of the judge left a pending change carrying a user hash, a model, a cost and a judge's explanation, one `git commit -a` away from being published. `data/judge/` is now ignored. A deployment is unaffected: its database lives in the `data` volume that docker-compose mounts at /data, not in the repository. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Signed-off-by: Oktie Hassanzadeh <oktie@oktie.com>



Release 1.6.0. The full notes are the
## [1.6.0]section ofCHANGELOG.md; the reasoning behind each item is indocs/attic/project-log.md, newest entry first.Highlights
Security
anyiofloored at 4.14.2 (GHSA-82r6-8w77-94w6, critical; GHSA-5p39-cfhj-2xmp). Merging closes the two open Dependabot alerts onmain.LLM judge
wxai:openai/gpt-oss-120b:llm_judge_default_config(with ground truth) andllm_judge_no_gt. They replace the four Llama-based configs, whose names still load their replacement. Chosen against a labelled set indata/judge_calibration/: on its 52-item holdout the default accepts 2 of 28 wrong predictions where the Llama config accepted 7.llm_judge_config_digest);llm_judge_reuse="any"keeps the old behaviour on request.v1.6.0results snapshot. Every LLM score falls. Re-evaluating also recomputed other metrics under current code for the five benchmarks last evaluated in August — see the changelog.Beaver
Deployment and release
workflow_dispatchruns the GitHub Release job as a draft).Checks
mkdocs build --strictpasses.scripts/ci/extract_changelog.py v1.6.0extracts the notes.v1.6.0results snapshot.After merging
main:git tag -a v1.6.0 -m "Release v1.6.0"and push the tag.release.ymlbuilds, checks the tag against the packaged version, waits on thepypienvironment's reviewer, publishes, and creates the GitHub Release.docs/attic/project-log.md, as the attic's rule asks.🤖 Generated with Claude Code