Skip to content

Release 1.6.0 - #17

Merged
oktie merged 32 commits into
mainfrom
1.6.0
Sep 20, 2026
Merged

oktie merged 32 commits into
mainfrom
1.6.0

Conversation

@oktie

@oktie oktie commented Sep 19, 2026 •

Copy link
Copy Markdown
Member

Release 1.6.0. The full notes are the ## [1.6.0] section of CHANGELOG.md; the reasoning behind each item is in docs/attic/project-log.md, newest entry first.

Highlights

Security

LLM judge

  • The packaged judges use wxai:openai/gpt-oss-120b: llm_judge_default_config (with ground truth) and llm_judge_no_gt. They replace the four Llama-based configs, whose names still load their replacement. Chosen against a labelled set in data/judge_calibration/: on its 52-item holdout the default accepts 2 of 28 wrong predictions where the Llama config accepted 7.
  • A stored verdict is reused only under the config that gave it (llm_judge_config_digest); llm_judge_reuse="any" keeps the old behaviour on request.
  • The published results are re-judged with the new default: 9,132 judge calls across the six benchmarks, in the v1.6.0 results snapshot. Every LLM score falls. Re-evaluating also recomputed other metrics under current code for the five benchmarks last evaluated in August — see the changelog.
  • Fixes: agentic predictions were judged without their schema or hints; a reasoning model's out-of-tokens reply was scored as a verdict; a batch run was refused by watsonx rate limits because every call built a new client.

Beaver

  • Its questions, SQL, schema and per-record results are gated, and were purged from this repository's history and the public results dataset on 2026-09-12. The dashboard shows Beaver's tile and overall scores to everyone and its details only to signed-in users.

Deployment and release

  • The dashboard no longer leaks SQLite connections until it runs out of files; the app container's open-file limit is raised.
  • A release can be rehearsed end to end (workflow_dispatch runs the GitHub Release job as a draft).

Checks

  • Hermetic suite: 1,192 passed. mkdocs build --strict passes. scripts/ci/extract_changelog.py v1.6.0 extracts the notes.
  • The public deployment runs this branch against the v1.6.0 results snapshot.

After merging

  1. Tag main: git tag -a v1.6.0 -m "Release v1.6.0" and push the tag. release.yml builds, checks the tag against the packaged version, waits on the pypi environment's reviewer, publishes, and creates the GitHub Release.
  2. Nothing else: the 1.6.0 plan is already retired into docs/attic/project-log.md, as the attic's rule asks.

🤖 Generated with Claude Code

oktie and others added 26 commits September 1, 2026 10:13
Seeded with the one thing 1.5.0's own release turned up: the release workflow's
last job could not run, and its dry-run mode could not have caught that.

`github-release` is gated on `refs/tags/v*`, so a `workflow_dispatch` from a
branch skips it -- the rehearsal path structurally excludes the only job that
had never run. That is the item: a dispatch that exercises `gh release create`
against real artifacts and the real token, and a deliberately broken step that
fails the rehearsal rather than the next tag.

Also records four things noticed during 1.5.0 and deliberately left alone, so
they are not lost: the tracked judge usage database that churns on every local
run, Playwright adopting a stray server as the system under test, Carbon's
always-mounted modal, and the docs directory being discovered rather than
configured. None investigated beyond the observation.

The attic README points at the open plan again.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Signed-off-by: Oktie Hassanzadeh <oktie@oktie.com>
…boxed

The question and db_id were set at 0.875rem in secondary text, smaller and
dimmer than the body copy around them, with the db_id run onto the end of the
same line. They are the subject of the whole view and were the hardest thing on
it to locate.

The question now has its own block: a QUESTION label, the text at 1.25rem in
primary colour, and the database as a labelled tag rather than an inline
fragment. Tinted with layer-02 and not layer-01 -- inside this content layer the
latter resolves to the same white as the page, so the first attempt had a
background in the code and none on screen.

Ground truth and predicted SQL each get a box. A bare Carbon text area carries
only an underline, so two side by side read as one undivided region instead of
as the two queries being compared. The border matches the panels below them.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Signed-off-by: Oktie Hassanzadeh <oktie@oktie.com>
The detail panels show a record read-only -- the question, both queries, the
results, the metrics. The playground is where that same record can be edited and
re-run. Getting between them meant reading the record id off the address bar and
assembling a /run/... URL by hand, which is a thing nobody does twice.

One button, in both places that open a record detail: Error Analysis and a
pipeline's detail view. It carries the benchmark, the record and the pipeline
being read, so the playground opens on the prediction you were just looking at
rather than picking its own default.

An anchor, not a button with a click handler. The playground address is worth
copying and worth opening in a new tab, and only a plain left click is
intercepted for single-page navigation -- every other gesture goes back to the
browser. Without an onNavigate the control stays a working link and the browser
navigates normally, so a caller that forgets to wire it degrades rather than
breaks.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Signed-off-by: Oktie Hassanzadeh <oktie@oktie.com>
Picking a judge config by name said nothing about what it would ask. The model
it runs and the prompt it sends are the whole of what a judge is, and finding
either meant leaving for the config editor and coming back.

The box names the config, tags its model, and puts the YAML behind a Show
prompt toggle. Collapsed by default: the prompt template is the bulk of every
config and runs to forty-odd lines, which would push the run controls and the
verdict off the screen for a reader who only wanted to know which model was
judging.

A `pre`, not the CodeMirror editor the config view uses. That editor's chunk is
416 KB against this view's 44 KB -- a great deal of syntax highlighting for a
read-only prompt template. This adds 4 KB.

Also fixes a sentence I broke adding Judge again: JSX drops a newline between an
element and the text after it, so "Judge again ignores the cache" rendered as
"Judge againignores the cache". And the model tag carries its full value as a
title, since Carbon truncates a long provider:model id.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Signed-off-by: Oktie Hassanzadeh <oktie@oktie.com>
…never closed

is_stale() and every JudgeStore method opened SQLite with
`with sqlite3.connect(...)`, which commits but does not close. A connection
refers to itself through its statement cache, so each stayed open until the
cyclic collector ran: six descriptors per landing-page load, one per /api/me
from a signed-in caller. On the public deployment that reached Docker's default
limit of 1024 and the app could no longer open the results it serves.

Both now close what they open. An index also reaps the connections of worker
threads that have exited, a smaller leak found on the way. Compose raises the
app's nofile limit to 65536, with a CI check, and the deployment guide gains a
troubleshooting row. Recorded as item 2 of the 1.6.0 plan and in the project log.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Signed-off-by: Oktie Hassanzadeh <oktie@oktie.com>
Items 2 to 5 and the items carried over from 1.5.0, as they stood before
work on item 5 began: the file-descriptor exhaustion on the deployment,
Archer's LLM-judge results, gpt-oss-120b baselines and judge configs, and
keeping Beaver behind a sign-in.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Signed-off-by: Oktie Hassanzadeh <oktie@oktie.com>
Plan item 5, in the dashboard. Every read route was public tier, and a
tier cannot tell an anonymous visitor from a signed-in read_only user, so
this is a check on identity, enforced in the tier middleware ahead of the
tier gate.

- A registry entry with "requires_sign_in": true (beaver, beaver_test_10)
  is hidden from anonymous callers on shared deployments. The flag is read
  from every registry copy, packaged ones included, because provision.sh
  never overwrites a data root's seeded benchmarks.json.
  TEXT2SQL_SIGN_IN_BENCHMARKS adds ids and is passed through compose.
- Every route whose declared parameters name such a benchmark answers 401
  with sign_in_required, the logo included; /api/benchmarks leaves it out.
  Ids are compared casefolded, and an id outside [A-Za-z0-9_-] is a 400 for
  a caller the wall applies to: /api/compare?left_id=charts/../beaver read
  Beaver's summary.
- X-Robots-Tag: noindex, nofollow on any address that mentions it, API and
  app shell alike.
- A Beaver link opened anonymously asks the reader to sign in and returns
  them to the same address; a signed-in listing tags it "Signed-in only".
  Editing a benchmark keeps its flag.

tests/test_benchmark_access.py parametrizes over the live route table and
fails on any route parameter not classified as naming a benchmark or not,
so a new route cannot serve one unnoticed.

The GitHub and Hugging Face copies of Beaver's data are untouched; how to
handle them is still an open question in the plan.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Signed-off-by: Oktie Hassanzadeh <oktie@oktie.com>
archer_en_dev was the only benchmark with no judge scores at all, which is
why its Metric Insights page showed no evidence under LLM judge comparison.
Run on 2026-09-01 with llm_judge_default_config
(wxai:meta-llama/llama-3-3-70b-instruct), the same model and prompt the
existing summary records.

1,093 of 1,144 predictions scored: 543 Yes, 550 No. The other 51 are
agentic predictions with no SQL (20 in baseline4, 31 in baseline5).

The eval file keeps HEAD's formatting and key order, so its diff is the
llm_score and llm_explanation lines and the trailing commas they need;
the summary, errors report and charts are regenerated from it.
archer_en_dev-predictions.json is unchanged.

Plan item 3, step 1. Not yet on the Hub or the deployment.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Signed-off-by: Oktie Hassanzadeh <oktie@oktie.com>
…ign-in

Beaver's questions, SQL and schema are distributed under gated access, and
only what its own leaderboard shows -- overall scores -- may be published.
The wall added the day before hid the whole benchmark from anonymous
callers; it now inverts to what may be shown.

- Anonymous callers get Beaver's tile, logo, overall summary, alias table
  and /api/compare. SUMMARY_ROUTES is an allowlist: every other route that
  names Beaver -- the breakdown by query category, error analysis, record
  detail, playground, insights, config, the judge -- answers 401
  sign_in_required, and a route added later is refused until it is added
  on purpose. Ids outside [A-Za-z0-9_-] are still refused first.
- /api/benchmarks lists Beaver for everyone with details_locked for
  anonymous callers. The summary page falls back to the overall table;
  every other Beaver view asks the reader to sign in, with a way back to
  the overall scores.
- The repository follows the history purge: beaver_test_10 unregistered,
  Beaver's summary report cut to overall scores, its per-category charts
  untracked, local copies of its files git-ignored. The MySQL read-only
  grant takes MYSQL_READONLY_DATABASES instead of naming Beaver's
  databases, and the Beaver database docs point at its gated distribution.
- Docs, changelog, plan item 5 and the project log record the decisions,
  the GitHub and Hugging Face purge, and what is still open: GitHub
  Support for PR refs and caches, forks and old clones, and a gated source
  for the deployment's Beaver details.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Signed-off-by: Oktie Hassanzadeh <oktie@oktie.com>
1.5.0's release published to PyPI and then failed on its last job, which
had never run: the dispatch path that looked like a rehearsal skipped it
with `if: startsWith(github.ref, 'refs/tags/v')`. Plan item 1.

- The Release job now runs on workflow_dispatch. It uses the same
  `gh release create` step a tag uses, with --draft under a throwaway
  name, and a following step deletes the draft even if a later check
  fails, and fails the job if the delete does.
- The build job extracts release notes on every run, for the version in
  pyproject.toml when there is no tag, and always uploads them.
- A rehearsal's TestPyPI upload skips a version already there, so it can
  be re-run; a real publish still refuses one.
- CONTRIBUTING.md says how to rehearse and what a rehearsal does not
  cover: the tag/version check, the pypi reviewer gate, PyPI itself.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Signed-off-by: Oktie Hassanzadeh <oktie@oktie.com>
…ark's details

The upload script sent everything under results/ except logs. Run from a
maintainer's checkout that would have published the derived query indices
in results/.index/ -- which carry every record's raw bytes -- backups, and
the local copies of Beaver's predictions, evaluation file and errors
report, which are distributed under gated access.

It now uploads an exact allow list. Hidden directories, bak/ and logs are
never included, and a benchmark flagged requires_sign_in in any registry
copy publishes only its overall summary (json, csv, md) and overall
chart. The manifest lists only what is uploaded, and an upload whose
restricted summary report still breaks results down by query category is
refused. tests/test_upload_results_script.py pins the file list and the
dry run's output.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Signed-off-by: Oktie Hassanzadeh <oktie@oktie.com>
The summaries for bird_mini_dev_postgres, bird_mini_dev_sqlite,
spider_dev and spider_realistic listed
gemini:gemini-3-flash-preview-greedy-zero-shot-chatapi, which none of
those benchmarks' predictions or eval files contain. Plan item 3, step 3.

The row is dropped from each summary JSON, with every other pipeline's
values unchanged, and the CSV and overall chart are regenerated from it
with the toolkit's own writers. The Markdown reports, rebuilt from the
eval records, came out identical, and no errors report mentioned the
pipeline. make_summary_report.py was not used: it rewrites the
hand-written data/results/README.md.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Signed-off-by: Oktie Hassanzadeh <oktie@oktie.com>
The first rehearsal (run 34702494907) built, published to TestPyPI and
created the draft Release with both artifacts, then failed deleting it:
--cleanup-tag asked for refs/tags/rehearsal-<id>, and a draft creates no
tag, so the API answered 422 after the draft itself was already gone.
Exactly the kind of failure the rehearsal exists to find before a tag.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Signed-off-by: Oktie Hassanzadeh <oktie@oktie.com>
Run 34702494907 found the draft-cleanup bug, run 34702714569 passed all
three jobs, and run 34702716866 -- GH_REPO removed on a throwaway branch,
as on 1.5.0 -- failed the Release job with 1.5.0's own error. No draft,
rehearsal tag or throwaway branch remains.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Signed-off-by: Oktie Hassanzadeh <oktie@oktie.com>
CHANGELOG: rehearsable releases and Archer's judge scores under Added;
the upload script's explicit file list and the stale Gemini summary rows
under Fixed. Plan item 3 records steps 1-4: the committed archer files,
the new v1.6.0 Hub tag, the dropped Gemini rows, and the verified upload.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Signed-off-by: Oktie Hassanzadeh <oktie@oktie.com>
The 1.6.0 branch is on the deployment: rebuilt image, re-provisioned onto
results v1.6.0, recreated with up -d. ulimit -n is 65536, every index is
current, archer's judge comparison has evidence for 11 of 11 pipelines,
and the documented health checks pass on the public domain. Item 3 is
done; item 2 keeps its restart cron until the log shows the descriptor
count flat across several twelve-hour periods, as its done-when asks.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Signed-off-by: Oktie Hassanzadeh <oktie@oktie.com>
Two packaged judge configs replace the four Llama-based ones, both
wxai:openai/gpt-oss-120b at max_new_tokens 8192: llm_judge_default_config
(with ground truth) and llm_judge_no_gt. The retired names resolve to their
replacement in the loader and in the dashboard, unless a user config has the
name.

They were chosen against data/judge_calibration/: 202 labelled execution
mismatches from BIRD, Spider, Spider Realistic and Archer, split into tune,
test and a holdout drawn and labelled before any judge saw it. On the
holdout the new default accepts 2 of 28 wrong predictions (Llama: 7), and
the no-ground-truth judge 7 (Llama: 15). scripts/analysis/judge_calibration.py
builds, runs and scores the set.

Two defects found on the way:
- An agentic prediction's trace was cut to 500 characters per message, task
  included, so the judge never saw its schema or hints. The task is now kept
  whole (evaluation/judge_inputs.py, shared with the calibration script).
- The watsonx client salvaged a reasoning model's thinking as its answer when
  it ran out of tokens, which the judge scored N/A. Text generation now
  raises; SQL generation still falls back.

Published scores are unchanged and still come from the Llama judge.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Signed-off-by: Oktie Hassanzadeh <oktie@oktie.com>
The batch judge reused any stored llm_score, whichever config had given
it. Evaluating Llama-judged results with the gpt-oss-120b default would
have kept every Llama score and recorded the new config in the summary.

- Each verdict the judge gives stores llm_judge_config_digest, the same
  digest the dashboard's verdict cache keys on (judge_config_digest). A
  stored verdict is reused only under that digest; verdicts from before
  1.6.0 carry none and are judged again.
- llm_judge_reuse="any" keeps stored verdicts from any config, each with
  its own digest, and the run warns that the summary records the current
  config. rerun_metrics.py --preserve-llm-judge uses it, so it still calls
  no judge for a stored verdict.
- Scores decided without the judge (subset match, no result) carry no
  digest and are worked out afresh; a stored "did not use LLM" note is
  never reused as a verdict.
- The judge is asked once per prediction, about the ground truth that
  decided the result, instead of once per ground truth with every answer
  but the last discarded.
- The digest is declared in metric_definitions and kept out of summary
  aggregation and the errors report's metric table.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Signed-off-by: Oktie Hassanzadeh <oktie@oktie.com>
All 710 Archer predictions the judge is asked about were judged again with
llm_judge_default_config (wxai:openai/gpt-oss-120b), with no judge errors.
The eval file's diff is exactly the 710 explanations, the 710
llm_judge_config_digest lines added, and 166 changed scores; every other
field is unchanged against a backup taken before the run, and every column
of the summary report other than LLM Score is identical.

The Llama judge accepted 217 of the 710; gpt-oss-120b accepts 85. Every
pipeline's LLM score falls, e.g. gpt-oss-120b zero-shot 0.558 -> 0.423.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Signed-off-by: Oktie Hassanzadeh <oktie@oktie.com>
Every LLM-judge call built a new ModelInference, which requests the
project's details and an IAM token; watsonx rate-limits both. Re-judging
the published results had most calls refused with "Exceeded limit of calls
to endpoint" or a /token rate limit before any inference, and repeating
the run could not get past it.

WXAIClientChatAPI now takes its handle from a bounded cache keyed on a
digest of the model, parameters and all credentials, built under a lock so
concurrent evaluation threads do not all miss at once. A failed build is
not cached, and clients with different keys never share a handle.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Signed-off-by: Oktie Hassanzadeh <oktie@oktie.com>
…udge

Spider Realistic, Spider Dev, BIRD SQLite, BIRD PostgreSQL and Beaver were
judged again with llm_judge_default_config (wxai:openai/gpt-oss-120b):
1,120, 2,051, 2,120, 2,190 and 941 judge calls, none failing. Archer was
committed earlier (710). Summaries, CSVs, summary reports, charts and
errors reports are regenerated from the new eval files; Beaver's summary
report stays overall-only and its details are not in the repository.

Every eval file was compared with a backup taken before the run. The
judge accepts far fewer predictions than the Llama judge did (on BIRD
PostgreSQL, 753 of its Yes verdicts become No), so every LLM score falls.

Evaluating again also recomputed the other metrics under current code and
dependencies, for the five benchmarks last evaluated in August:
syntactic-equivalence scores move (Spider Dev: 1,915 of 10,340 predictions
change sqlglot_equivalence), and 30 predictions that did not match their
reference now subset-match (BIRD PostgreSQL 24, BIRD SQLite 5, Beaver 1),
moving a pipeline's subset execution accuracy by at most 0.008.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Signed-off-by: Oktie Hassanzadeh <oktie@oktie.com>
The changelog replaces "published scores are unchanged" with what the
re-judge did: 9,132 judge calls across the six benchmarks, none failing,
every LLM score lower, and the other metrics recomputed under current code
for the five benchmarks last evaluated in August (syntactic equivalence
moved; 30 mismatches became subset matches, at most 0.008 on a pipeline).

Plan item 4 records the re-judge, the re-scoring change, the v1.6.0 results
tag moved to dataset commit f52b8d2 and checked file by file, the watsonx
rate limit it surfaced, and answers "replace or add?" with replace.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Signed-off-by: Oktie Hassanzadeh <oktie@oktie.com>
…2xmp)

Dependabot opened two alerts against anyio 4.12.1 in uv.lock: a critical
one, IDNA 2003 host-name encoding in TLSStream that can let a certificate
for another host be accepted, and a medium one, process-pool workers
blocking on undrained stderr. Both are fixed in 4.14.2.

anyio arrives through httpx, openai, google-genai and starlette, whose own
floors allow the vulnerable versions, so it joins the security floors in
pyproject.toml; the lockfile moves to 4.15.1 (typing-extensions 4.16.0
comes with it). The hermetic suite passes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Signed-off-by: Oktie Hassanzadeh <oktie@oktie.com>
- Item 4: no new baselines; the existing gpt-oss-120b zero-shot and
  agentic baseline0-5 predictions are the gpt-oss-120b baselines.
- Item 5: Beaver's details stay where they are for now -- served to
  signed-in users by the deployment from a copy placed on the host, with
  nowhere to download them from; distribution is to be worked out later.
  The GitHub Support purge request is dropped.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Signed-off-by: Oktie Hassanzadeh <oktie@oktie.com>
- pyproject.toml and uv.lock at 1.6.0; the changelog's Unreleased section
  becomes 1.6.0, dated 2026-09-19, with its link reference. The release
  notes extract cleanly for v1.6.0.
- Plan item 2: the descriptor count stayed between 15 and 27 across
  thirteen twelve-hour periods since the fix, so the leak is fixed; the
  twice-daily restart cron stays on the host by decision.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Signed-off-by: Oktie Hassanzadeh <oktie@oktie.com>
The plan is deleted, as the attic's rule says: a plan stops being true the
moment it is carried out. What survived it moves to where it will be met.

The log entry records what the release turned on rather than what it
shipped: that a split whose errors you read is a development set from then
on, which is why the holdout was drawn and why the prompt written from the
test split lost on it; that the judge's inputs were wrong before its prompt
was, so every agentic verdict was reached without the schema and a model
that ran out of tokens looked like one that said no; that a score has to
carry the judge that gave it; that re-judging re-evaluates everything, so
August's artifacts moved under September's dependencies; and that a rate
limit can belong to the client rather than the workload.

The attic's README no longer names an open plan.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Signed-off-by: Oktie Hassanzadeh <oktie@oktie.com>

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Judge inputs and config previews can misrepresent content, MySQL initialization mishandles valid secrets, and Beaver documentation remains contradictory.

Get a fresh assessment by requesting another Copilot review.

Review effort: Balanced
Findings: 1 High severity · 2 Medium severity · 1 Low severity

Open (4)
What changed in this PR

Prepares release 1.6.0 with safer benchmark publishing, improved LLM judging, connection lifecycle fixes, deployment hardening, and dashboard enhancements.

Changes:

  • Gates Beaver details behind sign-in and prevents restricted-result publication.
  • Updates judge models, provenance tracking, prompt construction, and watsonx client reuse.
  • Fixes SQLite leaks and adds release/deployment safeguards and dashboard navigation improvements.
File Description
tests/​test_watsonx_client_reuse.py Tests watsonx handle caching.
tests/​test_upload_results_script.py Tests restricted upload filtering.
tests/​test_security_hardening.py Updates packaged-config checks.
tests/​test_judge_reasoning_models.py Tests reasoning replies and judge inputs.
tests/​test_judge_endpoint.py Uses the new no-GT config.
tests/​test_judge_config_storage.py Updates packaged config expectations.
tests/​test_judge_config_retired.py Tests retired-config compatibility.
tests/​test_index_concurrency.py Tests connection reaping.
tests/​test_evaluate_prediction.py Tests digest-aware verdict reuse.
tests/​test_connection_leaks.py Tests SQLite connection closure.
src/​text2sql_eval_toolkit/​ui/​routers_judge.py Resolves retired judge names.
src/​text2sql_eval_toolkit/​ui/​routers_benchmarks.py Exposes and preserves access flags.
src/​text2sql_eval_toolkit/​ui/​models.py Adds benchmark lock metadata.
src/​text2sql_eval_toolkit/​ui/​middleware.py Enforces restricted access and noindex.
src/​text2sql_eval_toolkit/​ui/​judge_budget.py Closes ledger connections.
src/​text2sql_eval_toolkit/​inference/​inference_tools.py Reuses watsonx handles and rejects empty text.
src/​text2sql_eval_toolkit/​indexing/​store.py Reaps retired-thread connections.
src/​text2sql_eval_toolkit/​indexing/​builder.py Closes staleness-check connections.
src/​text2sql_eval_toolkit/​evaluation/​metric_definitions.py Documents judge-config digests.
src/​text2sql_eval_toolkit/​evaluation/​llm_judge_config/​llm_judge_no_gt.yaml Adds the new no-GT judge.
src/​text2sql_eval_toolkit/​evaluation/​llm_judge_config/​llm_judge_no_gt_v2.yaml Removes the retired config.
src/​text2sql_eval_toolkit/​evaluation/​llm_judge_config/​llm_judge_default_config.yaml Replaces the default judge.
src/​text2sql_eval_toolkit/​evaluation/​llm_as_judge.py Adds digests and retired-name mapping.
src/​text2sql_eval_toolkit/​evaluation/​judge_inputs.py Centralizes judge input rendering.
src/​text2sql_eval_toolkit/​data/​test-benchmarks.json Removes the Beaver test subset.
src/​text2sql_eval_toolkit/​data/​benchmarks.json Marks Beaver as restricted.
src/​text2sql_eval_toolkit/​analysis/​error_analysis.py Hides digest metadata from reports.
scripts/​evaluation/​rerun_metrics.py Adds configurable verdict reuse.
scripts/​analysis/​llm_judge_comparison.py Uses current judge configs.
pyproject.toml Bumps version and anyio floor.
docs/​guide/​llm-judge.md Documents new judges and provenance.
docs/​dashboard/​deployment.md Documents access and leak behavior.
docs/​dashboard/​capability-tiers.md Documents benchmark restrictions.
docs/​attic/​README.md Retires the 1.6.0 plan.
deploy/​sql/​mysql-init/​10-readonly-role.sh Creates configurable read-only grants.
deploy/​load-bird-postgres.sh Clarifies dump configuration.
deploy/​env.deploy.example Adds restriction and database settings.
deploy/​docker-compose.yml Raises file limits and passes settings.
data/​test-benchmarks.json Removes the Beaver test subset.
data/​results/​spider_realistic-predictions_eval_summary.md Refreshes published metrics.
data/​results/​spider_dev-predictions_eval_summary.md Refreshes published metrics.
data/​results/​bird_mini_dev_postgres-predictions_eval_summary.md Refreshes published metrics.
data/​judge_calibration/​README.md Documents judge calibration.
data/​judge_calibration/​configs/​llama_no_gt_v1.yaml Preserves a calibration baseline.
data/​judge_calibration/​configs/​llama_default.yaml Updates the calibration baseline.
data/​benchmarks/​test_benchmarks/​results/​README.md Removes Beaver test results.
data/​benchmarks/​README.md Documents gated Beaver access.
data/​benchmarks/​dbs/​README.md Updates Beaver database setup.
data/​benchmarks.json Marks Beaver as restricted.
dashboard/​src/​views/​SignInRequired.tsx Adds a sign-in-required view.
dashboard/​src/​views/​SignInRequired.test.tsx Tests sign-in navigation.
dashboard/​src/​views/​RunEvaluationView.tsx Improves playground layout.
dashboard/​src/​views/​PipelineDetailView.tsx Links records to the playground.
dashboard/​src/​views/​OpenInPlaygroundButton.tsx Adds a reusable playground link.
dashboard/​src/​views/​JudgePlayground.tsx Displays selected judge configuration.
dashboard/​src/​views/​ErrorAnalysis.tsx Links errors to the playground.
dashboard/​src/​views/​BenchmarkTiles.tsx Labels restricted benchmarks.
dashboard/​src/​views/​BenchmarkDetail.tsx Falls back to public overall scores.
dashboard/​src/​views/​BenchmarkDetail.test.tsx Tests summary fallback.
dashboard/​src/​types/​benchmark.ts Adds restriction fields.
dashboard/​src/​services/​benchmarks.ts Recognizes sign-in-required responses.
dashboard/​src/​pages/​App.tsx Routes locked views to sign-in.
dashboard/​src/​lib/​api.ts Adds a typed sign-in error.
dashboard/​src/​lib/​api.test.ts Tests sign-in error handling.
dashboard/​e2e/​shareable-urls.spec.ts Tests playground links.
dashboard/​dist/​index.html Points to the rebuilt frontend.
dashboard/​dist/​assets/​TableRow-Td9PgCqq.js Adds rebuilt table chunk.
dashboard/​dist/​assets/​TableRow-B81J3hyC.js Removes obsolete table chunk.
dashboard/​dist/​assets/​TableHeader-CH8CR35j.js Rebuilds table-header chunk.
dashboard/​dist/​assets/​TableContainer-O9G7CXka.js Adds rebuilt container chunk.
dashboard/​dist/​assets/​TableContainer-B7wsClud.js Removes obsolete container chunk.
dashboard/​dist/​assets/​swimlanesDiagram-VR7AAH4N-JB22Rs9l.js Adds rebuilt diagram chunk.
dashboard/​dist/​assets/​swimlanesDiagram-VR7AAH4N-D3ZXWkB1.js Removes obsolete diagram chunk.
dashboard/​dist/​assets/​stateDiagram-v2-MP3YSRHH-FQIrQuHa.js Removes obsolete state chunk.
dashboard/​dist/​assets/​stateDiagram-v2-MP3YSRHH-DwmnbYBy.js Adds rebuilt state chunk.
dashboard/​dist/​assets/​sizeCapture-INFHLROL-D3y8d2OJ.js Refreshes generated imports.
dashboard/​dist/​assets/​ResultTableView-D7zeyOYh.js Removes obsolete result chunk.
dashboard/​dist/​assets/​railroadDiagram-O6MQD6OU-D3WtqEwj.js Refreshes generated imports.
dashboard/​dist/​assets/​pieDiagram-E7YTZNPT-DF2dAtj3.js Refreshes generated imports.
dashboard/​dist/​assets/​pegDiagram-XKGWAZYB-C43s2JK-.js Refreshes generated imports.
dashboard/​dist/​assets/​OpenInPlaygroundButton-CCS_tmfs.js Adds rebuilt playground chunk.
dashboard/​dist/​assets/​metricInsightsSelect-Ca-YJ6Xf.js Refreshes generated imports.
dashboard/​dist/​assets/​infoDiagram-27XIBGKW-Bi6ru0tL.js Refreshes generated imports.
dashboard/​dist/​assets/​flowDiagram-HODETNUW-BZ_5n_mm.js Refreshes generated imports.
dashboard/​dist/​assets/​ebnfDiagram-PWID7BFC-u2JqIZgy.js Refreshes generated imports.
dashboard/​dist/​assets/​classDiagram-ZZMXUADV-Syw15DOs.js Removes obsolete class chunk.
dashboard/​dist/​assets/​classDiagram-ZZMXUADV-BUlv72DX.js Adds rebuilt class chunk.
dashboard/​dist/​assets/​classDiagram-v2-VYDZK3BY-Syw15DOs.js Removes obsolete class-v2 chunk.
dashboard/​dist/​assets/​classDiagram-v2-VYDZK3BY-BUlv72DX.js Adds rebuilt class-v2 chunk.
dashboard/​dist/​assets/​chunk-XXDRQBXY-B_g5D6EZ.js Refreshes generated chunk.
dashboard/​dist/​assets/​chunk-SVP7TREG-DLSceb9V.js Refreshes generated chunk.
dashboard/​dist/​assets/​chunk-POPQ4Y6H-DPe82o4u.js Refreshes generated chunk.
dashboard/​dist/​assets/​chunk-JWPE2WC7-B1bCvL2R.js Refreshes generated chunk.
dashboard/​dist/​assets/​chunk-F27PBJKO-Slp3VtyJ.js Refreshes generated chunk.
dashboard/​dist/​assets/​chunk-5VM5RSS4-CeSGRynF.js Refreshes generated chunk.
dashboard/​dist/​assets/​chunk-2Q5K7J3B-yhnntLfj.js Refreshes generated chunk.
dashboard/​dist/​assets/​channel-YHNFbJNY.js Removes obsolete channel chunk.
dashboard/​dist/​assets/​channel-luFgzhf3.js Adds rebuilt channel chunk.
dashboard/​dist/​assets/​bucket-0-Bdefr4Uj.js Refreshes generated imports.
dashboard/​dist/​assets/​abnfDiagram-VCTEODGH-D99xhN_E.js Refreshes generated imports.
CONTRIBUTING.md Documents release rehearsals.
.gitignore Excludes gated Beaver artifacts.
.github/​workflows/​ci.yml Checks the app file limit.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread deploy/sql/mysql-init/10-readonly-role.sh Outdated
Comment thread dashboard/src/views/JudgePlayground.tsx
Comment thread src/text2sql_eval_toolkit/evaluation/judge_inputs.py Outdated
Comment thread data/benchmarks/dbs/README.md
- The MySQL read-only role script interpolated the password and each
  database name straight into SQL. A password containing an apostrophe
  produced a malformed IDENTIFIED BY, and a backtick in a database name
  would have ended the quoted identifier. The password is now escaped for
  a string literal, and a database name that is not a plain identifier is
  refused rather than emitted.
- The judge playground kept the previous config's YAML and model on screen
  while the new one loaded, under the new config's name. Both are cleared
  before the request.
- Every later message in an agentic trace was suffixed with "...", whether
  or not it had reached the 500-character limit, telling the judge that
  what it had been given was incomplete. The marker is added only where
  something was removed. data/judge_calibration/README.md notes that the
  stored inputs predate this.
- data/benchmarks/dbs/README.md said Beaver's databases are gated and then
  described public dumps. The coverage figure stays; the contradiction is
  gone.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Signed-off-by: Oktie Hassanzadeh <oktie@oktie.com>

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Public uploads can expose environment-restricted benchmarks, and the sign-in page can incorrectly hide authentication.

Get a fresh assessment by requesting another Copilot review.

Review effort: Balanced
Findings: 1 High severity · 1 Medium severity

Open (2)
Resolved since last review (4)

Comment thread scripts/curation/upload_results_to_hub.py
Comment thread dashboard/src/views/SignInRequired.tsx Outdated
- The upload script decided what to publish from the registry flag alone,
  so a benchmark marked restricted only through
  TEXT2SQL_SIGN_IN_BENCHMARKS -- the other way a deployment marks one, and
  the one that needs no file edit -- would have had its predictions,
  evaluation, errors report and category charts uploaded from that same
  host. The environment's ids are now restricted too, and ids are compared
  without case, since on a case-insensitive filesystem `BEAVER-…` opens
  Beaver's files.
- SignInRequired treated "not asked yet" as "this server has no sign-in":
  the claim appeared on first render and stayed there if the deployment
  request failed, leaving no way to authenticate. It now distinguishes
  asking, answered and failed, claims sign-in is unavailable only when the
  server said so, and offers sign-in anyway when the request failed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Signed-off-by: Oktie Hassanzadeh <oktie@oktie.com>
@oktie
oktie requested a balanced review from Copilot September 20, 2026 05:17

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@oktie
oktie requested a balanced review from Copilot September 20, 2026 13:29

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

- A failed judge call erased the verdict already stored. Re-judging under a
  new config declines to reuse it, so a call the provider refused left the
  prediction with no llm_score -- and compute_summary averages a missing
  score as 0, which is how a run interrupted by a rate limit would have
  published scores far below the ones the judge gave. The stored verdict is
  kept, under the digest of the config that gave it, alongside the error,
  and a run with any judge errors says so.
- Keeping an agent's task whole put no bound on the judge's prompt: one
  Beaver trace runs to 160,000 characters, which a judge config on a
  smaller-context model cannot take -- and each such failure then scored 0.
  The task is kept to 40,000 characters, longer than 99% of the traces in
  data/judge_calibration/, and cut with a marker beyond that.
- The playground listed the 64-character judge digest as a metric, because
  its table is built from the evaluation's own keys. It is shown under the
  LLM judge section, shortened, with the whole value in the title.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Signed-off-by: Oktie Hassanzadeh <oktie@oktie.com>
@oktie

oktie commented Sep 20, 2026 •

Copy link
Copy Markdown
Member Author

Did another round of review. Three findings, all fixed in 8050cf1.

A failed judge call erased the verdict already stored. Re-judging under a new config declines to reuse the stored verdict, so a call the provider refused left the prediction with no llm_score at all — and compute_summary averages a missing score as 0 (one Yes plus one judge error gives 0.5). This release's own re-judge hit exactly that: the provider refused 1,760 of Spider Dev's 2,051 calls, and publishing at that point would have dropped its LLM score from ~0.94 to ~0.13 under the new config's name. The stored verdict is now kept, with the digest of the config that gave it, alongside the error, and a run with any judge errors logs a warning.

Keeping an agent's task whole left the judge's prompt unbounded. That fix was right — agentic predictions were being judged without their schema — but one Beaver trace runs to 160,000 characters, about 40k tokens. gpt-oss-120b takes it; a judge on a smaller-context model (the retired Llama config, still loadable by its old name, or a user's own) would fail on exactly the records with the richest traces, and each failure then scored 0. The task is now kept to 40,000 characters, longer than 99% of the traces in data/judge_calibration/, and cut with a marker beyond that.

The playground listed the judge digest as a metric. Its metric table is built from the evaluation's own keys, so every judged record showed a 64-character hex string next to llm_score. It now appears under the LLM judge section, shortened to 12 characters with the whole value in the title.

Checks: 1,197 Python tests, 212 dashboard tests, lint clean, bundle rebuilt. Tests were added for the first two.

Not changed: the published results keep the verdicts they have. The prompt cap changes what a rebuild of the calibration inputs would produce, which data/judge_calibration/README.md records.

Both came out of the last commit rather than the release:

- The judge prompt was still unbounded. The task cap applied per message,
  so a first step carrying three long messages rendered 120,000 characters
  against a 40,000 limit. The task's messages now share one budget, and
  each keeps at least TRACE_MESSAGE_CHARS so the question survives a schema
  that spent the budget before it.
- A summary could contradict itself. A verdict kept through a failed call
  is in the llm_score average, but num_correct_llm required the absence of
  an error, so a pipeline could read "1.00 average, 1 of 2 correct". The
  count now follows the score; a failure with nothing stored has no score
  and is counted by neither.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Signed-off-by: Oktie Hassanzadeh <oktie@oktie.com>
@oktie

oktie commented Sep 20, 2026

Copy link
Copy Markdown
Member Author

Second review pass, before merge. Two findings, both from the previous round of fixes rather than the release itself; fixed in e2fb3b3.

The judge prompt was still unbounded. The 40,000-character task cap applied per message, not to the task, so a first step carrying three long messages rendered 120,000 characters — verified by rendering one. The task's messages now share a single budget, and each keeps at least 500 characters so the question still appears after a schema that spent the budget.

A summary could contradict itself. A verdict kept through a failed call is in the llm_score average, but num_correct_llm required the absence of an error — so a pipeline could read "1.00 average, 1 of 2 correct", and the CSV's "Number of Correct Results According to LLM Judge" would disagree with the average beside it. The count now follows the score; a failure with nothing stored has no score and is counted by neither.

Also re-checked this pass, nothing found: the calibration script's report and resume paths, the CI compose assertions, the registry's requires_sign_in entries, and the dashboard digest change — verified in a browser against a judged record, which shows "Judged by config digest f8f45e240c42" with the full value in the title and no digest row in the metrics table.

Checks: 1,199 Python tests, 212 dashboard tests, lint clean, bundle rebuilt.

A stored `llm_judge_error` says the last judge call failed. It does not say
there is no verdict: a run that fails keeps the one an earlier run stored,
beside the error. `_is_judge_verdict` answered both questions at once, so:

- The second refusal in a row dropped the verdict the first one had kept, and
  the prediction ended with no `llm_score` at all -- which a summary counts as
  0. That is the collapse the handler exists to prevent, and it came back on
  the retry the run itself tells the operator to do.
- The "verdicts were kept from a different judge config" warning counted only
  verdicts with no error beside them, which is never how one of these is
  stored. Under the default `llm_judge_reuse="matching"` it never fired, so a
  run the provider refused published old-config scores labelled with the new
  one without a word.

Split the two questions with `past_failure`, defaulting to the old answer so
the reuse path still retries a failed prediction rather than settling for what
is stored.

Separately, `_check_restricted_summaries` built its path from the restricted
id's spelling while `_is_publishable` decides what to upload by the file's.
With `TEXT2SQL_SIGN_IN_BENCHMARKS=Beaver` on a case-sensitive filesystem the
check found nothing, passed, and let the real `beaver-...md` through with its
`## Category:` sections -- the one publishable artifact that can still carry
gated derived data. Casefold the ids at the source, as `ui/benchmark_access.py`
does, and match the reports on disk. The nested-layout manifest check casefolds
too, or a directory the upload skips would be named in the manifest and
`results fetch` would fail on it.

Six tests; five fail without these changes. The sixth pins that the first fix
did not over-correct into reusing a stale verdict instead of asking again.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Signed-off-by: Oktie Hassanzadeh <oktie@oktie.com>
@oktie

oktie commented Sep 20, 2026

Copy link
Copy Markdown
Member Author

Review round 3 — three findings, fixed in 67642bf

A full pass over the ~4,400 lines of actual code in the diff (Python source, scripts, dashboard sources, deploy/, CI), skipping regenerated result artifacts and dashboard/dist.

Findings

1. A second consecutive judge failure erased the verdict the first one kept — evaluation/evaluation_tools.py

_is_judge_verdict treated any evaluation carrying llm_judge_error as having no verdict. But a stored error says the last call failed; the verdict beside it is the one an earlier run kept.

So: re-judge under a new config, watsonx refuses N calls. Run 1 stores {llm_score, llm_explanation, llm_judge_config_digest: <old>, llm_judge_error} — the verdict survives, as intended — and logs "Run again to judge them." Run 2 does exactly that: the reuse path correctly declines (so the call is retried), the retry is refused again, and the except handler's _stored_verdict(..., "any") also declines, for the same reason. The prediction ends with no llm_score at all, which compute_summary sums as NaN over num_records — i.e. 0.

That is precisely the published-scores collapse the handler was written to prevent, reappearing on the retry the run recommends. Reproduced directly: run 1 → llm_score: 1.0; run 2 → no llm_score.

2. The gated-category guard missed when the restricted id's case differed — scripts/curation/upload_results_to_hub.py

_check_restricted_summaries built its path from the restricted id's spelling, while _is_publishable decides what to upload by the file's spelling (casefolded). With TEXT2SQL_SIGN_IN_BENCHMARKS=Beaver on a case-sensitive filesystem — a Linux release host or CI — the guard stats Beaver-predictions_eval_summary.md, finds nothing, and raises nothing; _publishable_files then matches and uploads the real beaver-predictions_eval_summary.md with its ## Category: sections, which are derived from the gated ground-truth SQL. That report is the one publishable artifact that can still carry gated derived data, and this refusal is the only thing standing in front of it.

On macOS the path lookup is case-insensitive, so the guard fires by accident and the bug is invisible locally.

3. The "kept from a different judge config" warning never fired under the default mode — evaluation/evaluation_tools.py

Same root cause as 1: the foreign counter filtered on _is_judge_verdict, which is false whenever llm_judge_error is present — and that is always how an error-kept verdict is stored. The comment claiming it was "only reachable with llm_judge_reuse='any'" was wrong; the default "matching" reaches it through the exception handler. Net effect: a run whose judge the provider refused averaged old-config scores into llm_score, recorded the new config in the summary, and the one warning whose job is to say so printed nothing. Lower severity than 1 — the judge-errors warning does mention it in prose — but the counter is the part that names the provenance problem.

The fixes

Findings 1 and 3 share a cause, so they share a fix: _is_judge_verdict was answering two different questions at once. They are now split by an explicit past_failure flag that defaults to the old answer, so the reuse path still retries a failed prediction rather than settling for what is stored — the failed-call handler and the foreign counter are the only callers that opt in.

Finding 2 is fixed at the source: _restricted_benchmarks now casefolds, as ui/benchmark_access.py already does, and _check_restricted_summaries globs the reports on disk and matches them instead of constructing a path — so it asks about the same file the uploader would publish. The nested-layout manifest comparison casefolds too; without that, casefolding the source would have let a differently-cased directory into manifest.json while the upload skipped it, and results fetch would fail on the missing file.

Verification

Six new tests. Five fail against the unfixed source (verified by stashing the two source files and re-running). The sixth, test_a_verdict_kept_through_a_failure_is_still_retried, passes either way by design — it pins that the fix did not over-correct into reusing a stale verdict instead of asking again.

Because macOS resolves paths case-insensitively, the case test asserts on the filename in the refusal message rather than on the refusal itself, which distinguishes old from new on any platform.

Full suite 1205 passed (was 1199); ruff check and black --check clean; tsc --noEmit and 212 vitest tests clean; dashboard/dist confirmed rebuilt from current sources.

Checked and found sound

  • The sign-in wall (ui/benchmark_access.py) is well built: it runs ahead of the tier gate, reads declared route parameters rather than URL text, fails closed on unlisted routes, casefolds, and rejects ids outside the id alphabet before they can reach a filename. unclassified_params plus the route-enumeration tests genuinely guard it rather than decorating it. /api/compare and /api/benchmarks/{id}/summary do return overall metrics only, and /api/static cannot reach result files.
  • Connection-leak fixes — closing() in builder.is_stale, the _connect contextmanager in judge_budget, thread-reaping in EvalIndex — are correct, with no lock-ordering or deadlock problem in reserve().
  • Both new judge prompt templates .format() cleanly against exactly the keys build_llm_judge_inputs supplies.
  • upload_large_folder with an explicit allow_patterns list is safe on the empty-list edge: filter_repo_objects checks is not None, not truthiness, so an empty results dir uploads nothing rather than everything.
  • Sharing one ModelInference across threads is fine — the SDK is on httpx.Client, which is thread-safe and which IBM documents as intended to be shared.

🤖 Generated with Claude Code

`data/judge/usage.sqlite` holds the dashboard's LLM-judge spend counters and
verdict cache, both written at runtime. `JudgeStore` creates the file, its
parent directory and its schema on first use, so the tracked copy seeded
nothing -- and it had gone stale, missing the `user_caps` table a fresh
database now creates.

Tracking it meant every local use of the judge left a pending change carrying
a user hash, a model, a cost and a judge's explanation, one `git commit -a`
away from being published. `data/judge/` is now ignored.

A deployment is unaffected: its database lives in the `data` volume that
docker-compose mounts at /data, not in the repository.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Signed-off-by: Oktie Hassanzadeh <oktie@oktie.com>
@oktie
oktie merged commit ddb9580 into main Sep 20, 2026
23 checks passed
@oktie
oktie deleted the 1.6.0 branch September 20, 2026 21:46
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants