Cherry pick AddFiles: Add dry_run (#40167) - #40255
Merged
Merged
Conversation
* AddFiles: dry_run reports what the schema pre-pass would do without committing or registering
SchemaEvolutionConfig.setDryRun(true) (provider key dry_run) turns
AddFiles into a report-only transform: the read side runs as usual
(footers, distinct schemas), then DryRunReport emits one Row per
distinct schema plus one summary row on a new output, dry_run_report.
Why: before enabling evolution on a large import, a user wants to know
what the options would do to the table and which files would be
refused, without touching anything. Running the real transform under
FAIL_PIPELINE answers only "would it fail", and only for the first
failure.
The verdicts come from CommitSchemaUnion.plan, the one step that
decides everything a commit does before writing anything. A Plan says
which distinct file schemas are merged (SchemaToMerge), which are
refused and why (IncompatibleSchema), and the schema the table ends
with. It is an EvolutionPlan against an existing table (base schema
snapshot, name-mapping repair) or a CreationPlan when the table is
missing (partition spec and sort order resolved against the union, or
the problem that blocks creation). commitOnce dispatches to evolve or
create, which fail or warn under the handling mode and then write; the
dry run turns the same plan into rows, so a check added to the plan
reaches both and the report cannot drift from the commit. A schema that
is fine against the table but conflicts with another schema of the
input is therefore reported with the blame a real run assigns.
Real-run changes that come with planning first: partition or sort
fields that do not fit the union fail with a message naming them,
before any catalog write and under either handling (the per-file
fallback creation throws the same error, so it could never be routed);
problems are reported before the transaction is opened; every
transaction is checked against the one base snapshot; planning retries
like committing; a window without schemas plans nothing. Settings
(config, the handling resolved for the mode, NewTableSettings) replaces
the three loose arguments both DoFns carried.
Report rows (REPORT_SCHEMA), told apart by row_type:
row_type schema | create | unreadable | unchecked | summary
schema_key short murmur3 key of the schema JSON on schema rows
schema the canonical schema JSON; on the create row, the
union the table would be created with
num_files files the row covers
changes ARRAY<STRING>: the SchemaDelta descriptions on schema
rows; "create <optional|required> <name> <type>" per
column on the create row (pins shown required); the
totals line and table-level changes on the summary
allowed whether a real run would accept it
reason why not, else ""; the consequence on the summary
would_create_table false whenever a real run would not create the
table, including when it would fail first
Provider: the dry_run_report output only exists when dry_run is set,
so existing YAML pipelines that enumerate outputs are unaffected.
* improve retry
* comments
* add untracked file
* move tests
* track file
* comments
Collaborator
Author
|
R: @Amar3tto |
Contributor
|
Stopping reviewer notifications for this pull request: review requested by someone other than the bot, ceding control. If you'd like to restart, comment |
Amar3tto
approved these changes
Sep 24, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
SchemaEvolutionConfig.setDryRun(true) (provider key dry_run) turns AddFiles into a report-only transform: the read side runs as usual (footers, distinct schemas), then DryRunReport emits one Row per distinct schema plus one summary row on a new output, dry_run_report.
Why: before enabling evolution on a large import, a user wants to know what the options would do to the table and which files would be refused, without touching anything. Running the real transform under FAIL_PIPELINE answers only "would it fail", and only for the first failure.
The verdicts come from CommitSchemaUnion.plan, the one step that decides everything a commit does before writing anything. A Plan says which distinct file schemas are merged (SchemaToMerge), which are refused and why (IncompatibleSchema), and the schema the table ends with. It is an EvolutionPlan against an existing table (base schema snapshot, name-mapping repair) or a CreationPlan when the table is missing (partition spec and sort order resolved against the union, or the problem that blocks creation). commitOnce dispatches to evolve or create, which fail or warn under the handling mode and then write; the dry run turns the same plan into rows, so a check added to the plan reaches both and the report cannot drift from the commit. A schema that is fine against the table but conflicts with another schema of the input is therefore reported with the blame a real run assigns.
Real-run changes that come with planning first: partition or sort fields that do not fit the union fail with a message naming them, before any catalog write and under either handling (the per-file fallback creation throws the same error, so it could never be routed); problems are reported before the transaction is opened; every transaction is checked against the one base snapshot; planning retries like committing; a window without schemas plans nothing. Settings (config, the handling resolved for the mode, NewTableSettings) replaces the three loose arguments both DoFns carried.
Report rows (REPORT_SCHEMA), told apart by row_type:
row_type schema | create | unreadable | unchecked | summary
schema_key short murmur3 key of the schema JSON on schema rows
schema the canonical schema JSON; on the create row, the
union the table would be created with
num_files files the row covers
changes ARRAY: the SchemaDelta descriptions on schema
rows; "create <optional|required> " per
column on the create row (pins shown required); the
totals line and table-level changes on the summary
allowed whether a real run would accept it
reason why not, else ""; the consequence on the summary
would_create_table false whenever a real run would not create the
table, including when it would fail first
Provider: the dry_run_report output only exists when dry_run is set, so existing YAML pipelines that enumerate outputs are unaffected.
improve retry
comments
add untracked file
move tests
track file
comments
Please add a meaningful description for your change here
Thank you for your contribution! Follow this checklist to help us incorporate your contribution quickly and easily:
addresses #123), if applicable. This will automatically add a link to the pull request in the issue. If you would like the issue to automatically close on merging the pull request, commentfixes #<ISSUE NUMBER>instead.CHANGES.mdwith noteworthy changes.See the Contributor Guide for more tips on how to make review process smoother.
To check the build health, please visit https://github.com/apache/beam/blob/master/.test-infra/BUILD_STATUS.md
GitHub Actions Tests Status (on master branch)
See CI.md for more information about GitHub Actions CI or the workflows README to see a list of phrases to trigger workflows.