Unofficial showcase. The authoritative EN 16931 code lists and validation artefacts are those the European Commission publishes in its Registry of supporting artefacts to implement EN16931. Where this repository or its reports differ, the Registry prevails.
Report: EN 16931 code lists: declared and implemented, on this repository's GitHub Pages site, svanteschubert.github.io/EU-Codelist-Normalizer. For every effective date it shows whether the spreadsheet, the Genericode files and the Index sheet's change notes agree, and whether the CEN validator implements the declared codes; each finding links to where it is written.
Normalize the EU's downloaded Genericode code lists into consistently ordered,
formatted Genericode XML, and extract and sort EN16931 spreadsheet sheets as CSV.
Releases are grouped by version and effective date under src/test/resources/.
Each release contains numbered revision subdirectories, each with extracted/
and normalized/ directories. Comparisons and change reports are future work.
EU-Codelist-Downloader
acquires and archives the official artefacts. This repository reads its registry, ZIP
files and workbooks locally, without changing or downloading anything in that sibling
repository. What this repository makes of them stays here: the extracted and
normalized releases in src/test/resources/, and the comparison report, which links
to their lines, in docs/, served by GitHub Pages.
Normalization uses Philip Helger's
com.helger:ph-genericode:8.1.0
as a released Maven dependency. Local Java classes add the EU-specific source
selection, ordering, directory layout and provenance. No locally installed fork
or changes to the upstream library are required.
Spreadsheet extraction uses org.apache.poi:poi-ooxml:5.5.1.
Every sheet is exported in its original row and column order, omitting completely
empty rows. A separate copy
sorts code-list rows using the same base-36 ordering as the Genericode output.
The starting point is Svante Schubert's and Philip Helger's
Genericode10EN16931CodeListMarshallerTest in
svanteschubert/ph-genericode,
including its 2024-05-15_13/EAS.gc output.
Use this sibling layout:
GitHub/
├── EU-Codelist-Downloader/
│ └── src/main/resources/
│ ├── downloaded-files.json
│ └── downloaded-files/
│ ├── EN 16931 code list - GeneriCode/*.zip
│ └── EN 16931 code list - XLSX/*.xlsx
└── EU-Codelist-Normalizer/
Maven 3.6+ and Java 25 are required. The project pins jenv to Java 25 through
.java-version, matching the downloader. The shell script resolves JAVA_HOME through jenv when installed.
./run-normalize.shThe script works from any working directory, builds and tests the application,
then reads ../EU-Codelist-Downloader. After building, processing can also run
directly, without Maven or network access:
java -jar target/eu-codelist-normalizer-all.jar
java -jar target/eu-codelist-normalizer-all.jar --helpOptional paths (relative to the normalizer repository when using the script, or the current working directory when invoking Java directly):
./run-normalize.sh --downloader ../EU-Codelist-Downloader --output src/test/resourcesMaven needs network access on the first build to resolve its dependencies. The normalizer itself performs no network downloads.
src/test/resources/
├── normalization.json
├── 01_2019-03-15/
│ └── r01/
│ ├── extracted/xlsx/...
│ └── normalized/xlsx/...
├── ...
├── 13_2024-05-15/
│ └── r01/
│ ├── extracted/
│ │ ├── gc/
│ │ │ ├── EAS.gc
│ │ │ ├── ...
│ │ │ └── source.json
│ │ └── xlsx/
│ │ ├── EAS.csv
│ │ ├── ...
│ │ └── source.json
│ └── normalized/
│ ├── gc/...
│ └── xlsx/...
├── ...
└── 17_2026-05-15/
├── r01/ # Original GC archive and v17 workbook
│ ├── extracted/{gc,xlsx}/...
│ └── normalized/{gc,xlsx}/...
└── r02/ # Replacement GC archive and v17b workbook
├── extracted/{gc,xlsx}/...
└── normalized/{gc,xlsx}/...
Directories use <bundle-version>_<effective-date>/r<revision>/, with numeric
version and revision numbers padded to at least two digits. Original version
values in source metadata and revision assignments remain unchanged.
The version is the EN16931 bundle version from the downloader's
registry, not the independent EAS or VATEX version. The date is the registry's
effective date, which can differ from the ZIP filename or internal XML version.
Original identification metadata inside the XML remains intact. Only downloaded
EN16931 Genericode ZIPs and EN16931 XLSX bundles are processed. Releases available
only as spreadsheets have an xlsx/ directory within each stage. The separate EAS and VATEX
workbook categories have independent version numbers and are not included;
their sheets within the EN16931 bundles are extracted.
Each revision separates original-order extracted/ files from sorted
normalized/ files. Each stage contains Genericode files in gc/ and spreadsheet
CSVs in the sibling xlsx/, with a source manifest in each format directory.
Extracted Genericode files are byte-for-byte copies of the original ZIP entries.
All downloaded records in these two categories are included, including
superseded records. Revisions are sibling subdirectories within their release; no hash-based
revisions/ or variants/ directories are generated.
src/main/resources/release-revisions.json
explicitly maps source SHA-256 values to a revision for each effective date,
bundle version and format. It is bundled in the executable JAR. For example:
{
"effective_date": "2026-05-15",
"version": "17",
"revision": 2,
"sources": { "gc": "<source-sha256>", "xlsx": "<source-sha256>" }
}The file has format_version: 1 and an assignments array containing these
entries. Revision numbers are local identifiers, not official attachment upload
numbers. The mappings pair release 17's original sources as r01 and their
replacements as r02. Version 6's original and "updated" workbooks are r01 and
r02, respectively; no official supersession relationship is invented for them.
Explicit assignments take precedence even when only one source is downloaded.
Unmapped releases with at most one source per format default to r01. If any
format has multiple sources, every source in that release must be assigned.
Missing or conflicting assignments fail before publication. Add explicit mappings
for new ambiguous releases and rebuild; upload numbers, timestamps and registry
order never determine the grouping. Missing formats are not copied from another
revision, and there is no duplicate latest directory.
Each source.json records the source URL, path relative to the downloader, source
SHA-256, version/date, archive entry or sheet names, output hashes and row counts.
Source manifests also retain original filenames, release and revision directories,
revision numbers and official superseded_by links when present. Spreadsheet manifests
record column counts, emitted row counts and whether they are normalized.
The root normalization.json (format version 5) identifies all generated archive and workbook
directories and their processing policies. Manifests contain no run timestamps
or absolute local paths.
Outputs are intentionally outside target/, survive mvn clean, and can be
versioned in Git. A rerun retains byte-identical files without rewriting them.
The normalizer validates all existing manifest-tracked files before generating
output in staging. After successful publication, it removes obsolete tracked
files and their empty directories, including previous flat revision and hash-based
layouts. Unrelated files and regression fixtures are preserved. Modified or
missing tracked files and untracked destination collisions fail before
publication; the error identifies the path that needs attention.
The branch code-history holds one commit per effective date, oldest first, so that
Git shows what changed from one release to the next:
./build-history-branch.sh
git diff code-lists-16_2025-11-15 code-lists-17_2026-05-15 -- gc/Currency.gc
git log --oneline code-history -- xlsx/ICD.csvEach commit holds the normalized/ files in force from its date at fixed paths,
xlsx/ and gc/: those of the last revision of its release, so 17_2026-05-15
holds the corrected r02. Its RELEASE.md names the original files with their URL
and SHA-256, and the earlier revisions they replace. It is dated by the effective
date and tagged code-lists-<release>. A release without Genericode files, such as
08_2021-11-15, keeps those of the last release that had them, as its RELEASE.md
and commit message state, so the Genericode diff spans the gap instead of deleting
and re-adding every file.
The branch is derived from the committed release data, so the script needs no
options and works out what to do by itself. It keeps every commit that is still
right, the same date from the same revisions with the same files, and never adds a
date twice. A new release is appended. A correction, such as a new revision of a
release, rebuilds the branch from that date on, and a branch built one commit per
revision is rebuilt as a whole. The script prints the commands that publish the
result: a plain push when it only appended, a force push and the tags to replace
when it rebuilt. To keep the branch up to date automatically, install a
post-commit hook; after every commit that changes release data it runs the
script, and stays quiet otherwise:
./build-history-branch.sh --install-hook
./build-history-branch.sh --remove-hookThe branch is built in a temporary worktree, never in your checkout, from committed
data only, and never pushed. The commits use your Git configuration and are signed
when commit.gpgsign is set.
- Select downloaded Genericode ZIPs from
downloaded-files.jsonand verify their stored SHA-256 before reading them. - Read each
.gcwith the Genericode 1.0 marshaller fromph-genericode. - Require a nonempty, unique
Codefor every row. - Sort alphanumeric codes by their base-36 value, following the earlier example. Use arbitrary precision to avoid integer overflow; break numeric ties using the original code string. Codes containing other characters follow, ordered lexically. This gives a consistent total order for mixed code formats.
- Write UTF-8, formatted Genericode XML using the
gcnamespace prefix.
Codes remain strings: leading zeros and case are preserved. Names, remarks, annotations, identification, column definitions and other data are retained; descriptions are not trimmed, rewritten or translated. This is application-level normalization, not W3C XML canonicalization (C14N).
For example, CD (445 in base 10) sorts before CBB (15959) and CEC (16068).
This numeric ordering applies to both normalized Genericode and spreadsheet
code lists. The extracted files retain their original order. Length alone only
determines base-36 order when codes have no leading zeros.
All archives and workbooks are processed in a temporary directory before output publication. Missing files, hash mismatches, invalid XML/workbooks, missing/duplicate Genericode codes and unresolved output collisions fail the run with a nonzero exit status. Invalid input does not publish partial results. Output inside the downloader repository is rejected.
- Verify each downloaded EN16931 XLSX workbook's SHA-256 and open it read-only with Apache POI.
- Export every sheet, including index, notes, hidden and empty sheets, to
<sheet-name>.csv. Keep the original sheet, row and column order, headings, duplicate values and cell whitespace inextracted/xlsx/, except for wholly empty rows as described below. - Write UTF-8 without a BOM, with commas between fields and double quotes around
every field, including empty fields. Escape a double quote as
""and preserve embedded commas and line breaks. CSV records use LF line endings. - Omit rows whose formatted cells are all empty or whitespace-only, including leading, internal and trailing empty rows and empty rows in documentation sheets. Retain partially empty rows, zero values and rows containing data even when the code is empty. Keep column positions and pad shorter rows with empty quoted fields so every row in a sheet has the same column count. A sheet with no retained rows produces an empty file. Embedded line breaks within populated cells are preserved; filtering operates on cells, not lines.
- Use POI's
DataFormatterwith a fixed US locale and CSV mode for displayed values, including leading-zero number formats and dates. Formula cells use their saved cached results without recalculation; those results may be stale if the source workbook was not recalculated before saving. CSV retains cell values, not workbook formatting, formulas, comments or images. - Record in
source.json, per sheet, the workbook rows left out as blank ("omitted_rows": [5, 26]), so that every CSV record can be traced back to its row, and a report can cite the cell of the original workbook.
normalized/xlsx/ contains a separate CSV for every extracted sheet, using the
same quoting and encoding. Empty-row removal happens once before writing either
stage; sorting operates on the retained rows. Extraction and spreadsheet
normalization policy identifiers both use v2. Between the two stages, only row
order changes:
- Find the code column in the first three header rows, recognizing
Code,Code Values,Alpha-2 code,Alphabetic Code,AESC,AESand2005 Code. This handles Country, Currency and Unit, whose codes are not in the first column, and sheets such as VAT ID and Time with multiple header rows. - Keep all rows through the code header in place. Sort subsequent rows by the code column using the Genericode comparator. For cross-syntax mapping tables, the first code column determines the order and the entire row moves together.
- Preserve every cell value, including leading zeros, whitespace and duplicates. Identical codes retain their original relative order. Partially populated rows with a blank code stay in their positions within the retained rows.
- Keep
Index,Main, empty sheets and other sheets without a recognized code header in the extracted order, since they have no code-list rows to sort.
The EN 16931 validation artefacts in
ConnectingEurope/eInvoicing-EN16931
enforce the same code lists, spelled out inline in the tests of their BR-CL
Schematron assertions. --validator extracts them for UBL and CII from every
tagged release, normalizes them and compares them with the Genericode files and
spreadsheets of the releases above:
./run-normalize.sh --validator --validator-repo ../eInvoicing-EN16931The run reads the validator checkout through git show <tag>:<path> only, so its
branch, local changes and untracked files do not matter, and it writes
src/test/resources/validator/ (or <--output>/validator/):
validator/
├── validator-index.json # tags, commits, effective dates, source hashes, rules
├── extracted/2026-05-15_validation-1.3.16/{ubl,cii}/EN16931-*-codes.sch
├── normalized/2026-05-15_validation-1.3.16/{ubl,cii}/BR-CL-01.csv
├── summary.csv # per effective date and syntax
├── rules.csv # per effective date, syntax and rule, with the differing codes
├── index.html # the report
└── about.html # how each comparison works, linked from the report
The report keeps the codes the Commission declares apart from the codes the validator implements. It lists every effective date, newest first, as one row of status chips that opens in place:
- Declared, for a date on which a code-list release took effect: the spreadsheet against the Genericode files of the same release, code by code; the Index sheet's change notes against what actually changed; the Index's business terms against EN 16931-1:2017.
- Implemented: each BR-CL rule of the validator in force, against the declared
codes of its list: the Genericode file, or the spreadsheet for the Time list, which
has none. Should the two ever declare different codes, the rule is shown against
each;
rules.csvkeeps both comparisons.
Four tiles above the timeline summarize the latest date. Each block opens on its own,
and the explanations live on about.html rather than on the report.
- Extraction: the codes of an assertion are the union of its
contains(' … ', concat(…))enumerations and its@attr = '…'comparisons (BR-CL-24). UBL's BR-CL-01 (invoice and credit note types) and BR-CL-10 (ICD plusSEPA) therefore compare as one list each. Literals are kept exactly, including case and stray spaces. - Normalization: one quoted
"Code"CSV per rule, codes de-duplicated and in the base-36 order of the normalized Genericode files. - Effective dates:
validator-releases.csvgives the date each tag applies from, taken from the Commission's registry where it lists the release, otherwise from the validator's README. Every date on which either the code lists or the validator changed is compared, each time with what was in force on both sides. The highest revision of a code-list release carrying a format is used. - Renames: a removed code next to the code that replaced it shows as one line,
STD → STN, with a link to the English Wikipedia article on the change. The pairs come fromcode-successions.csv(ISO 4217 redenominations, euro adoptions, the split of the Netherlands Antilles), because the old code is often in no EU release and has no name to match. Codes that differ only in punctuation or case (01'00and0100) and the only two codes with the same name are paired without it. Pairing changes the display, not the counts. - Mapping:
rule-catalog.csvties each rule to its Genericode file and sheet. BR-CL-06 reads the Time sheet's2005 Codecolumn for UBL and2475 Codefor CII. Unmapped rules are reported, not dropped. Before 2021 only spreadsheets were published; a missing component is reported as such, never as agreement.
The same run also checks the Index sheet of every EN16931 workbook revision,
whose table (row 6 of the workbook) states per tab whether and how the list
changed and which business terms use it:
validator/
├── index-claims.csv # per revision and tab: stated vs. actual changes
├── business-terms.csv # per revision and tab: BTs vs. EN 16931-1:2017
└── index-releases.csv # per revision: stated dates and structure
- Stated changes: the
Changesflag (Yes,No,Fixed) and the free-textRemark on updatesare parsed into added, removed, renamed and deprecated codes and counts ("adding 49 codes"). Words count as codes only when the list has them; missing leading zeros (Adding 0221 to 230) are restored, and code-like words the list lacks (2017for0217,VATEX-135-1) are reported as unresolved. - Spreadsheet against Genericode: within every revision, each sheet is compared with its Genericode file, code by code and name by name.
- Actual changes: every revision's TabName sheet is compared with the previous release, by column role, ignoring whitespace-only edits. Its Genericode file is compared with the latest earlier release that has Genericode, across all Index claims in between, and with the sheet over the same span.
- Business terms: the Index column "EN business terms where the code list is
used." is compared with
business-terms-2017.csv, derived from Table 2 of EN 16931-1:2017, and with the previous release. Genericode files and sheets name no business terms; they are searched forBT-nanyway. - Dates: the effective date the Index states is compared with the date the
release is filed under. The 2019 workbooks state only a publication date, on
Main.
The report is published to this repository's GitHub Pages folder (--report-folder PATH, which may be repeated, to change it):
docs/
├── index.html .nojekyll # hand-written landing page
└── en16931-code-list-comparison/
├── en16931-code-list-comparison.html # the report; open it in a browser
├── about.html # how it is compared, where links go, every file
├── manifest.json # every file with its size and SHA-256
├── summary.csv rules.csv index-claims.csv business-terms.csv index-releases.csv
└── configuration/validator-releases.csv rule-catalog.csv business-terms-2017.csv code-successions.csv
To serve it, enable Settings → Pages → Deploy from a branch, branch master,
folder /docs; the site is then at https://svanteschubert.github.io/EU-Codelist-Normalizer/.
Every finding links to where it is written, or to the file that lacks it:
- EU code lists: a code to its line in the release's
extracted/files on GitHub,src/test/resources/<release>/rNN/extracted/(sheets as CSV with?plain=1#Lnn, Genericode files with#Lnn), a remark or business terms to the tab's row of the Index. The title names the original file and the workbook cell, counted with the omitted rows. A code under What actually changed links instead to its line in thecode-historycommit of that date, dated by the effective date, whose diff shows the change: a new or renamed code on the new side (#diff-…R<n>), a removed one on the old (#diff-…L<n>). Build the branch before the report, so the links find its commits, and push it, so they resolve. Each release links its original XLSX and ZIP in the downloader'sdownloaded-files/. The GitHub addresses and branches come from theoriginof this checkout and of the downloader's; a release tree outside a checkout on GitHub is not linked. - Validator: a code to its line of the Schematron file at the commit the release tag names, so the link stays exact even if a tag moves.
Links are absolute, so the report folder still works when copied to any web server
or file share; index-claims.csv and rules.csv carry the same links.
The report folder, like the validator/ tree, is replaced as a whole on each run and is
byte-identical for the same inputs. A validator/ directory without
validator-index.json is refused, as is a published folder without this generator's
manifest.json of the same kind.
mvn verifyGenericodeNormalizer: the local row-ordering and serialization policy.NormalizationPipeline: registry input, archive verification, versioned output and source manifests.SpreadsheetExtractor: Apache POI sheet-to-CSV extraction.SpreadsheetNormalizer: header-aware sorting by the sheet's code column.ReleaseRevisions: explicit source-to-revision assignments.GeneratedOutputs: validation, publication and cleanup of manifest-tracked files.NormalizerMain: command-line entry point.ValidatorPipeline,SchematronCodeLists,CodeListReleases,ValidatorComparison,ValidatorReport: extraction of the validator's code lists and their comparison with the published ones.ReportFolder: the shareable copy of the report with its manifest.IndexSheet,ChangeClaims,ActualChanges,IndexCheck,BusinessTerms,IndexReport: the Index sheet checks.
Tests cover preservation of values, ordering, repeatability, historical revisions,
source integrity and failure handling. Spreadsheet tests cover quoted UTF-8
values, embedded delimiters/quotes/newlines, blank cells and empty-row removal, hidden and
empty sheets, number/date formats, cached formulas, source preservation and
explicit revision grouping, safe migration, code-column selection and normalization of CD, CBB,
CEC. A regression fixture checks compatibility with the earlier
2024-05-15_13/EAS.gc output. The fixed examples 13-EAS-input.gc and
13-EAS-expected.gc, and the normalization changes they demonstrate, are
explained in 13-EAS-README.md.
GNU Affero General Public License, version 3 or later (AGPL-3.0-or-later); see
LICENSE and NOTICE. The license covers this repository's code
and documentation, with one exception:
GenericodeNormalizer.java
is adapted from Genericode10EN16931CodeListMarshallerTest of
ph-genericode, by Svante Schubert
and Philip Helger, and keeps its copyright notices and its license, the Apache
License 2.0 (LICENSE-APACHE-2.0). The code lists and artefacts
the European Commission publishes, including the extracted and normalized copies
under src/test/resources/, retain their upstream terms, notices and provenance.
ph-genericode and Apache POI are used under the Apache License 2.0.