Skip to content

V3 support - #1

Open
jhamman wants to merge 9 commits into
mainfrom
v3-support
Open

jhamman wants to merge 9 commits into
mainfrom
v3-support

Conversation

@jhamman

@jhamman jhamman commented Sep 27, 2026 •

Copy link
Copy Markdown
Member

Part of apache#1818 (V3 tracking issue) and apache#1551 (writing V3 tables).

Rationale for this change

PyIceberg could parse Iceberg format-version 3 metadata but was not a v3 writer, and several v3 reads were broken against tables produced by the Java reference implementation. A gap analysis against Spark 4.0 + Iceberg 1.11 (Java Table API and the compose stack's REST fixture) found, among others: geometry(srid:3857) type strings written by Java could not be parsed (and PyIceberg's own quoted form was rejected or stored with embedded quotes by Java), geo columns could not be scanned, variant tables could not be loaded, predicates on timestamp_ns columns raised, _row_id / _last_updated_sequence_number were not readable, inspect omitted every v3 field, and all v3 writes failed behind three separate guards while leaving orphaned data files.

This branch closes those gaps in stages, one commit each, reusing and crediting the open community PRs where they existed (apache#3551, apache#3623, apache#3624, apache#3972, apache#3474/apache#2822, apache#3478, apache#3992, apache#3630, apache#3659, apache#4002):

  1. Fixtures, interop tests, docs — dev/provision_v3.scala provisions the v3 tables Spark SQL cannot create (nanosecond timestamps, geometry, column defaults, unknown, equality deletes) through the Java Table API; tests/integration/test_v3_interop.py compares PyIceberg against Spark/Java; new format-version-3.md docs page with a support matrix.
  2. Geometry read interop + nanosecond predicates — Java/spec geo type strings, binary-to-geo read resolution, geoarrow mapping, geo literals; TimestampNanoLiteral / TimestamptzNanoLiteral with string/long/date conversions so filters and partition pruning work on ns columns; ns JSON, Avro timestamp-nanos, nanosecond-exact bounds.
  3. v3 writes — metadata serialization, upgrade to v3, ManifestWriterV3 / ManifestListWriterV3 with first-row-id assignment, first-row-id / added-rows / next-row-id on commit, delete manifests, orphan cleanup on failed writes, format-version gating of v3-only schema changes.
  4. Row lineage on read, DV hardening, inspect — _row_id / _last_updated_sequence_number as reserved metadata columns (position-tracked through deletes and filters), first_row_id inheritance, range-based DV reads with length/magic/CRC/cardinality validation, DVs keyed by referenced_data_file with supersession of older position deletes, v3 columns in inspect tables.
  5. Geometry writes, bounds, defaults, unknown, promotions — Parquet GEOMETRY/GEOGRAPHY logical types via geoarrow, bounds in Java's GeospatialBound encoding from Parquet geospatial statistics, write-default on the Parquet path, unknown columns dropped from data files, date -> timestamp promotion gated to v3.
  6. Variant, equality deletes, multi-argument transforms — VariantType with unshredded reads and a pure-Python decoder (pyiceberg.variant), equality-delete reads with null-equal semantics, multi-arg transforms loaded as unknown transforms with spec source-ids serialization.
  7. Deletion-vector writes — PuffinWriter, DV serialization, and a _RowDelta producer so delete() on v3 merge-on-read tables writes one DV per data file, replacing the file's previous DV; a PyIceberg blob is byte-identical to Spark's for the same positions.
  8. Review fixes — defaults survive create_table, set_default_value is gated to v3, REST scan planning carries the v3 file fields, write-default only on writes, next-row-id required, strict parsing of malformed transforms, and copy-on-write rewrites preserve _row_id / _last_updated_sequence_number of copied rows.

Known limits (documented in format-version-3.md): writing variant and reading shredded variant wait on pyarrow's Parquet VARIANT support; overwrite() / upsert() remain copy-on-write; spatial predicates, geo file skipping and table encryption are out of scope. Spark 4.0 + Iceberg 1.11 cannot read timestamp_ns or geometry columns and Java's generic Parquet writer cannot write geometry, so those features are verified through the Java Table API, Parquet schema inspection and PyIceberg round trips rather than Spark reads.

Are these changes tested?

Yes.

  • Unit suite: 4748 passed (new coverage for types, literals, conversions, manifests v3, metadata v3, row lineage arithmetic, Puffin/DV round trips, delete file index, geo bounds, variant decoding, equality deletes, copy-on-write lineage).
  • Integration suite on dev/docker-compose-integration.yml (REST fixture, Hive, Spark Connect): full run 1401 passed, 9 skipped before the review fixes; the affected suites (interop, writes, partitioned writes, deletes, REST schema/scan planning) re-run green after them. Format-version parametrizations in the write, partitioned write, add_files and delete suites now cover [1, 2, 3].
  • Cross-engine checks in tests/integration/test_v3_interop.py: Spark reads PyIceberg-written v3 tables (appends, copy-on-write and merge-on-read deletes, overwrites, upgrades, defaults, unknown) with matching _row_id values and DV metadata; PyIceberg reads Spark/Java-written DV, variant, equality-delete, nanosecond, default, unknown and geometry tables with row-for-row parity.
  • Two standalone interop harnesses used for the gap analysis (local Spark + Java API, and REST) went from 27/63 and 4/13 passing to 62/62 and 13/13.

Are there any user-facing changes?

  • v3 tables can be created, upgraded to (upgrade_table_version(3)) and written on every catalog.
  • New reserved scan columns _row_id and _last_updated_sequence_number via selected_fields.
  • delete() on v3 tables with write.delete.mode=merge-on-read writes deletion vectors instead of rewriting files; v1/v2 tables keep the copy-on-write fallback (warning reworded).
  • geometry / geography type strings now serialize in the spec form (geometry(srid:3857)); the quoted form written by earlier releases is still parsed.
  • New types VariantType (read-only) and literals TimestampNanoLiteral, TimestamptzNanoLiteral, GeometryLiteral, GeographyLiteral; new modules pyiceberg.variant, pyiceberg.table.metadata_columns, pyiceberg.utils.geo.
  • Adding column defaults, set_default_value, and v3-only types or promotions on v1/v2 tables now raise unless allow_incompatible_changes is set; next-row-id is required in v3 metadata; malformed transform names such as bucket[abc] raise again.
  • inspect tables gain first_row_id, referenced_data_file, content_offset and content_size_in_bytes.
  • Docs: new "Format version 3" page, updated geospatial, API and configuration pages.

🤖 Generated with Claude Code

jhamman and others added 7 commits September 27, 2026 06:48
Provision v3 tables that Spark SQL cannot create (nanosecond timestamps,
geometry/geography, column defaults, unknown) through the Iceberg Java
Table API from dev/provision_v3.scala, run inside the Spark container by
dev/provision.py. Add tests/integration/test_v3_interop.py comparing
PyIceberg against the Java reference implementation; known v3 gaps are
strict xfails so each fix flips a test. Add a docs page tracking v3
support and correct the geospatial and nanosecond documentation.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…edicates

Geometry/geography: parse and serialize type strings in the Java/spec form
(geometry(srid:3857), geography(srid:4326, vincenty), case-insensitive,
quoted values still accepted for metadata written by earlier releases),
validate the geography algorithm, resolve binary Parquet storage to the geo
types on read (the Arrow dataset reader does not surface the GEOMETRY
logical type), map geoarrow.wkb extension input, skip bounds for comparison
predicates on geo columns in the metrics evaluators, add geo literals so
equality and null predicates bind, and stop inspect metadata tables from
failing on geo bound types.

Nanosecond timestamps: add TimestampNanoLiteral / TimestamptzNanoLiteral
with string, long, date and micros conversions so row filters and partition
pruning work on timestamp_ns / timestamptz_ns columns; JSON single-value
(de)serialization for defaults; Avro timestamp-nanos parsing; nanosecond
precise column bounds from raw Parquet statistics; upcast of microsecond
files under nanosecond columns. Includes the identity partition path fix
from apache#3659 by @anxkhn.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Serialize v3 table metadata (next-row-id starts at 0 on create and on
upgrade), allow upgrading v1/v2 tables to v3, and add ManifestWriterV3 and
ManifestListWriterV3. The manifest list writer assigns first-row-id to data
manifests per the spec and the snapshot producer records first-row-id and
added-rows so next-row-id advances; Spark reads spec-correct _row_id values
from tables PyIceberg writes. Data files in v3 manifests inherit
first_row_id on read so copy-on-write rewrites keep existing row ids.
Manifest writers take a content argument so delete manifests can be written
in v2 and v3. Data files written by a failed operation are removed.

Schema gating: schemas with v3-only types are rejected on v1/v2 tables,
union_by_name uses the table's format version, and column defaults require
v3 unless incompatible changes are allowed.

Builds on apache#3551 (@rambleraptor), apache#3623 and apache#3624 (@moomindani), apache#3972
(@rambleraptor) and apache#4002.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ds in inspect

Row lineage: scans accept the reserved _row_id and
_last_updated_sequence_number columns in selected_fields. They are derived
from the data file's inherited first_row_id plus the row's position in the
file and from the entry's data sequence number, with values stored in the
file taking precedence; positions are tracked through positional deletes and
row filters, and files added before lineage was enabled read as null.

Deletion vectors: DVs are read from their content_offset /
content_size_in_bytes range with length, magic, CRC-32 and cardinality
checks, only deletion-vector-v1 blobs are decoded, and a bitmap count larger
than the payload is rejected. DVs are indexed by referenced_data_file, the
latest DV for a data file replaces older position deletes for it, and
delete-file identity includes the content range so several DVs in one
Puffin file are kept. Delete files are only applied to the data files they
were planned for.

Inspect: files, entries and delete_files gain first_row_id,
referenced_data_file, content_offset and content_size_in_bytes, matching
Spark's metadata tables.

Builds on apache#3478 (@KaiqiJinWow) and apache#3992 (@ghoshp83).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…nd v3 type rules

Geo columns can now be appended from binary or GeoArrow input. With geoarrow-pyarrow the
Parquet GEOMETRY/GEOGRAPHY logical types (CRS, edges) are written; geography with a
non-spherical algorithm, which the PyArrow writer rejects, falls back to plain binary with a
warning. Raw lexicographic WKB min/max are no longer recorded as bounds: geometry bounds are
built from the Parquet geospatial statistics (also for add_files) and encoded as the
little-endian x:y[:z][:m] points that Java's GeospatialBound reads. Geography bounds are only
kept when they do not wrap the antimeridian.

The Parquet write path now fills missing columns from write-default instead of
initial-default, and a required column with a write-default may be missing from the
dataframe. Sanitized write schemas keep their defaults. Avro schema defaults use the Avro
encoding of the type, and optional fields keep a null default as Avro requires.

Unknown columns must be optional with null defaults and are not stored in data files;
unknown, geometry and geography columns reject non-null defaults. date -> timestamp and
timestamp_ns are read promotions, and v3-only evolutions (unknown -> any, date ->
timestamp*) are rejected on v1/v2 tables. NestedField pickling keeps defaults.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Variant: tables with variant columns load and scan. Unshredded values are
returned as a struct of the metadata and value binaries, with a pure-Python
decoder (pyiceberg.variant) for the Variant binary encoding. Shredded
files raise when the column is projected; writing variant columns is
rejected until pyarrow can annotate the Parquet VARIANT logical type.
Variant cannot be a partition or sort source, has no bounds, and must
default to null.

Equality deletes: the delete file index tracks equality deletes by
partition and globally, and scans apply them after positional deletes with
null-equal semantics, reading only the delete columns by field id. The
planning and REST guards that rejected equality deletes are removed.

Multi-argument transforms: partition and sort fields carry source_ids,
fields with several sources load as unknown transforms (no pruning), and
serialization emits source-id or source-ids per the spec. Builds on apache#3630
(@moomindani).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Add a PuffinWriter and DeletionVector serialization (portable 64-bit
roaring bitmaps framed with the deletion-vector-v1 length, magic and
CRC-32), and a _RowDelta snapshot producer that commits delete manifests.
On format-version 3 tables with write.delete.mode=merge-on-read, delete()
marks the matching positions of partially matching data files with one
deletion vector per file instead of rewriting the files: positions are
unioned with the file's previous deletion vector or position deletes, the
replaced delete entries are removed in the same commit, and the snapshot
summary records added/removed DVs. Surviving rows keep their row ids and
Spark reads the resulting tables; a PyIceberg DV blob is byte-identical to
the blob Spark writes for the same positions. v1/v2 tables keep the
copy-on-write fallback.

Builds on apache#3474 (@moomindani) and apache#2822 (@rambleraptor).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

@jhamman jhamman left a comment

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pre-merge review of the v3 work. I read the full non-test diff, ran the touched unit-test modules (2580 passed), and exercised row lineage, deletion vectors, copy-on-write, defaults, and nanosecond timestamps against an in-memory catalog and against the Arraylake REST catalog (create, append, MOR delete, DV replacement, v2→v3 upgrade, DuckDB read all pass end to end).

Verdict: nothing here writes files Java/Spark would reject, and the append/DV/lineage-on-append paths look solid. Two cheap should-fixes before merge, one spec gap to either fix or document, plus nits. Inline comments below carry the details; the two that could not be anchored to diff lines are:

  1. create_table silently drops column defaults — pyiceberg/schema.py:1356-1364 (_SetFreshIDs.struct) rebuilds each NestedField without initial_default/write_default. Repro: create_table(Schema(NestedField(2, "c", StringType(), required=True, initial_default="init", write_default="wd")), properties={"format-version": "3"}) → table.schema().find_field("c") has no defaults, and appending without c fails instead of filling "wd". Only the add_column/set_default_value path works today (which is what the interop tests use). Carry the two fields through when rebuilding.

  2. set_default_value has no format-version guard — pyiceberg/table/update/schema.py:320. add_column rejects defaults below v3 (line ~266) but set_default_value("c", "y") on a v2 table commits {"write-default": "y"} into v2 metadata (reproduced). Java's Schema.checkCompatibility rejects that on AddSchema, so a Java REST server refuses the commit and a file/SQL catalog table becomes inconsistent. Apply the same guard as add_column.

Verified OK (no action): row-lineage arithmetic (first-row-id, added-rows, next-row-id advance, retry recomputation from refreshed metadata), Puffin DV framing/CRC/one-DV-per-file/removed-dvs, v2 position-delete merging, v1/v2 serialization untouched, geo type strings in Java form, v3-only types rejected on v1/v2 on both create and AddSchemaUpdate, multi-arg transforms round-trip as unknown.

[This is Claude Code on behalf of Joe Hamman]

@@ -736,7 +749,7 @@ def overwrite(
table_metadata=self.table_metadata, write_uuid=append_files.commit_uuid, df=df, io=self._table.io
)
for data_file in data_files:

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

should-fix (spec "must"): copy-on-write reassigns _row_id to surviving rows. overwrite(), upsert(), and the COW branch of delete() all rewrite files through _dataframe_to_data_files, which never carries _row_id / _last_updated_sequence_number into the new file. The v3 spec requires writers to preserve _row_id when copying existing rows. Repro: v3 table, append [1,2,3] (row ids 0..2), append [4,5] (3,4), overwrite([6], "id = 5") → row 4 comes back with _row_id 5 and next-row-id 7. Either preserve the ids on rewrite (read the lineage columns in the COW scan and write them out), or at minimum state in the support matrix that lineage is preserved only for append and merge-on-read delete and track this as a known gap.

Comment thread tests/integration/test_v3_interop.py Outdated
@@ -398,6 +434,9 @@ def _set_column_default_value(self, path: str | tuple[str, ...], default_value:

field = self._schema.find_field(name, self._case_sensitive)

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

should-fix: this path is reachable from set_default_value with no format-version check, so a default lands in v2 metadata. add_column guards this at ~L266; apply the same guard here (and in update_column if it can set defaults).

@@ -2288,7 +2432,7 @@ def from_rest_response(

def _rest_file_to_data_file(rest_file: RESTContentFile) -> DataFile:
"""Convert a REST content file to a manifest DataFile."""

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

should-fix: DataFile.from_args here omits first_row_id, referenced_data_file, content_offset, and content_size_in_bytes, all of which RESTDataFile / RESTPositionDeleteFile carry. With REST server-side planning, _row_id / _last_updated_sequence_number are silently null (_inherit_row_lineage needs first_row_id) and DV reads fall back to reading the whole Puffin file rather than the blob range. Pass them through.

self.sequence_number = sequence_number

@staticmethod
def from_rest_response(

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Related: from_rest_response passes no sequence_number to FileScanTask, so _last_updated_sequence_number cannot be inherited on REST-planned scans either.

Comment thread pyiceberg/io/pyarrow.py Outdated
Comment thread pyiceberg/table/metadata.py Outdated
Comment thread pyiceberg/transforms.py Outdated
Comment thread mkdocs/docs/format-version-3.md Outdated
Comment thread mkdocs/docs/format-version-3.md Outdated
jhamman and others added 2 commits September 27, 2026 08:11
…REST v3 fields

- create_table no longer drops initial-default / write-default when it
  reassigns field ids.
- set_default_value requires format version 3 unless incompatible changes
  are allowed, matching add_column.
- REST scan planning passes first_row_id and the deletion-vector range
  fields into FileScanTask, so row lineage and range-based DV reads work
  with server-side planning. The REST spec carries no data sequence
  number, so _last_updated_sequence_number stays null on REST-planned scans.
- Writes fill missing columns from write-default only, never initial-default.
- next-row-id is required in v3 metadata instead of defaulting to 0, which
  could reuse assigned row ids.
- Malformed single-argument transforms (bucket[abc]) raise again; only
  unknown names are tolerated as unknown transforms.
- Docs: nanosecond timestamp writes are supported.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ables

The copy-on-write branch of delete(), which overwrite(), upsert() and
dynamic partition overwrite also go through, now reads _row_id and
_last_updated_sequence_number with the data and writes them into the
rewritten files as physical columns with their reserved field ids, as the
v3 spec requires for copied rows. Readers already prefer stored values
over inherited ones, so surviving rows keep their ids across rewrites;
rows changed by upsert get fresh ids.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant