Conversation
Provision v3 tables that Spark SQL cannot create (nanosecond timestamps, geometry/geography, column defaults, unknown) through the Iceberg Java Table API from dev/provision_v3.scala, run inside the Spark container by dev/provision.py. Add tests/integration/test_v3_interop.py comparing PyIceberg against the Java reference implementation; known v3 gaps are strict xfails so each fix flips a test. Add a docs page tracking v3 support and correct the geospatial and nanosecond documentation. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…edicates Geometry/geography: parse and serialize type strings in the Java/spec form (geometry(srid:3857), geography(srid:4326, vincenty), case-insensitive, quoted values still accepted for metadata written by earlier releases), validate the geography algorithm, resolve binary Parquet storage to the geo types on read (the Arrow dataset reader does not surface the GEOMETRY logical type), map geoarrow.wkb extension input, skip bounds for comparison predicates on geo columns in the metrics evaluators, add geo literals so equality and null predicates bind, and stop inspect metadata tables from failing on geo bound types. Nanosecond timestamps: add TimestampNanoLiteral / TimestamptzNanoLiteral with string, long, date and micros conversions so row filters and partition pruning work on timestamp_ns / timestamptz_ns columns; JSON single-value (de)serialization for defaults; Avro timestamp-nanos parsing; nanosecond precise column bounds from raw Parquet statistics; upcast of microsecond files under nanosecond columns. Includes the identity partition path fix from apache#3659 by @anxkhn. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Serialize v3 table metadata (next-row-id starts at 0 on create and on upgrade), allow upgrading v1/v2 tables to v3, and add ManifestWriterV3 and ManifestListWriterV3. The manifest list writer assigns first-row-id to data manifests per the spec and the snapshot producer records first-row-id and added-rows so next-row-id advances; Spark reads spec-correct _row_id values from tables PyIceberg writes. Data files in v3 manifests inherit first_row_id on read so copy-on-write rewrites keep existing row ids. Manifest writers take a content argument so delete manifests can be written in v2 and v3. Data files written by a failed operation are removed. Schema gating: schemas with v3-only types are rejected on v1/v2 tables, union_by_name uses the table's format version, and column defaults require v3 unless incompatible changes are allowed. Builds on apache#3551 (@rambleraptor), apache#3623 and apache#3624 (@moomindani), apache#3972 (@rambleraptor) and apache#4002. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ds in inspect Row lineage: scans accept the reserved _row_id and _last_updated_sequence_number columns in selected_fields. They are derived from the data file's inherited first_row_id plus the row's position in the file and from the entry's data sequence number, with values stored in the file taking precedence; positions are tracked through positional deletes and row filters, and files added before lineage was enabled read as null. Deletion vectors: DVs are read from their content_offset / content_size_in_bytes range with length, magic, CRC-32 and cardinality checks, only deletion-vector-v1 blobs are decoded, and a bitmap count larger than the payload is rejected. DVs are indexed by referenced_data_file, the latest DV for a data file replaces older position deletes for it, and delete-file identity includes the content range so several DVs in one Puffin file are kept. Delete files are only applied to the data files they were planned for. Inspect: files, entries and delete_files gain first_row_id, referenced_data_file, content_offset and content_size_in_bytes, matching Spark's metadata tables. Builds on apache#3478 (@KaiqiJinWow) and apache#3992 (@ghoshp83). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…nd v3 type rules Geo columns can now be appended from binary or GeoArrow input. With geoarrow-pyarrow the Parquet GEOMETRY/GEOGRAPHY logical types (CRS, edges) are written; geography with a non-spherical algorithm, which the PyArrow writer rejects, falls back to plain binary with a warning. Raw lexicographic WKB min/max are no longer recorded as bounds: geometry bounds are built from the Parquet geospatial statistics (also for add_files) and encoded as the little-endian x:y[:z][:m] points that Java's GeospatialBound reads. Geography bounds are only kept when they do not wrap the antimeridian. The Parquet write path now fills missing columns from write-default instead of initial-default, and a required column with a write-default may be missing from the dataframe. Sanitized write schemas keep their defaults. Avro schema defaults use the Avro encoding of the type, and optional fields keep a null default as Avro requires. Unknown columns must be optional with null defaults and are not stored in data files; unknown, geometry and geography columns reject non-null defaults. date -> timestamp and timestamp_ns are read promotions, and v3-only evolutions (unknown -> any, date -> timestamp*) are rejected on v1/v2 tables. NestedField pickling keeps defaults. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Variant: tables with variant columns load and scan. Unshredded values are returned as a struct of the metadata and value binaries, with a pure-Python decoder (pyiceberg.variant) for the Variant binary encoding. Shredded files raise when the column is projected; writing variant columns is rejected until pyarrow can annotate the Parquet VARIANT logical type. Variant cannot be a partition or sort source, has no bounds, and must default to null. Equality deletes: the delete file index tracks equality deletes by partition and globally, and scans apply them after positional deletes with null-equal semantics, reading only the delete columns by field id. The planning and REST guards that rejected equality deletes are removed. Multi-argument transforms: partition and sort fields carry source_ids, fields with several sources load as unknown transforms (no pruning), and serialization emits source-id or source-ids per the spec. Builds on apache#3630 (@moomindani). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Add a PuffinWriter and DeletionVector serialization (portable 64-bit roaring bitmaps framed with the deletion-vector-v1 length, magic and CRC-32), and a _RowDelta snapshot producer that commits delete manifests. On format-version 3 tables with write.delete.mode=merge-on-read, delete() marks the matching positions of partially matching data files with one deletion vector per file instead of rewriting the files: positions are unioned with the file's previous deletion vector or position deletes, the replaced delete entries are removed in the same commit, and the snapshot summary records added/removed DVs. Surviving rows keep their row ids and Spark reads the resulting tables; a PyIceberg DV blob is byte-identical to the blob Spark writes for the same positions. v1/v2 tables keep the copy-on-write fallback. Builds on apache#3474 (@moomindani) and apache#2822 (@rambleraptor). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
jhamman
left a comment
There was a problem hiding this comment.
Pre-merge review of the v3 work. I read the full non-test diff, ran the touched unit-test modules (2580 passed), and exercised row lineage, deletion vectors, copy-on-write, defaults, and nanosecond timestamps against an in-memory catalog and against the Arraylake REST catalog (create, append, MOR delete, DV replacement, v2→v3 upgrade, DuckDB read all pass end to end).
Verdict: nothing here writes files Java/Spark would reject, and the append/DV/lineage-on-append paths look solid. Two cheap should-fixes before merge, one spec gap to either fix or document, plus nits. Inline comments below carry the details; the two that could not be anchored to diff lines are:
-
create_tablesilently drops column defaults —pyiceberg/schema.py:1356-1364(_SetFreshIDs.struct) rebuilds eachNestedFieldwithoutinitial_default/write_default. Repro:create_table(Schema(NestedField(2, "c", StringType(), required=True, initial_default="init", write_default="wd")), properties={"format-version": "3"})→table.schema().find_field("c")has no defaults, and appending withoutcfails instead of filling"wd". Only theadd_column/set_default_valuepath works today (which is what the interop tests use). Carry the two fields through when rebuilding. -
set_default_valuehas no format-version guard —pyiceberg/table/update/schema.py:320.add_columnrejects defaults below v3 (line ~266) butset_default_value("c", "y")on a v2 table commits{"write-default": "y"}into v2 metadata (reproduced). Java'sSchema.checkCompatibilityrejects that onAddSchema, so a Java REST server refuses the commit and a file/SQL catalog table becomes inconsistent. Apply the same guard asadd_column.
Verified OK (no action): row-lineage arithmetic (first-row-id, added-rows, next-row-id advance, retry recomputation from refreshed metadata), Puffin DV framing/CRC/one-DV-per-file/removed-dvs, v2 position-delete merging, v1/v2 serialization untouched, geo type strings in Java form, v3-only types rejected on v1/v2 on both create and AddSchemaUpdate, multi-arg transforms round-trip as unknown.
[This is Claude Code on behalf of Joe Hamman]
| @@ -736,7 +749,7 @@ def overwrite( | |||
| table_metadata=self.table_metadata, write_uuid=append_files.commit_uuid, df=df, io=self._table.io | |||
| ) | |||
| for data_file in data_files: | |||
There was a problem hiding this comment.
should-fix (spec "must"): copy-on-write reassigns _row_id to surviving rows. overwrite(), upsert(), and the COW branch of delete() all rewrite files through _dataframe_to_data_files, which never carries _row_id / _last_updated_sequence_number into the new file. The v3 spec requires writers to preserve _row_id when copying existing rows. Repro: v3 table, append [1,2,3] (row ids 0..2), append [4,5] (3,4), overwrite([6], "id = 5") → row 4 comes back with _row_id 5 and next-row-id 7. Either preserve the ids on rewrite (read the lineage columns in the COW scan and write them out), or at minimum state in the support matrix that lineage is preserved only for append and merge-on-read delete and track this as a known gap.
| @@ -398,6 +434,9 @@ def _set_column_default_value(self, path: str | tuple[str, ...], default_value: | |||
|
|
|||
| field = self._schema.find_field(name, self._case_sensitive) | |||
|
|
|||
There was a problem hiding this comment.
should-fix: this path is reachable from set_default_value with no format-version check, so a default lands in v2 metadata. add_column guards this at ~L266; apply the same guard here (and in update_column if it can set defaults).
| @@ -2288,7 +2432,7 @@ def from_rest_response( | |||
|
|
|||
| def _rest_file_to_data_file(rest_file: RESTContentFile) -> DataFile: | |||
| """Convert a REST content file to a manifest DataFile.""" | |||
There was a problem hiding this comment.
should-fix: DataFile.from_args here omits first_row_id, referenced_data_file, content_offset, and content_size_in_bytes, all of which RESTDataFile / RESTPositionDeleteFile carry. With REST server-side planning, _row_id / _last_updated_sequence_number are silently null (_inherit_row_lineage needs first_row_id) and DV reads fall back to reading the whole Puffin file rather than the blob range. Pass them through.
| self.sequence_number = sequence_number | ||
|
|
||
| @staticmethod | ||
| def from_rest_response( |
There was a problem hiding this comment.
Related: from_rest_response passes no sequence_number to FileScanTask, so _last_updated_sequence_number cannot be inherited on REST-planned scans either.
…REST v3 fields - create_table no longer drops initial-default / write-default when it reassigns field ids. - set_default_value requires format version 3 unless incompatible changes are allowed, matching add_column. - REST scan planning passes first_row_id and the deletion-vector range fields into FileScanTask, so row lineage and range-based DV reads work with server-side planning. The REST spec carries no data sequence number, so _last_updated_sequence_number stays null on REST-planned scans. - Writes fill missing columns from write-default only, never initial-default. - next-row-id is required in v3 metadata instead of defaulting to 0, which could reuse assigned row ids. - Malformed single-argument transforms (bucket[abc]) raise again; only unknown names are tolerated as unknown transforms. - Docs: nanosecond timestamp writes are supported. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ables The copy-on-write branch of delete(), which overwrite(), upsert() and dynamic partition overwrite also go through, now reads _row_id and _last_updated_sequence_number with the data and writes them into the rewritten files as physical columns with their reserved field ids, as the v3 spec requires for copied rows. Readers already prefer stored values over inherited ones, so surviving rows keep their ids across rewrites; rows changed by upsert get fresh ids. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Part of apache#1818 (V3 tracking issue) and apache#1551 (writing V3 tables).
Rationale for this change
PyIceberg could parse Iceberg format-version 3 metadata but was not a v3 writer, and several v3 reads were broken against tables produced by the Java reference implementation. A gap analysis against Spark 4.0 + Iceberg 1.11 (Java Table API and the compose stack's REST fixture) found, among others:
geometry(srid:3857)type strings written by Java could not be parsed (and PyIceberg's own quoted form was rejected or stored with embedded quotes by Java), geo columns could not be scanned,varianttables could not be loaded, predicates ontimestamp_nscolumns raised,_row_id/_last_updated_sequence_numberwere not readable,inspectomitted every v3 field, and all v3 writes failed behind three separate guards while leaving orphaned data files.This branch closes those gaps in stages, one commit each, reusing and crediting the open community PRs where they existed (apache#3551, apache#3623, apache#3624, apache#3972, apache#3474/apache#2822, apache#3478, apache#3992, apache#3630, apache#3659, apache#4002):
dev/provision_v3.scalaprovisions the v3 tables Spark SQL cannot create (nanosecond timestamps, geometry, column defaults, unknown, equality deletes) through the Java Table API;tests/integration/test_v3_interop.pycompares PyIceberg against Spark/Java; newformat-version-3.mddocs page with a support matrix.TimestampNanoLiteral/TimestamptzNanoLiteralwith string/long/date conversions so filters and partition pruning work on ns columns; ns JSON, Avrotimestamp-nanos, nanosecond-exact bounds.ManifestWriterV3/ManifestListWriterV3with first-row-id assignment,first-row-id/added-rows/next-row-idon commit, delete manifests, orphan cleanup on failed writes, format-version gating of v3-only schema changes._row_id/_last_updated_sequence_numberas reserved metadata columns (position-tracked through deletes and filters),first_row_idinheritance, range-based DV reads with length/magic/CRC/cardinality validation, DVs keyed byreferenced_data_filewith supersession of older position deletes, v3 columns ininspecttables.GeospatialBoundencoding from Parquet geospatial statistics,write-defaulton the Parquet path, unknown columns dropped from data files,date -> timestamppromotion gated to v3.VariantTypewith unshredded reads and a pure-Python decoder (pyiceberg.variant), equality-delete reads with null-equal semantics, multi-arg transforms loaded as unknown transforms with specsource-idsserialization.PuffinWriter, DV serialization, and a_RowDeltaproducer sodelete()on v3 merge-on-read tables writes one DV per data file, replacing the file's previous DV; a PyIceberg blob is byte-identical to Spark's for the same positions.create_table,set_default_valueis gated to v3, REST scan planning carries the v3 file fields,write-defaultonly on writes,next-row-idrequired, strict parsing of malformed transforms, and copy-on-write rewrites preserve_row_id/_last_updated_sequence_numberof copied rows.Known limits (documented in
format-version-3.md): writing variant and reading shredded variant wait on pyarrow's Parquet VARIANT support;overwrite()/upsert()remain copy-on-write; spatial predicates, geo file skipping and table encryption are out of scope. Spark 4.0 + Iceberg 1.11 cannot readtimestamp_nsor geometry columns and Java's generic Parquet writer cannot write geometry, so those features are verified through the Java Table API, Parquet schema inspection and PyIceberg round trips rather than Spark reads.Are these changes tested?
Yes.
dev/docker-compose-integration.yml(REST fixture, Hive, Spark Connect): full run 1401 passed, 9 skipped before the review fixes; the affected suites (interop, writes, partitioned writes, deletes, REST schema/scan planning) re-run green after them. Format-version parametrizations in the write, partitioned write, add_files and delete suites now cover[1, 2, 3].tests/integration/test_v3_interop.py: Spark reads PyIceberg-written v3 tables (appends, copy-on-write and merge-on-read deletes, overwrites, upgrades, defaults, unknown) with matching_row_idvalues and DV metadata; PyIceberg reads Spark/Java-written DV, variant, equality-delete, nanosecond, default, unknown and geometry tables with row-for-row parity.Are there any user-facing changes?
upgrade_table_version(3)) and written on every catalog._row_idand_last_updated_sequence_numberviaselected_fields.delete()on v3 tables withwrite.delete.mode=merge-on-readwrites deletion vectors instead of rewriting files; v1/v2 tables keep the copy-on-write fallback (warning reworded).geometry/geographytype strings now serialize in the spec form (geometry(srid:3857)); the quoted form written by earlier releases is still parsed.VariantType(read-only) and literalsTimestampNanoLiteral,TimestamptzNanoLiteral,GeometryLiteral,GeographyLiteral; new modulespyiceberg.variant,pyiceberg.table.metadata_columns,pyiceberg.utils.geo.set_default_value, and v3-only types or promotions on v1/v2 tables now raise unlessallow_incompatible_changesis set;next-row-idis required in v3 metadata; malformed transform names such asbucket[abc]raise again.inspecttables gainfirst_row_id,referenced_data_file,content_offsetandcontent_size_in_bytes.🤖 Generated with Claude Code