Skip to content

Release v3.3.0: real-resource linkers, and policies at cohort and whole-genome scale - #30

Open
ecrum19 wants to merge 7 commits into
mainfrom
release/v3.3.0
Open

ecrum19 wants to merge 7 commits into
mainfrom
release/v3.3.0

Conversation

@ecrum19

@ecrum19 ecrum19 commented Oct 1, 2026

Copy link
Copy Markdown
Owner

v3.3.0 makes linking reach real resources, and makes the policy plug-in work at cohort and whole-genome scale. Nothing changes what a v3.2.0 conversion writes. Release metadata comes from python3 scripts/release.py 3.3.0, following scripts/RELEASING.md.

Commits

  1. Evaluate and check policies against SPARQL endpoints, at cohort scale. vcf-rdfizer-policy 0.2.0:
    • evaluate and check read the graph from a SPARQL endpoint (QLever) instead of loading it into memory.
    • The release view is streamed in parallel, one gzip member per worker.
    • Each record's decision is computed once, in place of once per rule.
    • The check runs against a view endpoint and the VCF-text oracle.
    • Also new: vcfp:LinkedSelector, and list-valued selector parameters as inline VALUES.
  2. Add the SPDI, Ensembl-gene and MyVariant linkers, and read RDF inputs with SPARQL.
    • spdi gives a variant the same NCBI SPDI IRI in every file, whatever each file calls its chromosome.
    • ensembl-genes-grch38 links calls to the Ensembl 116 genes they overlap, using the new vcfl:contigAliases.
    • rsid-myvariant confirms rsIDs against MyVariant.info: batch POST, 1 request/s, capped per run.
    • vcf-rdfizer-link run --endpoint reads the input from an endpoint. Otherwise only the join fields are loaded into a temporary Oxigraph store. pyoxigraph >= 0.3.18 is a new dependency.
    • read_tsv was dead code and is removed.
  3. Let the oracle command write gzip.
  4. Check an empty view without a view endpoint.
  5. Release v3.3.0. Updates the version markers and the policy, linking, limitations and roadmap docs. It also documents Gate the shape layer on the graph, and let Node's heap be raised #29's --shacl-max-triples and --node-heap-mb in docs/cli-reference.md.

What v3.3.0 ships (since v3.2.0)

Verified

  • Unit suite: 1,007 tests passed on vcf-bench-2 at the pre-rebase tip (5d3adb1). Rebasing onto Gate the shape layer on the graph, and let Node's heap be raised #29 touched no file this branch changes. CI on this PR covers the rebased tip.

  • Coverage: the linking package has 102 tests, at 99% line coverage.

  • release.py --check-tag v3.3.0: release metadata matches.

  • End-to-end: evaluated in vcf-rdfizer-testing experiment 17 at three scales:

    • five heterogeneous genomes plus ClinVar;
    • 104 participants from 1000 Genomes;
    • HG005's whole genome, 668M triples.

    Every requester's carrier list equals a bcftools baseline's at all three scales. MyVariant was called 21 times in total, and every call returned HTTP 200.

After merge

  1. Run git tag -a v3.3.0 -m "VCF-RDFizer v3.3.0" on the merge commit, then push the tag. This triggers PyPI and Docker Hub (3.3.0, v3.3.0, latest).
  2. Run python3 scripts/release.py 3.3.0 --fetch-conda-sha256 and open a follow-up PR.
  3. Open the feedstock PR. It adds pyoxigraph to run.
  4. Record the published image digest.

🤖 Generated with Claude Code

ecrum19 and others added 5 commits October 1, 2026 08:11
The policy plug-in now runs on QLever (or any SPARQL 1.1 endpoint) instead
of an in-memory graph, which refused anything above 5M triples.

- Selector parameters, including lists, are injected as an inline VALUES
  block opening the outer WHERE group. A trailing VALUES fails open.
- vcfp:LinkedSelector (shipped in the VCF Core profile) selects records by
  what their calls link to. Its violations query refuses a graph without the
  links (FILTER NOT EXISTS: rdflib misreads OPTIONAL + !BOUND here).
- evaluate --endpoint streams the view byte for byte, in input order. Inputs
  are filtered in parallel, one gzip member per input. Each line costs a few
  set lookups (the binding rules' selections merged once), not a scan of
  every rule.
- check --endpoint needs --view-endpoint serving the view alone. It confirms
  the served triple count against the view's lines, then finds dangling
  references with one FILTER NOT EXISTS query. --oracle-endpoint serves the
  new `oracle` command's VCF-text graph.
- decide stays rule by rule, so the in-memory view and the oracle remain an
  independent implementation. It now finds a term's ancestors once.

On arm 2 (104 genomes, 70M triples) a requester governs in ~1-1.5 min
(evaluate) plus ~3-5 min (check), from ~23-27 plus ~9-22 min, with
byte-identical views.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ARQL

Linking now reaches real resources, and files join on shared identifiers.

- AlleleJoin with a SequenceMap reference: the shipped `spdi` linker emits
  trimmed NCBI SPDI IRIs, the same for a variant whatever each file calls
  its chromosome.
- vcfl:contigAliases lets an IntervalJoin resolve contig names through a
  SequenceMap. `ensembl-genes-grch38` links calls to the Ensembl 116 genes
  they overlap (GFF3 fetched once by digest).
- `rsid-myvariant`: a tier-3 resolver confirming rsIDs against MyVariant.info
  in batches of 1,000 (Ensembl's REST service returned persistent HTTP 500s).
  vcfl:requestTimeout sets a service's read timeout.
- Linking from RDF reads the graph with SPARQL through any store:
  `vcf-rdfizer-link run --endpoint`, or a file loaded into a temporary
  on-disk Oxigraph store of just the join fields. The reader's assumptions
  are declared queries. pyoxigraph (>= 0.3.18) becomes a dependency.
- read_tsv removed: it had no caller.
- test_linking_edges_unit.py pins every refusal and edge path. Linking line
  and branch coverage is 99%, resolvers included. test_generality_unit.py
  keeps experiment names out of shipped code and profiles.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
`oracle -o x.nt.gz` writes the VCF-text oracle compressed. A whole genome's
oracle is ~31M triples, several gigabytes as plain N-Triples. Endpoints and
the streaming executor already read .nt.gz.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
No SPARQL index can hold an empty file, and an empty view names nothing, so
nothing in it can dangle. check_stream now takes view_store=None for an empty
view, and refuses None for any other. check --view-endpoint is optional
accordingly.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A minor: linking reaches real resources and the policy plug-in reaches cohort
and whole-genome scale. Nothing changes what a v3.2.0 conversion writes.

- Linkers:
  - `spdi` (AlleleJoin over a SequenceMap): the same NCBI SPDI IRI for a
    variant in every file, whatever each calls its chromosome.
  - `ensembl-genes-grch38`: Ensembl 116 gene overlaps, through the new
    vcfl:contigAliases.
  - `rsid-myvariant`: rsIDs confirmed against MyVariant.info.
  - vcfl:requestTimeout sets a live service's read timeout.
- vcf-rdfizer-policy 0.2.0:
  - evaluate and check against a SPARQL endpoint (QLever), streaming the view
    in parallel and checking it against a view endpoint and the VCF-text
    oracle (`oracle`, gzip-capable);
  - list-valued selector parameters as inline VALUES;
  - vcfp:LinkedSelector (records selected by what their calls link to).
- Linking from RDF reads the graph with SPARQL: from an endpoint
  (`vcf-rdfizer-link run --endpoint`) or a temporary Oxigraph store of just
  the join fields. pyoxigraph (>= 0.3.18) is a new dependency.
- Validation (#29): the shape layer is gated on the decoded graph's size
  (`--shacl-max-triples`, default 50M, recorded as a skip), and
  `--node-heap-mb` raises Node's heap for the Comunica-backed engines.
- Evaluated in vcf-rdfizer-testing experiment 17: five heterogeneous genomes,
  104 1000 Genomes participants, and HG005's whole genome (668M triples). Each
  agreed with a bcftools baseline for every requester.

The conda sha256 stays a placeholder until the tag exists. Populate it with
`python3 scripts/release.py 3.3.0 --fetch-conda-sha256` after pushing v3.3.0.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
ecrum19 and others added 2 commits October 1, 2026 08:37
CI's coverage job failed on test_linking_edges_unit.RdfInputs.test_refusals:

  RuntimeError: Invalid argument: ingestion arg list is empty
    inputs.py in read_rdf -> store.bulk_extend(wanted(triples))

The test feeds a triple carrying no VCF vocabulary, so wanted() filters
everything out and bulk_extend receives an empty batch. An input like that is
legitimate -- it is a graph that simply does not use the VCF-RDFizer
vocabulary -- and the caller is owed read_store's "No VCFFile/hasRecord"
ValueError rather than a RuntimeError from the Rust loader. read_rdf now pulls
one triple first and loads only when there is something to load, translating a
lazy-parse SyntaxError at either point.

Two things hid this. RdfInputs is the only linking test class without the
rdflib skip guard, so it runs where the others skip; and the behaviour is
platform-dependent -- pyoxigraph 0.5.11 rejects the empty batch on Linux and
accepts it on macOS, so the whole file passes locally. Only the coverage job,
which pulls in rdflib through its extra installs, met both conditions.

test_a_graph_with_no_vcf_triples_never_reaches_the_bulk_loader therefore spies
on bulk_extend and asserts it is never handed an empty batch, which fails on
macOS against the unfixed code. Without that the regression stays invisible to
anyone developing on a Mac, which is how it got here.
test_the_first_kept_triple_is_not_dropped guards the obvious way to get the fix
wrong -- loading the remainder and losing the triple that was peeked at.

Verified by reintroducing each failure mode: dropping the peeked triple errors,
simulating the Linux empty-batch RuntimeError reproduces the CI failure
exactly, and reverting the guard fails the new test.

Full suite: 1062 tests, OK.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
… temp file

With the test error fixed, CI's coverage step reached its last command and
failed there instead:

  No source for code: '/tmp/.../resolver.py'
  ##[error]Process completed with exit code 1

The repository had no coverage configuration, so coverage measured whatever the
interpreter executed -- including a resolver.py that the linking tests write
into a TemporaryDirectory, import, and let the cleanup delete. By the time
coverage xml runs, the file is gone. This was unreachable before: the step runs
under bash -eo pipefail, so the earlier test failure aborted at the first
coverage run and coverage xml never executed.

Scoping to the repository is the honest fix rather than ignore_errors. A
throwaway script in a temp directory is not this project's code and was never
meant to be in the report, and silencing the error would also hide a genuinely
missing project source later.

Like the bug it was hiding behind, this is platform-split: the same coverage
7.16.2 exits 1 on the Linux runner and only warns on macOS, so it cannot be
reproduced locally. The checkable outcomes are that the "No source" warning
goes from one to none and the XML report holds 73 files, all project code and
no temp paths.

CI's exact three commands, run locally: 1062 tests OK, XML written, all exit 0.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@codecov-commenter

Copy link
Copy Markdown

⚠️ Please install the 'codecov app svg image' to ensure uploads and comments are reliably processed by Codecov.

Codecov Report

❌ Patch coverage is 83.23144% with 192 lines in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
vcf_rdfizer_policies/check.py 0.00% 74 Missing ⚠️
vcf_rdfizer_policies/release.py 0.00% 30 Missing ⚠️
vcf_rdfizer_policies/engine.py 73.40% 25 Missing ⚠️
vcf_rdfizer_policies/store.py 51.06% 23 Missing ⚠️
vcf_rdfizer_policy.py 47.22% 19 Missing ⚠️
vcf_rdfizer_policies/vcf_oracle.py 77.77% 8 Missing ⚠️
vcf_rdfizer_policies/policy.py 16.66% 5 Missing ⚠️
test/test_linking_edges_unit.py 98.90% 4 Missing ⚠️
vcf_rdfizer_linking/inputs.py 96.29% 2 Missing ⚠️
test/test_generality_unit.py 92.85% 1 Missing ⚠️
... and 1 more

📢 Thoughts on this report? Let us know!

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants