Conversation
The policy plug-in now runs on QLever (or any SPARQL 1.1 endpoint) instead of an in-memory graph, which refused anything above 5M triples. - Selector parameters, including lists, are injected as an inline VALUES block opening the outer WHERE group. A trailing VALUES fails open. - vcfp:LinkedSelector (shipped in the VCF Core profile) selects records by what their calls link to. Its violations query refuses a graph without the links (FILTER NOT EXISTS: rdflib misreads OPTIONAL + !BOUND here). - evaluate --endpoint streams the view byte for byte, in input order. Inputs are filtered in parallel, one gzip member per input. Each line costs a few set lookups (the binding rules' selections merged once), not a scan of every rule. - check --endpoint needs --view-endpoint serving the view alone. It confirms the served triple count against the view's lines, then finds dangling references with one FILTER NOT EXISTS query. --oracle-endpoint serves the new `oracle` command's VCF-text graph. - decide stays rule by rule, so the in-memory view and the oracle remain an independent implementation. It now finds a term's ancestors once. On arm 2 (104 genomes, 70M triples) a requester governs in ~1-1.5 min (evaluate) plus ~3-5 min (check), from ~23-27 plus ~9-22 min, with byte-identical views. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ARQL Linking now reaches real resources, and files join on shared identifiers. - AlleleJoin with a SequenceMap reference: the shipped `spdi` linker emits trimmed NCBI SPDI IRIs, the same for a variant whatever each file calls its chromosome. - vcfl:contigAliases lets an IntervalJoin resolve contig names through a SequenceMap. `ensembl-genes-grch38` links calls to the Ensembl 116 genes they overlap (GFF3 fetched once by digest). - `rsid-myvariant`: a tier-3 resolver confirming rsIDs against MyVariant.info in batches of 1,000 (Ensembl's REST service returned persistent HTTP 500s). vcfl:requestTimeout sets a service's read timeout. - Linking from RDF reads the graph with SPARQL through any store: `vcf-rdfizer-link run --endpoint`, or a file loaded into a temporary on-disk Oxigraph store of just the join fields. The reader's assumptions are declared queries. pyoxigraph (>= 0.3.18) becomes a dependency. - read_tsv removed: it had no caller. - test_linking_edges_unit.py pins every refusal and edge path. Linking line and branch coverage is 99%, resolvers included. test_generality_unit.py keeps experiment names out of shipped code and profiles. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
`oracle -o x.nt.gz` writes the VCF-text oracle compressed. A whole genome's oracle is ~31M triples, several gigabytes as plain N-Triples. Endpoints and the streaming executor already read .nt.gz. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
No SPARQL index can hold an empty file, and an empty view names nothing, so nothing in it can dangle. check_stream now takes view_store=None for an empty view, and refuses None for any other. check --view-endpoint is optional accordingly. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A minor: linking reaches real resources and the policy plug-in reaches cohort
and whole-genome scale. Nothing changes what a v3.2.0 conversion writes.
- Linkers:
- `spdi` (AlleleJoin over a SequenceMap): the same NCBI SPDI IRI for a
variant in every file, whatever each calls its chromosome.
- `ensembl-genes-grch38`: Ensembl 116 gene overlaps, through the new
vcfl:contigAliases.
- `rsid-myvariant`: rsIDs confirmed against MyVariant.info.
- vcfl:requestTimeout sets a live service's read timeout.
- vcf-rdfizer-policy 0.2.0:
- evaluate and check against a SPARQL endpoint (QLever), streaming the view
in parallel and checking it against a view endpoint and the VCF-text
oracle (`oracle`, gzip-capable);
- list-valued selector parameters as inline VALUES;
- vcfp:LinkedSelector (records selected by what their calls link to).
- Linking from RDF reads the graph with SPARQL: from an endpoint
(`vcf-rdfizer-link run --endpoint`) or a temporary Oxigraph store of just
the join fields. pyoxigraph (>= 0.3.18) is a new dependency.
- Validation (#29): the shape layer is gated on the decoded graph's size
(`--shacl-max-triples`, default 50M, recorded as a skip), and
`--node-heap-mb` raises Node's heap for the Comunica-backed engines.
- Evaluated in vcf-rdfizer-testing experiment 17: five heterogeneous genomes,
104 1000 Genomes participants, and HG005's whole genome (668M triples). Each
agreed with a bcftools baseline for every requester.
The conda sha256 stays a placeholder until the tag exists. Populate it with
`python3 scripts/release.py 3.3.0 --fetch-conda-sha256` after pushing v3.3.0.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
CI's coverage job failed on test_linking_edges_unit.RdfInputs.test_refusals:
RuntimeError: Invalid argument: ingestion arg list is empty
inputs.py in read_rdf -> store.bulk_extend(wanted(triples))
The test feeds a triple carrying no VCF vocabulary, so wanted() filters
everything out and bulk_extend receives an empty batch. An input like that is
legitimate -- it is a graph that simply does not use the VCF-RDFizer
vocabulary -- and the caller is owed read_store's "No VCFFile/hasRecord"
ValueError rather than a RuntimeError from the Rust loader. read_rdf now pulls
one triple first and loads only when there is something to load, translating a
lazy-parse SyntaxError at either point.
Two things hid this. RdfInputs is the only linking test class without the
rdflib skip guard, so it runs where the others skip; and the behaviour is
platform-dependent -- pyoxigraph 0.5.11 rejects the empty batch on Linux and
accepts it on macOS, so the whole file passes locally. Only the coverage job,
which pulls in rdflib through its extra installs, met both conditions.
test_a_graph_with_no_vcf_triples_never_reaches_the_bulk_loader therefore spies
on bulk_extend and asserts it is never handed an empty batch, which fails on
macOS against the unfixed code. Without that the regression stays invisible to
anyone developing on a Mac, which is how it got here.
test_the_first_kept_triple_is_not_dropped guards the obvious way to get the fix
wrong -- loading the remainder and losing the triple that was peeked at.
Verified by reintroducing each failure mode: dropping the peeked triple errors,
simulating the Linux empty-batch RuntimeError reproduces the CI failure
exactly, and reverting the guard fails the new test.
Full suite: 1062 tests, OK.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
… temp file With the test error fixed, CI's coverage step reached its last command and failed there instead: No source for code: '/tmp/.../resolver.py' ##[error]Process completed with exit code 1 The repository had no coverage configuration, so coverage measured whatever the interpreter executed -- including a resolver.py that the linking tests write into a TemporaryDirectory, import, and let the cleanup delete. By the time coverage xml runs, the file is gone. This was unreachable before: the step runs under bash -eo pipefail, so the earlier test failure aborted at the first coverage run and coverage xml never executed. Scoping to the repository is the honest fix rather than ignore_errors. A throwaway script in a temp directory is not this project's code and was never meant to be in the report, and silencing the error would also hide a genuinely missing project source later. Like the bug it was hiding behind, this is platform-split: the same coverage 7.16.2 exits 1 on the Linux runner and only warns on macOS, so it cannot be reproduced locally. The checkable outcomes are that the "No source" warning goes from one to none and the XML report holds 73 files, all project code and no temp paths. CI's exact three commands, run locally: 1062 tests OK, XML written, all exit 0. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
Codecov Report❌ Patch coverage is 📢 Thoughts on this report? Let us know! |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
v3.3.0 makes linking reach real resources, and makes the policy plug-in work at cohort and whole-genome scale. Nothing changes what a v3.2.0 conversion writes. Release metadata comes from
python3 scripts/release.py 3.3.0, followingscripts/RELEASING.md.Commits
vcf-rdfizer-policy0.2.0:evaluateandcheckread the graph from a SPARQL endpoint (QLever) instead of loading it into memory.vcfp:LinkedSelector, and list-valued selector parameters as inlineVALUES.spdigives a variant the same NCBI SPDI IRI in every file, whatever each file calls its chromosome.ensembl-genes-grch38links calls to the Ensembl 116 genes they overlap, using the newvcfl:contigAliases.rsid-myvariantconfirms rsIDs against MyVariant.info: batch POST, 1 request/s, capped per run.vcf-rdfizer-link run --endpointreads the input from an endpoint. Otherwise only the join fields are loaded into a temporary Oxigraph store.pyoxigraph >= 0.3.18is a new dependency.read_tsvwas dead code and is removed.--shacl-max-triplesand--node-heap-mbindocs/cli-reference.md.What v3.3.0 ships (since v3.2.0)
--node-heap-mbraises Node's heap for the Comunica-backed engines.Verified
Unit suite: 1,007 tests passed on
vcf-bench-2at the pre-rebase tip (5d3adb1). Rebasing onto Gate the shape layer on the graph, and let Node's heap be raised #29 touched no file this branch changes. CI on this PR covers the rebased tip.Coverage: the linking package has 102 tests, at 99% line coverage.
release.py --check-tag v3.3.0: release metadata matches.End-to-end: evaluated in vcf-rdfizer-testing experiment 17 at three scales:
Every requester's carrier list equals a bcftools baseline's at all three scales. MyVariant was called 21 times in total, and every call returned HTTP 200.
After merge
git tag -a v3.3.0 -m "VCF-RDFizer v3.3.0"on the merge commit, then push the tag. This triggers PyPI and Docker Hub (3.3.0,v3.3.0,latest).python3 scripts/release.py 3.3.0 --fetch-conda-sha256and open a follow-up PR.pyoxigraphtorun.🤖 Generated with Claude Code