Replies: 1 comment
|
This behavior is expected, although the current wording around "handle" is easy to read differently.
There is no supported way for several worker contexts to read another context's uncommitted changes as one shared transaction. Work that requires that state must run through the parent connection and will be serialized. The practical options are:
For a concurrent read phase followed by a write-back phase, I would commit setup, run workers with thread-local cursors, collect their results, then start a new transaction on the parent for the write-back. If one atomic transaction across all three phases is mandatory, concurrent DuckDB client contexts are not the right boundary. The threading guide's phrase "thread-local connection" matches the observed semantics: DuckDB multiple Python threads. |
Uh oh!
There was an error while loading. Please reload this page.
Context:
I'm currently using python + multiple threads to concurrently process many isolated read-only queries.
The parent then gathers responses and writes back to duckdb. All of this happens in the context of a single transaction.
Each worker thread uses a db.cursor() as explained in the documentation.
A new requirement has been introduced that makes the main process write to duckdb before the workers start, but after the transaction has started.
Issue:
I'm observing the unexpected (for me) behavior that the worker threads have no visibility over the uncommitted data introduced before .cursor() was invoked.
Here's a short demo without all the threading abstracted:
I'm unsure if this is expected behavior.
Either the cursor shares the full context to the db connection, in which case it should be able to read the uncommited data.
Or the cursor does not share is state with the parent, and behaves more as a separate connection. If this is the case then the documentation could be more explicit about this.
All reactions