Sync (graph synchronization)
Sync makes a graph’s contents exactly the payload you supply, committing only the delta. It is the “full replacement” verb for data whose source of truth lives outside Fluree — an ontology maintained in an editor, a reference table regenerated by a pipeline — where you want the graph to match the latest export and the commit to show exactly what changed.
| Insert | Upsert | Sync | |
|---|---|---|---|
| Asserts new triples | ✓ | ✓ | ✓ |
| Retracts changed values | — | for supplied (subject, predicate) pairs | ✓ |
| Retracts triples absent from the payload | — | — | ✓ |
| Scope | payload | payload’s subjects/predicates | one whole graph |
Semantics
Given the graph’s current contents A and the payload B:
- retract
A − B - assert
B − A A ∩ Bproduces no flakes — unchanged facts do not appear in the commitA = Bproduces no commit (committed: false)
Sync is transactional and history-preserving: one normal commit at
t = current + 1; queries as-of an earlier t still see the previous
contents. SHACL validation, modify-policy enforcement, and novelty
backpressure apply exactly as for any other transaction.
The scope is exactly one graph, named by the caller — never inferred from the payload: a named graph, or the ledger’s default graph. The payload may not address other named graphs, and reserved system graphs are rejected.
The Graph Store Protocol’s PUT is sync under the
W3C protocol’s URL and status-code conventions.
HTTP endpoint
curl -X POST "http://localhost:8090/v1/fluree/sync?ledger=mydb:main&graph=http://example.org/graphs/ontology" \
-H "Content-Type: application/json" \
-d '{
"@context": { "ex": "http://example.org/" },
"@graph": [
{ "@id": "ex:alice", "ex:name": "Alice", "ex:role": "engineer" },
{ "@id": "ex:bob", "ex:name": "Bob" }
]
}'
Query parameters:
| Parameter | Meaning |
|---|---|
ledger | Target ledger (name:branch) |
graph | Target named graph IRI — the sync scope. Omit it to sync the default graph |
default | Target the default graph explicitly (a bare key: ?default); the same as omitting graph. Passing both is a 400 |
dryRun=true | Stage and report the delta (asserted/retracted counts) without committing — under the same policy, header, and inline-constraint inputs the real run uses, so it reports what the real run would do (or fails the way it would) |
allowEmpty=true | Confirm an empty payload ("@graph": [], or an RDF document with no triples), which clears the graph |
Payload formats
| Content-Type | Payload |
|---|---|
application/json | Insert-shaped JSON-LD |
text/turtle | Turtle: its triples are the graph’s contents |
application/n-triples | N-Triples (read as Turtle, of which it is a subset) |
application/trig | TriG: the graph’s contents in GRAPH <graph> { … } (or <graph> { … }) blocks, or as default-graph triples |
All four stage the same flakes for the same triples, so a graph loaded from
one format and resynced from another commits nothing. RDF 1.2 annotations
({| … |}, ~ reifier, << s p o >>) sync in every format, and sync
anchors each reifier bundle to the target graph, so a claims file syncs like
any other data.
A TriG body is still one graph’s contents:
- Every block must name the
graphparameter’s IRI; a body syncing the default graph cannot contain blocks. A block for another graph is a400. Several blocks for the target are fine; their triples are combined, and a blank-node label shared between them is one node. - Default-graph triples beside a block are a
400. In TriG they belong to the default graph, which sync does not write. - A
GRAPH <#txn-meta> { … }block annotates the commit, as on/upsert. - Block contents get the full Turtle grammar, including
[ … ]blank nodes and( … )collections.
Turtle and TriG bodies carry no opts, so policy inputs come from the
fluree-identity / fluree-policy* headers, as on the other Turtle routes.
curl -X POST "http://localhost:8090/v1/fluree/sync?ledger=mydb:main&graph=http://example.org/graphs/ontology" \
-H "Content-Type: text/turtle" \
--data-binary @ontology.ttl
# No graph parameter: replace the default graph
curl -X POST "http://localhost:8090/v1/fluree/sync?ledger=mydb:main" \
-H "Content-Type: text/turtle" \
--data-binary @data.ttl
The payload is a set: a fact it states more than once (the JSON-LD parallel-annotation shape repeats the base edge once per annotation) is one assertion, so an unchanged payload is a no-op even when it carries several claims on one edge.
A dry run responds with the delta report:
{ "ledger": "mydb:main", "graph": "http://example.org/graphs/ontology",
"asserted": 2, "retracted": 2, "committed": false, "dryRun": true, "t": 7 }
Rust API
#![allow(unused)]
fn main() {
use fluree_db_api::SyncGraphOpts;
let report = fluree
.sync_named_graph("mydb:main", "http://example.org/graphs/ontology",
&payload, SyncGraphOpts::default())
.await?;
assert!(report.committed || (report.asserted == 0 && report.retracted == 0));
}
SyncGraphOpts { dry_run, allow_empty } mirror the query parameters;
sync_named_graph_with additionally takes explicit TxnOpts and a
PolicyContext, and sync_named_graph_rdf_with is the same call for
Turtle / N-Triples / TriG text. sync_graph_with(ledger, &GraphSel, GraphPayload, …)
is the general form they delegate to: GraphSel::Default or
GraphSel::Graph(iri), and GraphPayload::JsonLd(&json) or
GraphPayload::Rdf(text). The report’s graph_iri is None for the default
graph. The builder form fluree.stage(&handle).sync_graph(graph_iri, &payload)
is also available, with sync_graph_payload(graph, payload, allow_empty) for
any target and payload (note: the builder applies the allow_empty gate to an
RDF payload, but not to JSON-LD). Its consensus
terminal build_commit() returns Ok(None) for a no-change sync. Target
validation (absolute IRI, no system graphs) is enforced at staging, so it
applies on every entry point.
Safety rails
- Empty payload requires opt-in.
"@graph": [], or a Turtle / TriG document with no triples, means “the graph’s desired contents are empty” — i.e. clear the graph. WithoutallowEmpty, it is rejected, so a truncated export cannot silently wipe the graph. - The target is always one whole graph. Without
graph, that is the default graph: a sync meant for a named graph that leaves outgraphreplaces the default graph’s contents instead. Subjects in the payload never widen or narrow the scope. - Dry run first. For a periodic pipeline, a
dryRuncall that reports an unexpectedly largeretractedcount is a cheap tripwire before the real run. - Policy model. Like
CLEAR/COPY/MOVE(and unlike DELETE-WHERE), the current-contents scan is not view-policy filtered: sync is an authoritative whole-graph replacement, and a view-filtered scan would leave rows the caller cannot see in place, breaking “the graph now equals the payload”. Modify-policy is still enforced on the resulting delta.
Blank nodes
Sync skolemizes the payload’s blank nodes with a deterministic, graph-scoped key: the same blank-node label in the same target graph mints the same skolem IRI on every sync. Exporters that keep labels stable (hand-maintained Turtle, most pipeline generators) therefore resync bnode-rooted structures (OWL restrictions, RDF lists) with zero churn.
Exporters that regenerate labels on every save (e.g. Protégé’s genid…)
still churn those structures: the triples are isomorphic but the labels — and
therefore the skolemized identities — differ. The result is correct, just
noisier commits. Structural (RDFC 1.0-style) canonicalization is the designed
follow-up for label-unstable exporters.
Scale
Staging a sync materializes the target graph’s current flakes plus the
payload’s flakes in memory — the same profile as CLEAR/COPY/MOVE
(chunked staging for whole-graph operations is a known follow-up). A huge
payload with a small delta is fine; the commit only carries the delta. A huge
delta is bounded by novelty backpressure (reindex_max_bytes): the commit
fails with NoveltyWouldExceed rather than overrunning memory, and the graph
is left unchanged.
A huge graph is bounded by FLUREE_MAX_GRAPH_SCAN_FLAKES (default
10,000,000; 0 disables): the whole-graph scan stops there and the
transaction fails with a resource-limit error naming the knob, instead of
exhausting memory — an identical resync of a large graph is the worst case,
since it materializes everything and commits nothing. A dry run trips the
cap the same way the real run would, and CLEAR/DROP/COPY/MOVE share
it. The follow-up that streams the diff instead of materializing the graph
is #1691.