Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Sync (graph synchronization)

Sync makes a graph’s contents exactly the payload you supply, committing only the delta. It is the “full replacement” verb for data whose source of truth lives outside Fluree — an ontology maintained in an editor, a reference table regenerated by a pipeline — where you want the graph to match the latest export and the commit to show exactly what changed.

InsertUpsertSync
Asserts new triples✓✓✓
Retracts changed values—for supplied (subject, predicate) pairs✓
Retracts triples absent from the payload——✓
Scopepayloadpayload’s subjects/predicatesone whole graph

Semantics

Given the graph’s current contents A and the payload B:

  • retract A − B
  • assert B − A
  • A ∩ B produces no flakes — unchanged facts do not appear in the commit
  • A = B produces no commit (committed: false)

Sync is transactional and history-preserving: one normal commit at t = current + 1; queries as-of an earlier t still see the previous contents. SHACL validation, modify-policy enforcement, and novelty backpressure apply exactly as for any other transaction.

The scope is exactly one graph, named by the caller — never inferred from the payload: a named graph, or the ledger’s default graph. The payload may not address other named graphs, and reserved system graphs are rejected.

The Graph Store Protocol’s PUT is sync under the W3C protocol’s URL and status-code conventions.

HTTP endpoint

curl -X POST "http://localhost:8090/v1/fluree/sync?ledger=mydb:main&graph=http://example.org/graphs/ontology" \
  -H "Content-Type: application/json" \
  -d '{
    "@context": { "ex": "http://example.org/" },
    "@graph": [
      { "@id": "ex:alice", "ex:name": "Alice", "ex:role": "engineer" },
      { "@id": "ex:bob", "ex:name": "Bob" }
    ]
  }'

Query parameters:

ParameterMeaning
ledgerTarget ledger (name:branch)
graphTarget named graph IRI — the sync scope. Omit it to sync the default graph
defaultTarget the default graph explicitly (a bare key: ?default); the same as omitting graph. Passing both is a 400
dryRun=trueStage and report the delta (asserted/retracted counts) without committing — under the same policy, header, and inline-constraint inputs the real run uses, so it reports what the real run would do (or fails the way it would)
allowEmpty=trueConfirm an empty payload ("@graph": [], or an RDF document with no triples), which clears the graph

Payload formats

Content-TypePayload
application/jsonInsert-shaped JSON-LD
text/turtleTurtle: its triples are the graph’s contents
application/n-triplesN-Triples (read as Turtle, of which it is a subset)
application/trigTriG: the graph’s contents in GRAPH <graph> { … } (or <graph> { … }) blocks, or as default-graph triples

All four stage the same flakes for the same triples, so a graph loaded from one format and resynced from another commits nothing. RDF 1.2 annotations ({| … |}, ~ reifier, << s p o >>) sync in every format, and sync anchors each reifier bundle to the target graph, so a claims file syncs like any other data.

A TriG body is still one graph’s contents:

  • Every block must name the graph parameter’s IRI; a body syncing the default graph cannot contain blocks. A block for another graph is a 400. Several blocks for the target are fine; their triples are combined, and a blank-node label shared between them is one node.
  • Default-graph triples beside a block are a 400. In TriG they belong to the default graph, which sync does not write.
  • A GRAPH <#txn-meta> { … } block annotates the commit, as on /upsert.
  • Block contents get the full Turtle grammar, including [ … ] blank nodes and ( … ) collections.

Turtle and TriG bodies carry no opts, so policy inputs come from the fluree-identity / fluree-policy* headers, as on the other Turtle routes.

curl -X POST "http://localhost:8090/v1/fluree/sync?ledger=mydb:main&graph=http://example.org/graphs/ontology" \
  -H "Content-Type: text/turtle" \
  --data-binary @ontology.ttl

# No graph parameter: replace the default graph
curl -X POST "http://localhost:8090/v1/fluree/sync?ledger=mydb:main" \
  -H "Content-Type: text/turtle" \
  --data-binary @data.ttl

The payload is a set: a fact it states more than once (the JSON-LD parallel-annotation shape repeats the base edge once per annotation) is one assertion, so an unchanged payload is a no-op even when it carries several claims on one edge.

A dry run responds with the delta report:

{ "ledger": "mydb:main", "graph": "http://example.org/graphs/ontology",
  "asserted": 2, "retracted": 2, "committed": false, "dryRun": true, "t": 7 }

Rust API

#![allow(unused)]
fn main() {
use fluree_db_api::SyncGraphOpts;

let report = fluree
    .sync_named_graph("mydb:main", "http://example.org/graphs/ontology",
                      &payload, SyncGraphOpts::default())
    .await?;
assert!(report.committed || (report.asserted == 0 && report.retracted == 0));
}

SyncGraphOpts { dry_run, allow_empty } mirror the query parameters; sync_named_graph_with additionally takes explicit TxnOpts and a PolicyContext, and sync_named_graph_rdf_with is the same call for Turtle / N-Triples / TriG text. sync_graph_with(ledger, &GraphSel, GraphPayload, …) is the general form they delegate to: GraphSel::Default or GraphSel::Graph(iri), and GraphPayload::JsonLd(&json) or GraphPayload::Rdf(text). The report’s graph_iri is None for the default graph. The builder form fluree.stage(&handle).sync_graph(graph_iri, &payload) is also available, with sync_graph_payload(graph, payload, allow_empty) for any target and payload (note: the builder applies the allow_empty gate to an RDF payload, but not to JSON-LD). Its consensus terminal build_commit() returns Ok(None) for a no-change sync. Target validation (absolute IRI, no system graphs) is enforced at staging, so it applies on every entry point.

Safety rails

  • Empty payload requires opt-in. "@graph": [], or a Turtle / TriG document with no triples, means “the graph’s desired contents are empty” — i.e. clear the graph. Without allowEmpty, it is rejected, so a truncated export cannot silently wipe the graph.
  • The target is always one whole graph. Without graph, that is the default graph: a sync meant for a named graph that leaves out graph replaces the default graph’s contents instead. Subjects in the payload never widen or narrow the scope.
  • Dry run first. For a periodic pipeline, a dryRun call that reports an unexpectedly large retracted count is a cheap tripwire before the real run.
  • Policy model. Like CLEAR/COPY/MOVE (and unlike DELETE-WHERE), the current-contents scan is not view-policy filtered: sync is an authoritative whole-graph replacement, and a view-filtered scan would leave rows the caller cannot see in place, breaking “the graph now equals the payload”. Modify-policy is still enforced on the resulting delta.

Blank nodes

Sync skolemizes the payload’s blank nodes with a deterministic, graph-scoped key: the same blank-node label in the same target graph mints the same skolem IRI on every sync. Exporters that keep labels stable (hand-maintained Turtle, most pipeline generators) therefore resync bnode-rooted structures (OWL restrictions, RDF lists) with zero churn.

Exporters that regenerate labels on every save (e.g. Protégé’s genid…) still churn those structures: the triples are isomorphic but the labels — and therefore the skolemized identities — differ. The result is correct, just noisier commits. Structural (RDFC 1.0-style) canonicalization is the designed follow-up for label-unstable exporters.

Scale

Staging a sync materializes the target graph’s current flakes plus the payload’s flakes in memory — the same profile as CLEAR/COPY/MOVE (chunked staging for whole-graph operations is a known follow-up). A huge payload with a small delta is fine; the commit only carries the delta. A huge delta is bounded by novelty backpressure (reindex_max_bytes): the commit fails with NoveltyWouldExceed rather than overrunning memory, and the graph is left unchanged.

A huge graph is bounded by FLUREE_MAX_GRAPH_SCAN_FLAKES (default 10,000,000; 0 disables): the whole-graph scan stops there and the transaction fails with a resource-limit error naming the knob, instead of exhausting memory — an identical resync of a large graph is the worst case, since it materializes everything and commits nothing. A dry run trips the cap the same way the real run would, and CLEAR/DROP/COPY/MOVE share it. The follow-up that streams the diff instead of materializing the graph is #1691.