fluree sweep
Reclaim index artifacts that no index chain references.
Usage
fluree sweep [LEDGER] [--dry-run] [--remote <NAME>]
Arguments
| Argument | Description |
|---|---|
[LEDGER] | Ledger name, without a branch suffix (defaults to active ledger) |
Options
| Option | Description |
|---|---|
--dry-run | Report what would be reclaimed without deleting anything |
--remote <NAME> | Execute against a remote server (by remote name, e.g. origin) |
Description
Index builds leave superseded artifacts behind. The garbage collector normally reclaims them: each index root carries a manifest naming what the previous version replaced, and the collector releases exactly those. That only reaches artifacts some manifest records.
A sweep finds the rest. It enumerates what storage actually holds, subtracts everything reachable from a live index chain, and releases the remainder. Two sources account for most of it: a reindex published by Fluree 4.1.4 or earlier, which severed the chain and left every earlier index version unreachable, and dictionary blobs whose manifests were consumed on Fluree 4.2.0 or earlier, which left every dictionary blob to the sweep. A ledger with a long history on one of those versions can hold many times its live index in blobs no retention policy can reach.
A sweep covers every branch of a ledger, which is why LEDGER names the ledger rather than a branch — a branch-qualified alias is rejected. Dictionary blobs live in a namespace shared by all of a ledger’s branches, so releasing one is only safe with every branch’s index accounted for. Background GC accounts for the other branches by reading their chains before it releases a dictionary blob; a sweep holds every branch and unions their reachable sets, which is what lets it release blobs no manifest names any more. Soft-dropped branches count as live, so dropping a branch without purging it stays reversible.
Only index artifacts are considered. Commits, transactions, and config blobs are reachable through the commit chain rather than the index chain, so a sweep cannot establish that they are unreferenced and never touches them.
Index builds are held off for the duration. If any index root cannot be read, or any storage prefix cannot be listed, the sweep aborts without deleting anything — an incomplete picture of what is reachable would classify live artifacts as orphans.
Single-process deployments only. The hold that keeps index builds from writing during a sweep excludes the server’s own indexer. An external or second-process indexer writing to the same storage is not excluded, and its in-flight artifacts would be indistinguishable from orphans.
Examples
# See what would be reclaimed, without deleting
fluree sweep mydb --dry-run
# Reclaim them
fluree sweep mydb
# Against a remote server
fluree sweep mydb --remote origin
Output
mydb would reclaim 2847 of 5216 index artifacts (2369 still referenced)
fluree:file://mydb/main/index/roots/3f2a....fir6
fluree:file://mydb/@shared/dicts/9c1b....dict
...
Run without --dry-run to reclaim them.
Reclaimed 2847 artifacts from mydb
Artifacts that resist deletion are reported rather than treated as failures — they stay in storage and the next sweep retries them.
When to Use
- Disk usage far exceeds the data — the ledger directory is many times the size of its commits.
- After reindexing on Fluree 4.1.4 or earlier — those reindexes orphaned every earlier index version, and only a sweep reclaims them.
- After upgrading from Fluree 4.2.0 or earlier — those versions left every dictionary blob the collector replaced in storage, often the largest part of a ledger’s index footprint. One sweep clears the backlog; afterwards the collector keeps up and a sweep is a repair rather than routine maintenance. Running it when nothing is reclaimable is a safe no-op.
See Also
- reindex - Full rebuild from commit history (rebuilds; does not reclaim)
- index - Incremental index build
- Background indexing - Retention and reclamation in the server