Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Datatypes and Typed Values

Fluree enforces strong typing for all literal values, ensuring data consistency and enabling efficient indexing and querying. Every literal value has an explicit datatype, following RDF and XSD standards.

Core Principle: No Untyped Literals

Unlike some databases that allow “plain” strings, Fluree requires every literal to have a datatype. This design provides:

  • Type Safety: Prevents type confusion in queries and applications
  • Consistent Comparisons: Typed values compare predictably
  • Standards Compliance: Follows RDF and SPARQL specifications
  • Query Optimization: Enables efficient indexing and query planning

XSD Datatypes

Fluree supports the core XML Schema Definition (XSD) datatypes:

String Types

{
  "@context": {
    "xsd": "http://www.w3.org/2001/XMLSchema#",
    "ex": "http://example.org/ns/"
  },
  "@graph": [
    {
      "@id": "ex:book1",
      "ex:title": "The Great Gatsby",
      "ex:author": {
        "@value": "F. Scott Fitzgerald",
        "@type": "xsd:string"
      }
    }
  ]
}

xsd:string is the default for plain string literals when no type is specified.

Numeric Types

{
  "@graph": [
    {
      "@id": "ex:product1",
      "ex:price": {
        "@value": "29.99",
        "@type": "xsd:decimal"
      },
      "ex:quantity": {
        "@value": "100",
        "@type": "xsd:integer"
      },
      "ex:rating": {
        "@value": "4.5",
        "@type": "xsd:double"
      }
    }
  ]
}

Supported numeric types:

  • xsd:integer: Whole numbers (-∞, ∞)
  • xsd:long: 64-bit integers
  • xsd:int: 32-bit integers
  • xsd:short: 16-bit integers
  • xsd:byte: 8-bit integers
  • xsd:decimal: Arbitrary precision decimals
  • xsd:double: 64-bit floating point
  • xsd:float: 32-bit floating point

Boolean Type

{
  "@graph": [
    {
      "@id": "ex:user1",
      "ex:isActive": {
        "@value": "true",
        "@type": "xsd:boolean"
      },
      "ex:hasVerifiedEmail": {
        "@value": "false",
        "@type": "xsd:boolean"
      }
    }
  ]
}

xsd:boolean accepts: true, false, 1, 0.

Date and Time Types

{
  "@graph": [
    {
      "@id": "ex:event1",
      "ex:startDate": {
        "@value": "2024-01-15",
        "@type": "xsd:date"
      },
      "ex:startTime": {
        "@value": "14:30:00Z",
        "@type": "xsd:time"
      },
      "ex:createdAt": {
        "@value": "2024-01-15T14:30:00Z",
        "@type": "xsd:dateTime"
      }
    }
  ]
}

Temporal types:

  • xsd:date: Dates without time (e.g., 2024-01-15)
  • xsd:time: Times without date (e.g., 14:30:00Z)
  • xsd:dateTime: Full timestamps (e.g., 2024-01-15T14:30:00Z)

Other XSD Types

{
  "@graph": [
    {
      "@id": "ex:resource1",
      "ex:homepage": {
        "@value": "https://example.com",
        "@type": "xsd:anyURI"
      },
      "ex:duration": {
        "@value": "PT1H30M",
        "@type": "xsd:duration"
      }
    }
  ]
}

Additional types include:

  • xsd:anyURI: Web addresses and identifiers
  • xsd:duration: Time periods (ISO 8601 format)
  • xsd:dayTimeDuration, xsd:yearMonthDuration: The two xsd:duration subtypes. Unlike xsd:duration these are totally ordered, so they sort and compare. xsd:dayTimeDuration is also the type queries produce when subtracting two temporal values — see Date/Time Arithmetic.
  • xsd:gYear, xsd:gMonth, xsd:gDay: Partial date components

RDF Datatypes

Beyond XSD, Fluree supports RDF-specific datatypes:

Language-Tagged Strings

{
  "@graph": [
    {
      "@id": "ex:book1",
      "ex:title": {
        "@value": "The Great Gatsby",
        "@language": "en"
      },
      "ex:titel": {
        "@value": "Der große Gatsby",
        "@language": "de"
      }
    }
  ]
}

rdf:langString represents strings with language tags. This is distinct from plain strings and enables language-aware queries.

Matching a string literal in a query. "bob", "bob"@en and "bob"@fr are three different RDF terms. In SPARQL a constant object matches only its own term: ?s ex:name "bob"@en returns just the English value, ?s ex:name "bob" (an xsd:string) does not match tagged values, and a variable bound to a string joins only against its own term — the same language tag, or the same datatype, which holds for xsd:anyURI, xsd:token and customer-defined datatypes exactly as it does for xsd:string. In a JSON-LD query, {"@value": "bob", "@language": "en"} and {"@value": "bob", "@type": "xsd:string"} are likewise exact, while a bare JSON string as a constant object ("ex:name": "bob") matches the lexical value under any string datatype or language tag. That leniency is specific to the constant-object position: the same bare string in a values cell is an xsd:string term and matches only xsd:string rows. An explicitly typed literal is exact for its datatype in both surfaces: "25"^^xsd:int (or {"@value": "25", "@type": "xsd:int"}) matches xsd:int rows only, not the same number stored as xsd:long, and "25"^^xsd:integer written out matches xsd:integer rows only. Bare numeric literals (25, 25.0, and a JSON number in a JSON-LD query) are the one lenient form — they match an equal value under any numeric datatype, so 25 finds xsd:integer, xsd:long, xsd:int and xsd:double values alike. Language tags are case-insensitive (BCP 47) and are stored and compared in lowercase: "chat"@FR is written, matched, and returned by LANG() as fr.

Annotating literal values. Any literal — plain, typed, or language-tagged — can carry statement-level metadata (source, confidence, timestamp) via an edge annotation. Because a JSON scalar has no room for sibling keys, an annotated literal must be written in value-object form (@value plus @annotation); language-tagged annotations are language-pinned, so "chat"@fr and "chat"@en annotate independently. See Edge annotations → Annotating literal-valued edges.

JSON Data

{
  "@graph": [
    {
      "@id": "ex:config1",
      "ex:settings": {
        "@value": "{\"theme\": \"dark\", \"notifications\": true}",
        "@type": "@json"
      }
    }
  ]
}

rdf:JSON stores JSON data as typed literals. This is useful for storing complex structured data that doesn’t fit the RDF model.

A JSON literal is stored in canonical form, per RFC 8785. JSON-LD 1.1 requires this. Members are sorted by key, insignificant whitespace is dropped, and numbers are rendered as ECMAScript renders them.

A literal’s value is its text. Without canonical form, {"a":1,"b":2} and {"b":2,"a":1} would be two different facts. A delete, an upsert, or a merge that matched one would miss the other.

Two consequences are worth knowing:

  • A value reads back canonicalized, not as written. Members come back in key order. A number written as 1.0 reads back as 1. The JSON is the same. Only its spelling changes.
  • Integers keep their exact value. RFC 8785 treats every number as a double, which rounds integers past 2^53. Fluree keeps them exact instead.

@value may be a JSON document or a string holding one. Either is canonicalized. A string that is not valid JSON is stored as written.

Ledgers written before canonicalization keep the text their writers produced. Reindexing does not change that: it rebuilds the index from the commits, and the commits hold the original text. Those values stay readable, and new writes are canonical.

Two things follow for a ledger with older values.

  • Naming a value no longer matches it. A literal in a where, a delete, or a FILTER is canonicalized before it is compared, so it misses a value stored under another spelling. Bind the value with a variable instead.
  • Re-asserting a value can leave two of them. An upsert writes the canonical spelling. If its delete names the old text, the delete misses and both spellings end up on the subject. Any duplicates the missing canonicalization already created stay as they are.

Repair a subject by binding the old value and writing it back:

{
  "@context": {"ex": "http://example.org/ns/"},
  "where": {"@id": "ex:config", "ex:items": "?old"},
  "delete": {"@id": "ex:config", "ex:items": "?old"},
  "insert": {
    "@id": "ex:config",
    "ex:items": {"@value": [{"name": "alpha", "qty": 1}], "@type": "@json"}
  }
}

The where clause binds every spelling on that property, so this collapses duplicates as well. History keeps the original text; only the current state changes. fluree validate finds the properties worth repairing wherever a shape constrains them.

Geographic Data

{
  "@context": {
    "geo": "http://www.opengis.net/ont/geosparql#",
    "ex": "http://example.org/"
  },
  "@graph": [
    {
      "@id": "ex:location1",
      "ex:coordinates": {
        "@value": "POINT(2.3522 48.8566)",
        "@type": "geo:wktLiteral"
      }
    }
  ]
}

geo:wktLiteral stores geographic data in Well-Known Text (WKT) format. POINT geometries are automatically converted to an optimized binary encoding, while other geometry types (polygons, lines) are stored as strings.

See Geospatial for complete documentation.

Vector Data

{
  "@context": {
    "ex": "http://example.org/"
  },
  "@graph": [
    {
      "@id": "ex:doc1",
      "ex:embedding": {
        "@value": [0.1, 0.2, 0.3, 0.4],
        "@type": "@vector"
      }
    }
  ]
}

@vector (full IRI: https://ns.flur.ee/db#embeddingVector, prefix form: f:embeddingVector) stores numeric arrays as embedding vectors. Values are quantized to IEEE-754 f32 at ingest for compact storage and SIMD-accelerated similarity computation. In Turtle/SPARQL, use f:embeddingVector with the ^^ typed-literal syntax.

Without this type annotation, plain JSON arrays are decomposed into individual RDF values where duplicates may be removed and ordering is lost.

See Vector Search for complete documentation.

Fulltext Data

{
  "@context": {
    "ex": "http://example.org/"
  },
  "@graph": [
    {
      "@id": "ex:article-1",
      "ex:content": {
        "@value": "Rust is a systems programming language focused on safety and performance",
        "@type": "@fulltext"
      }
    }
  ]
}

@fulltext (full IRI: https://ns.flur.ee/db#fullText, prefix form: f:fullText) marks a string value for full-text search indexing. Values annotated with @fulltext are automatically analyzed (tokenized, stemmed, stopword-filtered) and indexed into per-predicate fulltext arenas during background index builds. This enables BM25-ranked relevance scoring via the fulltext() query function.

Because the annotation is the stored datatype, tagged values are f:fullText literals rather than xsd:string — a custom datatype that external RDF consumers won’t recognize. For standards-conformant data, prefer declaring searchable properties via f:fullTextDefaults in the ledger’s #config graph, which full-text indexes plain xsd:string / rdf:langString values (with per-language analysis) without changing their datatypes.

Without this type annotation, strings are stored as plain xsd:string values and support only exact matching and prefix queries – not relevance-ranked full-text search.

See Inline Fulltext Search for complete documentation.

Custom Datatypes

Any IRI can serve as a datatype. A literal with a datatype Fluree does not recognize is stored with that datatype and returned exactly as written.

{
  "@context": {"unit": "http://example.org/unit/"},
  "@id": "ex:room1",
  "ex:area": {"@value": "42.5", "@type": "unit:SquareMetre"}
}

In Turtle and SPARQL, use the ^^ syntax: "42.5"^^unit:SquareMetre.

Datatype Limit

A ledger holds at most 16,369 distinct datatypes, plus 15 reserved ones that never count toward the limit. The reserved datatypes are @id, xsd:string, xsd:boolean, xsd:integer, xsd:long, xsd:decimal, xsd:double, xsd:float, xsd:dateTime, xsd:date, xsd:time, rdf:langString, rdf:JSON, @vector, and @fulltext. Every other datatype counts, including other XSD types such as xsd:int and xsd:anyURI.

A datatype counts from the first write that uses it. It still counts after its data is retracted, because the index never releases a datatype’s ID.

A write that would pass the limit is refused, and the ledger is left unchanged:

  • Transactions and SPARQL updates fail with HTTP 422 and err:db/DatatypeLimitExceeded. See Common Errors.
  • Pushed commits are refused with the same error.
  • A bulk import (fluree create --from) fails with a “datatype limit exceeded” error.

Upgrading and Downgrading

Fluree 4.2.1 and earlier cannot index a ledger that holds more than 241 non-reserved datatypes. Those releases still accept writes past that point, so such a ledger may already exist. Later releases index it with no migration.

Rolling back to 4.2.1 or earlier affects any ledger that holds more than 241 non-reserved datatypes. The older release can read the ledger’s existing index, but its index builds can fail. When they do, new data stays in novelty, where queries are slower, and writes are refused once novelty reaches its size limit. Moving the ledger back to a newer release clears both.

Vocabularies of units or currencies can define hundreds of datatypes. If data needs more distinct datatypes than the limit allows, record the unit in its own property instead of in the datatype:

ex:room1 ex:area "42.5"^^xsd:decimal ;
         ex:areaUnit unit:SquareMetre .

Type Coercion and Compatibility

Automatic Type Promotion

Fluree handles type compatibility intelligently:

# This works - integer can be used where decimal is expected
SELECT ?price
WHERE {
  ?product ex:price ?price .
  FILTER(?price > 10.0)  # decimal comparison
}

Comparisons Between Incompatible Types

When a filter compares values of incompatible types (e.g., a number and a string), the behavior depends on the operator:

  • Equality (=) returns false — values of different types are never equal
  • Inequality (!=) returns true — values of different types are never equal
  • Ordering (<, <=, >, >=) raises an error — ordering between incompatible types is undefined

Numeric types (long, double, bigint, decimal) are mutually comparable via automatic promotion, so cross-numeric comparisons work as expected. Similarly, temporal types can be compared with string representations that parse to the same temporal type.

Text That Is Not a Value of Its Datatype

A literal such as "2024-02-30"^^xsd:date or "300"^^xsd:byte names a built-in datatype but is not one of its values. JSON-LD transactions and SPARQL UPDATE reject it, naming the datatype in the error. Turtle transactions and bulk import keep it, as RDF allows: the literal is stored with its text and its datatype and reads back exactly as written, but it is not a date or a number, so it never equals one.

Type Casting in Queries

SPARQL provides functions for type conversion:

SELECT ?name (xsd:string(?id) AS ?idString)
WHERE {
  ?person ex:name ?name ;
          ex:id ?id .
}

Best Practices

Choosing Datatypes

  1. Be Specific: Use the most appropriate type for your data

    • Use xsd:integer for whole numbers that will be used in calculations
    • Use xsd:string for identifiers and labels
    • Use xsd:dateTime for timestamps
  2. Consider Query Patterns: Choose types that support your intended queries

    • Numeric types enable range queries and aggregations
    • Date types enable temporal queries
    • String types support text search
  3. Standards Alignment: Use standard datatypes where possible

Type Consistency

  1. Consistent Usage: Use the same datatype for equivalent properties across your data
  2. Change Planning: Plan for type changes as your data model evolves
  3. Validation: Validate data types at ingestion time

Performance Considerations

  1. Index Efficiency: Different types have different indexing characteristics

    • Numeric types support efficient range queries
    • String types support prefix and substring matching
    • Date types enable temporal range queries
  2. Storage Size: Some types are more storage-efficient than others

    • xsd:integer is more compact than xsd:string
    • xsd:boolean is more efficient than string representations

Type System Architecture

Internal Representation

Fluree stores all typed values with their datatype information:

  • Value Storage: The literal value as a string
  • Type Metadata: The datatype IRI
  • Comparison Logic: Type-aware comparison functions

Query Processing

The type system affects query processing:

  • Type Checking: Ensures type compatibility in filters and joins
  • Index Selection: Chooses appropriate indexes based on types
  • Result Formatting: Formats results according to datatype rules

Standards Compliance

Fluree’s type system is fully compliant with:

  • RDF 1.1 Concepts: Literal typing requirements
  • SPARQL 1.1: Type promotion and compatibility rules
  • XSD 1.1: Datatype definitions and constraints
  • JSON-LD 1.1: Typed value syntax

This strong typing foundation ensures data consistency, enables optimization, and maintains interoperability with the broader semantic web ecosystem.