Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

Native bitemporal support

Open
#4,954 0 comments 0 reactions 0 assignees View on GitHub

Maintainers usually reply within 1 day

Nobody has claimed this yet.

Assessment

Difficulty
5/5
Estimated time
Over a week
Newbie friendliness
25/100
Issue type
Feature
Clarity
Needs clarification
Activity status
Active
Tech stack
cassandra, java

Research direction

Read the companion document's Section 9.5 and review the proposed measurement prototype for the in-memory and BerkeleyJE backends. Run the TinkerPop test suites against an all-bitemporal graph, measure storage, read, write, and backfill costs, and audit paths that assume one id maps to one stored column; the prototype branch and results should resolve the design before shippable work begins.

Written by the indexing model from the issue text.

Description

Proposal: Native Bitemporal Support in JanusGraph

This is a proposal for a major new feature under the Development Decisions process. This issue is the reference point; a [DISCUSS] Native bitemporal support thread on janusgraph-dev will follow, and a VOTE will only be called after the measurement prototype described at the end has results and we have general consensus on the design implementation. Nothing here is committed to; every part of the design is open.

Summary

We propose adding bitemporal support to JanusGraph as a per-type schema feature, modeled on SQL:2011 temporal tables. A user would mark an edge label, property key, or vertex label as bitemporal; JanusGraph would then keep every version of those elements, with both the time a fact was true and the time the database recorded it, and let any traversal be evaluated "as of" either. Nothing changes for types or graphs that do not opt in. This proposal does not require changes to TinkerPop. We are asking whether this belongs in JanusGraph, and if so, whether this is the right shape for it.

Problem

Two questions come up repeatedly in systems that track people, organizations, assets, or contracts:

What was true at a point in time? "What was Alice's title on 1 March 2024?" or "Which suppliers will be active next quarter?"

What did we believe at a point in time? "What did the report we ran on 1 March 2024 say Alice's title was, and when was it corrected?"

The first is valid time: when a fact holds in the real world. The second is system time: when the database knew it. Audit, compliance, and reproducibility need both.

JanusGraph users approximate valid time today by adding validFrom/validTo properties and filtering at every step of every traversal, which is what our own indexing documentation recommends. That approach is fragile (one missed filter silently returns the wrong state), cannot use the column sort order, and cannot be checked by the engine.

System time cannot be approximated at all. JanusGraph deletes physically: an update to a single-valued property is a delete plus a write, and a removed edge is gone. The storage read API has no version parameter, and the HBase backend returns only the latest cell even where versions are retained. No plugin, strategy, or application library can recover what the store no longer holds. Only the commit path can record it, and only the engine owns the commit path.

The demand is real enough that a funded research team (TEMPO, University of Patras) forked JanusGraph and Gremlin to add valid-time support, reporting a six-month delay learning the codebase; the fork implements one of the two axes and nothing came back upstream.

Requirements

Opt-in per type. Bitemporality is declared on individual edge labels, property keys, and vertex labels. Types that do not opt in are unaffected in storage, behavior, and performance.

Existing traversals keep their meaning. A traversal with no time options returns the state that is valid now and known now, on bitemporal and non-bitemporal types alike.

Both axes are queryable, together or separately, from every Gremlin client without changes to TinkerPop.

Valid time is user-controlled; system time is not. Users may record facts about the past or the future. The engine alone assigns when something was recorded.

Identity is stable. An edge keeps its id across versions and corrections.

Existing graphs upgrade with no data change and remain openable by older releases until an operator explicitly enables the feature on that graph.

Existing types can become bitemporal in place, without recreating the graph.

The cost is paid only by types that use it, and history retention is configurable.

Proposal

Reference model: SQL:2011
SQL:2011 defines application-time period tables (valid time) and system-versioned tables (system time), implemented in SQL Server, MariaDB, Db2 and Oracle as a current table plus a history table. We follow that model directly, so the semantics are already specified and documented rather than invented here.

SQL:2011 JanusGraph proposal
Application-time period Valid time: ~validFrom, ~validTo, user-set
System-time period System time: ~systemFrom, ~systemTo, engine-set at commit
Primary key stable across versions Logical relation id stable across versions; g.E(id) returns the version visible in the current view
PRIMARY KEY (id, period WITHOUT OVERLAPS) Per-type overlap policy; reject by default for single-valued types
Current table + history table Existing edgestore holds current knowledge; a new history store holds every version
FOR VALID_TIME AS OF / FOR SYSTEM_TIME AS OF g.with('janusgraph.valid-time', t) / g.with('janusgraph.system-time', t)
Enabling versioning on an existing table starts history at that moment Enabling on an existing type runs a backfill; history begins at that "horizon"

What a user sees

mgmt = graph.openManagement()
mgmt.enableBitemporal()            // once per graph; explicit, one-way, warns about downgrade
mgmt.makePropertyKey('salary').dataType(Integer.class).cardinality(Cardinality.SINGLE).bitemporal().make()
mgmt.makeEdgeLabel('worksFor').multiplicity(Multiplicity.MANY2ONE).bitemporal().make()
mgmt.commit()

g.V().has('name','alice').values('salary')                                       // valid now, known now
g.with('janusgraph.valid-time', datetime('2024-01-01T00:00:00Z'))
 .V().has('name','alice').values('salary')                                       // what was true then
g.with('janusgraph.valid-time',  datetime('2024-01-01T00:00:00Z'))
 .with('janusgraph.system-time', datetime('2024-03-01T00:00:00Z'))
 .V().has('name','alice').values('salary')                                       // what we believed then about then
g.with('janusgraph.valid-time', datetime('2026-10-01T00:00:00Z'))
 .V(alice).property('salary', 120000)                                            // scheduled; invisible until October
g.with(key, value) already exists in every Gremlin client and reaches the provider; TinkerPop itself uses the same mechanism for its own options.

How it works

Storage. The edgestore is unchanged and continues to hold current knowledge. For a bitemporal type, it holds every currently-known version with its valid interval, using the existing multi-valued column layout with validFrom in the sort key. A new store, opened through the same backend API as the edgestore, holds every version ever recorded with both time intervals. Reads with no time options, and valid-time reads, use the edgestore; system-time reads use the history store. History is retained under a configurable TTL.

Writes. A write to a bitemporal type writes the new version to both stores and closes the prior version in the history store with the commit time. A removal ends a fact's validity and closes it; nothing is physically deleted from history except by retention. A user-supplied commit time is rejected for transactions that touch bitemporal types, because system time must not be forgeable.

Enabling. enableBitemporal() moves the graph to storage version 3. Graphs that never call it stay at version 2, upgrade with zero writes, and remain openable by older releases. Graphs that call it can no longer be opened by releases that predate the feature, in the same way the 0.3.0 storage version bump worked. The call requires the existing graph.allow-upgrade=true opt-in and exclusive access, and logs a warning saying exactly this. We chose a hard refusal over a transparent downgrade because a downgraded release would write to the edgestore without recording history, leaving gaps that could not be detected afterward.

Existing types. enableBitemporal(type) follows the index lifecycle: registered, backfill job, enabled. Existing rows receive a system-time start equal to the enable instant and valid-time bounds from named existing keys if the user has them. Queries before that horizon raise an error by default so that "no history" is never mistaken for "did not exist."

What does not change

The byte layout of non-bitemporal types; the wire format; existing indexes; results of any existing traversal; the storage version of any graph that does not opt in; TinkerPop.

Alternatives considered

  • Leave it to applications (validFrom/validTo properties, filtered per step). This is today's answer for valid time. It cannot provide system time at all, because the engine deletes physically and nothing outside the commit path can observe what was overwritten.

  • Store every version inline in the existing edgestore columns. Changes the byte layout of bitemporal types, puts version filtering on the hot path of every read of those types, and turns retention into a scan job. The separate history store keeps the edgestore layout and read path as they are.

  • Use backend cell versions (HBase, Cassandra timestamps). The storage read API has no version parameter, the HBase backend already collapses multiple versions to the latest, CQL and BerkeleyJE retain nothing queryable, and cell versions cannot express valid time.

  • Derive history from the transaction log or CDC. Both are event streams for replay, not indexed state; answering "state at (valid time, system time)" would mean replaying, logs expire, and neither carries valid time. Both stay useful for downstream consumers.

  • Wait for TinkerPop to define temporal semantics. TinkerPop defines no temporal model today; g.with() is the provider-option channel TinkerPop itself uses. A first-class Gremlin surface can be proposed to TinkerPop later, after the semantics have been exercised in JanusGraph.

Where we need input

  • Is this a problem JanusGraph should solve natively? We have prior art and a fork, but no census of users who would adopt it. If you would use this, or have built the property-based workaround, please say so here.

  • Is the per-type, SQL:2011-shaped model the right one for a property graph?

  • Is the one-way storage-version bump acceptable, given it is lazy and explicit?

  • The write API (transaction-scoped valid time plus writable ~validFrom/~validTo) is the least settled part.

Next step

Before proposing any shippable code, we will build a measurement prototype on the in-memory and BerkeleyJE backends to put numbers on storage growth, write cost, read cost with many versions, and backfill throughput, run the TinkerPop test suites against an all-bitemporal graph, and audit the code paths that assume one id maps to one stored column. The measurement plan is published in the companion document (Section 9.5) so it can be challenged before the results exist; the prototype branch will be linked from this issue when it exists. The first shippable change would be small and independent: accepting storage versions 2 and 3 plus a compatibility test that opens a graph from the current release and proves nothing was written.

Dominant language
Java
Stars
5.8k
Forks
1.2k
Avg merge
20h 36m
Merged PRs (30d)
25

Getting set up

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from JanusGraph/janusgraph

All issues in JanusGraph/janusgraph

Similar issues

More Java issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.