Migrating between Datomic and datahike

August 15, 2026 · View on GitHub

datahike.migrate.datomic moves a database in either direction between Datomic Pro and datahike, over the same record seam dumps use (see Migration).

Experimental, and deliberately at arm's length: it lives on its own source path, takes no dependency datahike itself carries, and is reachable only if you put the Datomic peer jar on your classpath. It may well move out of datahike into a library of its own — nothing in datahike depends on it, which is what would make that move cheap. Treat the API as likely to change.

(require '[datahike.migrate.datomic :as dtm])

(dtm/import-from-datomic! dh-conn datomic-conn)   ; Datomic  -> datahike
(dtm/export-to-datomic!   @dh-conn datomic-conn)  ; datahike -> Datomic

Getting the namespace

datahike takes no Datomic dependency. The namespace requires datomic.api, so it loads only when you put the peer jar on your own classpath:

{:deps {org.replikativ/datahike {:mvn/version "…"}
        com.datomic/peer        {:mvn/version "1.0.7622"}}}

Nothing else in datahike requires it, so if you never add the dependency you never pay for it.

Scope: Pro, via the peer API

The peer API (datomic.api) is Pro only. Datomic Cloud and Datomic Local speak the client API (datomic.client.api), which has no d/log — and the transaction log is what the source reads to reproduce history. Supporting them means a different namespace and a different strategy for history, not a different require.

datomic:mem:// works, which is what the test suite uses: no licence key, no transactor, no container.

What survives, and what cannot

Values, history and transaction times survive. Entity and transaction ids do not. That is not a shortcut; it is arithmetic.

Datomicdatahike
user entity id~1.76e13emax = 2 147 483 647
transaction entity id~1.3e13txmax = 2 147 483 647

Datomic's ids are four orders of magnitude above what datahike can represent, so {:eids :preserve} is refused rather than silently downgraded — datahike does not range-check an incoming eid, it simply reallocates, so passing them through would look like it had worked. :eids still accepts a map or a function if you know how the two id spaces should relate.

In the other direction Datomic refuses explicit entity ids outright (:db.error/invalid-entity-id), so a datahike → Datomic export always reallocates too.

datahike also assigns its own t: a stream is renumbered sequentially in source order. What is preserved is the transaction order and every :db/txInstant, which is what history actually depends on.

Provenance: which transaction was which

Because the ids do not survive, the correspondence is recorded as data. Each imported transaction carries two datoms on the transaction entity:

attributevalue
:datomic/tthe source t (e.g. 1004)
:datomic/tx-eidthe source transaction entity id

So the question is a query rather than arithmetic:

(d/q '[:find ?tx :in $ ?t :where [?tx :datomic/t ?t]] @conn 1004)
;=> #{[536870917]}

Pass {:provenance? false} to leave them out.

Round trips

Both directions are covered by tests that compare against the original rather than merely checking the result is non-empty.

Datomic → datahike → Datomic is datom-for-datom identical to the source, with exactly one difference: one extra transaction, and its timestamp — however many schema transactions the source had.

Datomic will not let an attribute be used — appear as the attribute of a datom — in the same transaction that installs it. datahike will, and the import relies on that: it emits the schema for :datomic/t and :datomic/tx-eid with the log's first transaction and stamps that same transaction with them. So it is the provenance that forces the split, not anything the Datomic source did — a Datomic source could not have produced such a transaction, since Datomic would have refused it. {:provenance? false} removes the cause and the extra transaction with it.

The schema half is stamped one millisecond before the source instant, so the two still ascend; its content and ordering are exact, its installation time is approximate.

Every other schema transaction comes back whole. The sink splits on a transaction that installs an attribute and uses it, not on the mere presence of schema — an earlier version split on presence, which cost one wasted transaction per schema transaction (measured: four came back as six).

datahike → Datomic → datahike preserves current values and history.

Limits worth knowing before you start

Export into a fresh Datomic database. :db/txInstant must be monotonic and at or after the database basis. A fresh database's basis instant is 1970-01-01, so real historical times replay from the first transaction onward. Appending into a populated database whose basis is newer than your oldest record cannot work, and Datomic says so with :db.error/past-tx-instant.

Re-asserting a :db.unique/identity value upserts. Exporting into a non-empty Datomic database merges onto existing entities rather than duplicating, which may or may not be what you want.

datahike-only schema is stripped. Datomic answers :db/maxLength, :db.valid/from, :db.valid/to, :db.secondary/* and datahike's extra value types with :db.error/not-an-entity, which would fail the transaction. {:strip-datahike-schema? false} keeps them and lets Datomic refuse. Unknown application attributes are fine either way.

Datomic-only vocabulary is dropped, with a warning: transaction functions (:db/fn, :db.fn/cas), :db/fulltext, :db/lang, and the excision vocabulary have no datahike equivalent. :db.type/uri values arrive as their string form, losing the type (#135).

Excised data cannot be recovered by any reader — it is absent from the log by definition.

Importing into a database that already has data

{:merge? true} lifts the empty-target refusal. It is append-only: transact-entities-directly does not resolve :db.unique/identity, so importing the same Datomic database twice adds the entities twice rather than upserting onto the existing ones.

Memory

Descriptors are t windows, not records, and :read fetches one window at a time, so neither the descriptor list nor a read holds the log. Measured on 3000 transactions / 12 012 datoms under -Xmx700m: 60 descriptors, heap 61 → 63 MB across the whole import, and 0 MB to build the source.

:window (default 100) sets transactions per chunk. A chunk is whole transactions by construction, so transaction alignment is free.

Verification

import-from-datomic! counts the source by default so the import can be verified, which costs one extra pass over the log. {:count? false} skips it — and then requires {:verify? false}, so an unverified import is always something you asked for rather than something that quietly happened.

Running the tests

bb test datomic

Needs the :datomic alias for the peer jar. The tests live on their own source path (test-datomic) precisely because the namespace fails to load without Datomic rather than skipping, so it must be unreachable from every other tier.