Troubleshooting
Common failures and how to resolve them.
"Failed to connect to socket: …"
The cardano-node Unix socket isn't reachable.
- Confirm
cardano-nodeis running. - Check that the path you pass to
--socket-pathmatches the path the node was started with. - Check file permissions — dbsync needs read access. The easiest fix is to run both as the same OS user.
If the socket file exists but the connection still fails, the node
is probably still starting up — its LedgerDB replay can take minutes
on a populated chain. Wait until you see Chain extended in the
node's log, or query the tip via cardano-cli.
"could not connect to server: Connection refused" (PostgreSQL)
PG isn't reachable on the configured host/port.
- On a local install:
systemctl status postgresql(Linux) orbrew services list | grep postgres(macOS). - Check the host / port in your
--pg-configfile. - For remote PG, confirm
pg_hba.confallows the dbsync host.
"FATAL: role "..." does not exist"
The PG user named in the pg-config file doesn't exist. Create it:
createuser --createdb dbsync
Or set user: "" in the pg-config file to fall back to the OS user
(peer authentication on a local install).
"permission denied for database "cexplorer""
The PG user exists but lacks CREATEDB or grants on an existing
database. For a fresh sync, the simplest fix is to grant CREATEDB
and have dbsync create the database itself; for a constrained
production user, the operator should pre-create the database and
grant CONNECT, CREATE, and USAGE on the public schema.
"… extractor requires … to be enabled"
Config validation rejected the combination and named the missing
dependency. Add it to db_profile, or enable ledger. The enforced
rules:
multi_asset→ needsutxo.off_chain_pools→ needspool.off_chain_votes→ needsgovernance.epoch_boundary,pool_stats,stake_delegation_ledger,current_state→ needledger.enabled = true.
See Custom configs for the dependency table.
"Schema mismatch — refusing to start"
You changed the config against an existing database. This is the profile immutability guard. Your options:
- Revert the config to match what's in the database (resume the existing sync).
- Re-sync against a fresh database (drop the existing one).
- Pass
--resync-from-genesisto wipe and start over (destructive).
"Cannot resume: the database was synced against a different network"
On first run dbsync records the network in the database
(dbsync_sync_state.network_magic / network_name); every later boot
compares that against the genesis reachable through --node-config
and refuses to interleave two chains. The message names both sides:
Database : preview (magic 2)
This run : mainnet (magic 764824073)
- Wrong
--node-config(the common case): point it at theconfig.jsonof the network this database was synced against. - Actually switching networks: use a fresh database, or pass
--resync-from-genesisto wipe this one (destructive).
"WARN: wal_level is replica; consider 'minimal' during initial sync"
dbsync emitted this at boot because PG is configured with
wal_level = replica (the default). It's not an error — the sync
will work — but the UNLOGGED → LOGGED flip in
PreparingForVolatileTail writes
every row to WAL on replica, materially increasing wall-clock time
for a large profile.
If you don't have replicas to worry about, set wal_level = minimal
in postgresql.conf for the sync duration. See
scripts/postgres-tuning.conf in the repo for a tuned starting
point — typically a 20–30% reduction in Prep wall-clock time on a
large profile.
"FATAL: too many open files" / "EMFILE"
Hit on Linux during Ingest. dbsync opens one libpq connection per extractor table, plus connections for the control channel and the TxOut worker. On a large profile this can exceed the default per-user fd limit (1024 on most distros).
Raise the limit:
ulimit -n 8192
Add it to /etc/security/limits.conf for a permanent fix:
your-user soft nofile 8192
your-user hard nofile 16384
dbsync also checks the limit at boot
(Phase.Ingest.FdLimit)
and aborts early if it's clearly too low.
Disk full during Ingest
PG runs out of space partway through the bulk-load.
- UNLOGGED tables don't shrink — once allocated, the space is held.
Recovery is a
DROP DATABASEand a re-sync against a larger disk. - Index builds during
PreparingForVolatileTailtemporarily allocate roughly the size of the indexed table. If Ingest itself fit but Prep doesn't, dropping the smallest disabled extractor's tables and re-running can sometimes recover; usually it's cleaner to re-sync to a bigger disk.
The size table in Prerequisites gives rough working figures; budget 50% headroom over the listed sizes.
OOM during ledger replay
ledger.enabled = true puts an in-RAM ledger state on top of the
sync. On mainnet at tip that's around 8 GB of resident memory; less
on testnets.
If the process is OOM-killed:
- Confirm you have at least 16 GB of RAM total (8 for the ledger,
plus PG's
shared_buffers, plus the sync's working set). - If memory is genuinely tight, disable the ledger and use
everything-no-ledger.jsoninstead. You lose rewards / deposits / protocol-param tables; everything else still works.
Slow Ingest — diagnostics
The expected throughput on mainnet is roughly:
- 100–500 blocks/sec early in Byron (small blocks, mostly UTxO).
- 30–80 blocks/sec late in Alonzo/Babbage (large blocks, Plutus scripts, lots of metadata).
- 10–30 blocks/sec late in Conway (governance + large NFT mints).
If you're consistently below these, in order of likelihood:
- PG isn't tuned. Check
wal_level,shared_buffers,maintenance_work_mem,max_parallel_maintenance_workers. The shippedscripts/postgres-tuning.confis a reasonable starting point. - Disk is slow. Even on SSD, a heavily-fragmented filesystem or a network-mounted PG data directory hurts a lot. Local NVMe is the sweet spot.
- Too many extractors enabled for your machine. Drop to a smaller profile and re-sync.
liburingisn't installed on Linux and the build fell back to+serialblockio. The fallback is correct but noticeably slower for the LSM dedup stores and (if enabled) the LedgerDB.
Set logging.level = "debug" and re-start; you'll get per-epoch
timing breakdowns showing exactly where the time goes (LSM
compaction vs COPY vs tx-out worker drain).
"rollback exceeds k blocks (depth N > 2160)"
The Follow loop received a MsgRollback deeper than the protocol
security parameter k. This is either a node bug or operator error
(rolling back manually to a slot more than 2160 blocks behind the
tip). dbsync panics rather than silently corrupting the database.
The recovery path is --rollback-to-slot SLOT with a slot at or
after the new chain's intersection point, or --resync-from-genesis
if you can't determine one. See Recovery.
"Ingest scratch state was wiped — refusing to resume"
The ingest LSM session at <state-dir>/dbsync-ledger/ingest-lsm/ is
missing or corrupt while the database still shows Ingest as
incomplete. Recovery: pass --resync-from-genesis to start over.
This is unusual — the only way to hit it is to manually delete the scratch directory mid-sync, which you shouldn't do.