2026-07-24
PostgreSQL PITR with WAL archiving and native incremental backups
Build and test PostgreSQL 18 physical recovery from a base backup and archived WAL, and verify native incremental base backup chains.
Physical point-in-time recovery depends on a usable base backup and an unbroken sequence of archived WAL from that backup through the recovery point. PostgreSQL restores the files, replays WAL to make them consistent, and can stop at a requested target. That is the model described in the PostgreSQL 18 continuous archiving documentation.
WAL archiving supplies the ordered change stream needed to replay a physical backup. PostgreSQL 18 native incremental base backups are a separate mechanism with their own prerequisites and restore step. Both can belong in one recovery design.
Make the WAL archive reliable
Set wal_level appropriately for physical backup and replication use, enable archiving, and make archive_command succeed only when the segment is safely present at its destination. PostgreSQL substitutes %p with the source path and %f with the filename. The command is retried after failure, so a retry-safe command must not report success for a missing, partial, or different archived file.
archive_mode = on
archive_command = 'test ! -f /archive/%f && cp %p /archive/%f || cmp %p /archive/%f'This is the disposable local-copy command used in the drill below. It is useful for demonstrating the required success semantics: copy a previously absent segment, or accept an existing byte-identical segment. It is not a production archive-storage design. Production storage needs its own authentication, durability, access-control, integrity, and failure-handling design. The documented command interface and placeholders are in continuous archiving.
Archived WAL normally arrives when a WAL segment completes. SELECT pg_switch_wal(); closes the current segment early, which is useful before checking that a restore point has reached the archive. archive_timeout can force a switch during quiet periods. A shorter timeout reduces the maximum time a quiet server can wait before sending a segment, but it creates more partially filled archive files and more archive work. Do not use it as a substitute for monitoring or for a tested recovery objective. The settings and tradeoffs are documented under WAL configuration, and pg_switch_wal is documented among the administration functions.
Watch pg_stat_archiver for failures and archive recency, and watch local pg_wal pressure. An archiver that cannot make progress can retain WAL locally until the filesystem fills. Copies must be separate from the primary's storage domain; a copy that fails with the primary is not a recovery dependency you can rely on.
Take and verify a physical base backup
pg_basebackup makes a physical cluster backup and writes backup_manifest by default. In this drill, -X stream included the WAL needed to make the copied cluster consistent, and pg_verifybackup checked the manifest and files afterward:
pg_basebackup -h 127.0.0.1 -U postgres \
-D /backup/full -Fp -X stream -c fast --manifest-checksums=SHA256
pg_verifybackup /backup/fullpg_verifybackup is worth making routine, but it verifies backup integrity against the manifest. It does not prove that archive retrieval, target selection, timelines, permissions, configuration, or a real server startup will work. Only a restore drill proves those dependencies together. See pg_basebackup and pg_verifybackup.
Use PostgreSQL 18 incremental base backups precisely
PostgreSQL 18 native incremental backups are not WAL archiving. They are pg_basebackup outputs containing changed relation blocks since an earlier backup. The server must have summarize_wal = on; WAL summary files in pg_wal/summaries must cover the interval from the reference backup to the new backup. Those summaries are separate from archived WAL. An archive may have every replay segment and still lack the summaries needed to create a later incremental backup.
Point --incremental at the earlier backup_manifest:
pg_basebackup -h 127.0.0.1 -U postgres \
-D /backup/inc -Fp -X stream -c fast \
--incremental=/backup/full/backup_manifest --manifest-checksums=SHA256
pg_verifybackup /backup/incAn incremental directory is not directly usable as a PostgreSQL data directory. Materialize a synthetic full backup from a valid oldest-to-newest chain before restoring it:
pg_combinebackup /backup/full /backup/inc -o /backup/combined
pg_verifybackup /backup/combinedpg_combinebackup documents that ordering and output. Retain the required WAL summaries for as long as future incrementals depend on them, using settings such as wal_summary_keep_time as part of the policy. They are not a replacement for the archived WAL needed to replay either a full or combined backup. The prerequisites and manifest reference are in the pg_basebackup documentation.
Do not assume incremental means smaller. It reflects changed blocks, non-relation files, and the workload. In the drill, deliberate change volume made the full backup 39M, the incremental 51M, and the combined backup 67M.
Recover with an explicit target
To start archive recovery, copy a base or combined backup into a new data directory, provide a restore_command, configure any target before startup, and create recovery.signal. PostgreSQL removes that signal when recovery completes and normal operation resumes.
Before the first startup, quarantine the restored snapshot from the primary and application networks, and keep ordinary clients out. Use an isolated network with no route to production, set listen_addresses to a loopback-only address where possible, and make pg_hba.conf permit only a dedicated local validation connection while rejecting broader client access. Keep those restrictions in place through promotion and validation; do not let applications reach the recovered instance.
restore_command = 'cp /archive/%f %p'
recovery_target_name = 'article_cut'
recovery_target_action = 'promote'
recovery_target_timeline = 'current'restore_command is another command boundary: it must retrieve the exact requested archived file or fail. The cp above again belongs only to the isolated local drill. The recovery setup is specified in continuous archiving.
PostgreSQL accepts one recovery target at a time. The choices are recovery_target = 'immediate', recovery_target_name, recovery_target_time, recovery_target_xid, and recovery_target_lsn. immediate stops at the first consistent state, usually the end of an online backup. Time, XID, and LSN targets are inclusive by default through recovery_target_inclusive; set it deliberately when the boundary matters. A named target comes from pg_create_restore_point('name'), which writes a named WAL marker. If recovery reaches the end of available WAL without the requested target, PostgreSQL treats that as a failure rather than silently claiming success. These target controls are documented in WAL configuration and pg_create_restore_point.
When a target is reached, recovery_target_action is pause, promote, or shutdown. pause is the default and leaves a verification window. Resuming with pg_wal_replay_resume() ends the pause and recovery proceeds to completion; it does not continue toward another target. promote ends recovery and accepts writes. shutdown stops the server at the target and leaves recovery.signal in place, so change the recovery configuration or remove the signal before restarting. Pick the action as part of the drill, not during an incident.
Treat timelines as history, not overwrites
Promotion after PITR creates a new timeline. PostgreSQL keeps the old history and does not overwrite its WAL, which permits another recovery attempt from the same base backup. The history file is itself a recovery dependency.
recovery_target_timeline = 'current' follows the timeline recorded by the backup, while latest follows the newest reachable timeline found in the archive. latest is useful only when the archive contains the needed WAL and timeline history and that is the branch you actually intend to follow. The timeline section of continuous archiving explains why recovery branches instead of replacing earlier WAL.
The PostgreSQL 18.4 drill
I ran this locally with postgres:18.4-bookworm, pinned in the evidence to postgres@sha256:1961f96e6029a02c3812d7cb329a3b03a3ac2bb067058dec17b0f5596aca9296. The disposable Docker script configured archive_mode=on, summarize_wal=on, the local archive command above, and wal_summary_keep_time=7d. Its exit trap removed the containers, volumes, and network.
- I inserted
base_backup, took and verified the full streamed-WAL backup, then insertedbefore_target. - I created the named restore point
article_cut, committedafter_target, generated change volume, switched WAL, and observed nine WAL summaries. - I took and verified a true incremental against the full backup manifest, combined full plus incremental, and verified the combined output. The archiver reported nine archived segments and zero failed attempts before recovery.
- From the full backup, I restored with
recovery_target_name = 'article_cut',recovery_target_action = 'promote', and timelinecurrent. The recovered table containedbase_backup,before_target, and excludedafter_target. - The promoted instance was on timeline 2. I inserted
timeline_two, switched and archived WAL, and observed00000002.historyin the archive. - A second restore from the same full backup with
recovery_target_timeline = 'latest'containedbase_backup,before_target,timeline_twoand excludedafter_targetfrom abandoned timeline 1.
The drill checked target selection, promotion, archived timeline history, and the difference between current and latest. It did not test S3 or any other object-storage implementation.
Keep the closure, not just old files
Retention is dependency closure, not age-only deletion. For every recovery window you advertise, retain:
- Each retained full base backup and its manifest. For an unmaterialized incremental, retain that incremental, its manifest, and every ordered ancestor backup and manifest in the chain. Alternatively, materialize and verify a combined backup before deleting any of those ancestors.
- All archived WAL required from that backup's replay start through the end of its recovery window, including the right timeline branches.
- Required timeline history files.
- WAL summaries required to make future native incrementals from retained references.
Removing any element can make an otherwise healthy-looking backup unrestorable. Model the oldest retained recovery point for each backup chain, then prove the model with a restore test. Test a target, not merely a server that starts.
Keep logical dumps in their lane
pg_dump creates a logical snapshot: SQL or an archive representation of database objects and data at dump time. It is valuable for selective restore, schema and data movement, reviewable exports, and migrations where logical compatibility is the goal. It is not an incremental physical backup, does not preserve a continuous WAL replay chain, and cannot restore an entire physical cluster to an arbitrary point between dumps. The scope and restore model are covered in the PostgreSQL backup and restore documentation.
Restore-drill checklist
- Verify every new full, incremental, and combined backup with
pg_verifybackup, then restore a representative chain. - Before declaring a target recoverable, force or wait for a WAL switch and confirm the required segment archived without archiver failures.
- Before starting the restored snapshot, quarantine its network and restrict
listen_addressesandpg_hba.confto the dedicated validation connection; keep applications excluded through promotion and validation. - Restore into isolated storage, use
recovery.signal, and test the exact target andrecovery_target_actionyou would use operationally. - After a promoted drill, archive new timeline WAL and its
.historyfile, then test whethercurrentandlatestproduce the intended branch. - Recalculate retention as backup plus WAL plus history plus summary dependencies before deleting anything.