Skip to content
fomoxx.

Tencent Cloud Lighthouse Has No Automatic Snapshots: My Two-Slot Rotation with systemd

中文版

Tencent Cloud Lighthouse Has No Automatic Snapshots: My Two-Slot Rotation with systemd

Tencent Cloud Lighthouse does not provide an automatic snapshot feature for the server workflow I was using, and my instance only has two snapshot slots.

If I rely on manual snapshots, I either consume the available recovery points quickly or simply forget to create them.

I ended up running a script on the Lighthouse server itself. The script calls the Tencent Cloud API, deletes the oldest automatic snapshot, creates a new one, and is triggered by a systemd timer every Sunday at 04:30.

I will put the final design first for readers who only want the working setup.

Quick solution

  • Use both snapshot slots for rotating automatic snapshots; do not reserve a fixed manual baseline.
  • Give managed snapshots the Auto-Snapshot- prefix. The script only deletes snapshots with that prefix, so differently named manual snapshots remain protected by default.
  • Run once per week at Sunday 04:30 with a systemd timer, not Cloudflare Workers and not traditional cron.
  • Set Persistent=false. If the server is not running at the scheduled time, skip that week’s run instead of creating a catch-up snapshot after reboot.
  • Run the script on the Lighthouse instance being backed up. If the server is damaged badly enough that the local script cannot run, an external scheduler cannot continue replacing older recovery points with new snapshots of a potentially damaged system.

Environment

ItemValue
Instancelhins-xxxxxxxx (replace with your own instance ID)
Regionna-siliconvalley (US West / Silicon Valley)
Script path/root/tencent-snapshot.sh
Automatic snapshot prefixAuto-Snapshot-
Number retained2

The important configuration at the top of the script is:

TENCENT_SECRET_ID=""
TENCENT_SECRET_KEY=""
TENCENT_REGION="na-siliconvalley"
INSTANCE_ID="lhins-xxxxxxxx"

SNAPSHOT_PREFIX="Auto-Snapshot-"
KEEP_SNAPSHOTS=2

The SecretId and SecretKey live only in the local script on this server. I do not sync them into cloud notes or version control.

Script permissions:

chmod 700 /root/tencent-snapshot.sh

CAM permissions

The snapshot script uses a dedicated CAM API credential with only snapshot-related permissions:

{
  "version": "2.0",
  "statement": [
    {
      "effect": "allow",
      "action": [
        "lighthouse:DescribeAllSnapshotsOverview",
        "lighthouse:DescribeSnapshots",
        "lighthouse:DescribeSnapshotsDeniedActions",
        "lighthouse:CreateInstanceSnapshot",
        "lighthouse:DeleteSnapshots"
      ],
      "resource": [
        "*"
      ]
    }
  ]
}

I hit one permission problem during testing.

The create-snapshot call returned:

UnauthorizedOperation.NoPermission

The CAM policy was missing lighthouse:CreateInstanceSnapshot. After adding that action, snapshot creation succeeded.

Snapshot rotation logic

Each run performs these steps in order:

  1. Run the optional health check.
  2. Call DescribeSnapshots for the current instance.
  3. If any snapshot is not in NORMAL state, stop without deleting anything.
  4. If two snapshots already exist, find the oldest one whose name starts with Auto-Snapshot-.
  5. Call DeleteSnapshots for that oldest managed snapshot.
  6. Wait until the old snapshot has actually disappeared.
  7. Run the health check again.
  8. Run sync to flush filesystem buffers as far as possible.
  9. Call CreateInstanceSnapshot.
  10. Wait until the new snapshot reaches NORMAL, then exit.

Because deletion is restricted to the managed prefix, manually named snapshots are protected by default and will not be rotated away.

With two slots, the snapshots look roughly like this:

Auto-Snapshot-20260906T203000Z
Auto-Snapshot-20260913T203000Z

The next weekly run deletes the older one and creates a new recovery point.

Health checks

The script supports PRECHECK_COMMAND. Leaving it empty skips the check:

PRECHECK_COMMAND=""

You can replace it with something simple, such as confirming that the filesystem is writable:

PRECHECK_COMMAND='touch /tmp/.snapshot-check && rm -f /tmp/.snapshot-check'

The design rule is: if the health check fails, stop the rotation. I do not want to delete a historical recovery point while the server is obviously unhealthy.

Immediately before creating the snapshot, the script also runs a fixed command to flush filesystem buffers:

PRE_SNAPSHOT_COMMAND="sync"

systemd service

File:

/etc/systemd/system/tencent-snapshot.service

Configuration:

[Unit]
Description=Tencent Lighthouse Automatic Snapshot
After=network-online.target
Wants=network-online.target

[Service]
Type=oneshot
ExecStart=/root/tencent-snapshot.sh
TimeoutStartSec=30min

This service runs the script once and exits. It does not stay resident in the background.

systemd timer

File:

/etc/systemd/system/tencent-snapshot.timer

Configuration:

[Unit]
Description=Run Tencent Lighthouse Snapshot Weekly

[Timer]
OnCalendar=Sun *-*-* 04:30:00
AccuracySec=1min
Persistent=false

[Install]
WantedBy=timers.target

Persistent=false is intentional.

If the server is powered off or broken at Sunday 04:30, the job is skipped. When the machine later recovers, systemd does not immediately create a catch-up snapshot. The existing recovery points remain in place until the next normal weekly rotation.

That is the opposite of how many backup timers are configured, and it is a deliberate trade-off for this two-slot design.

Enable the timer

systemctl daemon-reload
systemctl enable --now tencent-snapshot.timer

Confirm it is enabled:

systemctl is-enabled tencent-snapshot.timer

Expected output:

enabled

Confirm it is waiting:

systemctl is-active tencent-snapshot.timer

Expected output:

active

Check the next scheduled run:

systemctl list-timers --all | grep tencent-snapshot

Manual testing

Check the Bash syntax first:

bash -n /root/tencent-snapshot.sh

No output means the syntax check passed.

Then run the script directly:

/root/tencent-snapshot.sh

You can also test the same path the timer will use:

systemctl start tencent-snapshot.service

Check the service result:

systemctl status tencent-snapshot.service

A successful Type=oneshot service becomes inactive (dead) after it finishes. That is normal. The important part is the exit result: 0/SUCCESS.

Logs and troubleshooting

Last 100 lines:

journalctl -u tencent-snapshot.service -n 100 --no-pager

Last seven days:

journalctl -u tencent-snapshot.service --since "7 days ago"

Follow live:

journalctl -u tencent-snapshot.service -f

A normal rotation progresses through logs roughly like this:

Start processing instance
Health check passed
Query current snapshots
Delete oldest automatic snapshot
Wait for deletion to finish
Run sync
Create new snapshot
Wait for snapshot to enter NORMAL
Snapshot rotation complete

Why I schedule it locally instead of using Cloudflare Workers

My first idea was to put the schedule in Cloudflare Workers.

The obvious advantage is that if the server itself fails, the backup scheduler remains alive.

With only two snapshot slots, that turns into a different risk.

Imagine the server’s data is already damaged, but Tencent Cloud’s control plane can still create a system snapshot. An external scheduler keeps running on schedule.

Each run deletes one old snapshot and writes a new snapshot of the already-damaged state.

After two rotations, both old recovery points can be gone.

Running the scheduler on the server itself creates a fail-safe behavior instead: if the server is damaged badly enough that the local script cannot run, snapshot rotation stops naturally and the two existing recovery points are no longer overwritten by automation.

Persistent=false supports the same goal. Downtime does not cause an automatic catch-up run after recovery.

Why weekly instead of daily

This is another consequence of having only two slots.

The more frequently I rotate, the shorter the time distance between the two remaining recovery points.

A weekly rotation normally leaves something close to “this week” and “last week.”

That is more useful for my purpose, which is disaster recovery rather than fine-grained historical versioning.

Important limitations

  • Tencent Cloud snapshots are system-disk recovery points. They should not be treated as the only backup for important data.
  • Important databases and application data should still have independent data-level backups, such as database dumps, configuration archives, or object-storage copies.
  • If you change SNAPSHOT_PREFIX, make sure the existing snapshots you want managed use the same prefix. Otherwise the script treats them as manual snapshots and refuses to delete them.
  • Never put the Tencent Cloud SecretKey into cloud notes, Git repositories, or logs.

FAQ

Does Tencent Cloud Lighthouse support automatic server snapshots?

Not for the Lighthouse server snapshots covered here. My workaround runs a script on the server itself, calls the Lighthouse API, deletes the oldest managed snapshot, and creates a new one on a weekly systemd timer.

Why not schedule the two-slot rotation from Cloudflare Workers or another external scheduler?

With only two slots, an external scheduler can keep running even after the server’s data has already been corrupted. If the cloud control plane can still create snapshots, each run can delete an older good recovery point and replace it with a snapshot of the damaged state. Running the scheduler locally makes failure stop the rotation.

Why is Persistent=false set on the systemd timer?

It is deliberate. If the server is offline or unhealthy at Sunday 04:30, the timer does not create a catch-up snapshot immediately after recovery. Existing recovery points remain untouched until the next normal rotation.

Why rotate weekly instead of daily?

Because there are only two snapshot slots. More frequent rotation shortens the time distance between the two recovery points. Weekly rotation usually leaves a current-week and previous-week point, which better matches my disaster-recovery goal.

What should I check for UnauthorizedOperation.NoPermission when creating a snapshot?

Check whether the CAM policy contains lighthouse:CreateInstanceSnapshot. My first test failed because that action was missing; adding it allowed snapshot creation to succeed.

Continue reading

Comments