TESTUDO / COMMUNITY

Open-source Community Edition · Commercial support

Commercial · Open source

Cloud-Native Backup and Disaster Recovery Platform

Backup, recovery, migration, and business continuity for Kubernetes workloads.

Testudo is developed and maintained by SOFTC, with an open-source Community Edition and commercial support.

01
Why it matters

Move from having backups to verified recovery

Cloud-native applications combine resources, data, and dependencies; protecting one element does not prove recoverability. Testudo brings protection, replication, validation, and recovery into one consistent workflow.

Core value

Turn disaster recovery from backup jobs into verifiable business continuity

01

Application consistency

Protect resources, data, and recovery order around application boundaries.

02

Verified recovery

Validate protection policies through recurring drills and recovery records.

03

Open and controlled

Combine the open-source Testudo technical foundation with SOFTC professional services.

Capabilities

Eight capabilities built around the Kubernetes application DR loop

01

Application-level protection

Use DisasterInstance to define source and target clusters, namespaces, synchronization, and recovery policies.

Create a single source of truth for application protection.
02

Persistent data synchronization

Use AppBackup, Velero Backup, and AppRestore to recover persistent data in the target cluster.

Keep target-side data continuously recoverable.
03

Kubernetes resource synchronization

Synchronize application resource skeletons and keep the target in standby through modifiers.

Protect resources and data on one cadence.
04

Observable failover

Run pre-check, pause schedules, final sync, source scale-down, target scale-up, and role switching.

Turn failover into an executable workflow.
05

Reprotection

Confirm new roles, resume schedules, and establish reverse protection.

Restore continuous protection after cutover.
06

Multi-application disaster groups

Use levels for ordered execution, same-level parallelism, timeouts, retries, and failure policy.

Recover complex services in dependency order.
07

Non-disruptive recovery drills

Run instance or group drills using standby resources or isolated drill namespaces.

Continuously validate the full protection chain.
08

Status and event observability

Expose status, conditions, history, Kubernetes Events, Watch, and statistics APIs.

Keep protection, failover, and drills traceable.

Architecture

A five-layer control system with CRDs as the source of truth and Operators as executors

  1. 01

    User entry layer

    Manage disaster recovery objects through web consoles, CLI, automation, or APIs.

    • Web console
    • CLI
    • Automation platforms
  2. 02

    API aggregation layer

    disaster-server provides authentication, REST, Watch, event, and statistics endpoints.

    • Authentication
    • REST / Watch
    • Events and statistics
  3. 03

    CRD source of truth

    Kubernetes API stores desired state, runtime status, conditions, history, and events.

    • DisasterInstance
    • DisasterOperation
    • DisasterGroup
  4. 04

    Operator execution layer

    disaster-operator reconciles synchronization, failover, reprotection, and drill state machines.

    • DataSync
    • ResourceSync
    • DisasterDrill
  5. 05

    Runtime dependency layer

    Connect source and target clusters and use Velero with compatible object storage for backup and restore.

    • Kubernetes clusters
    • Velero
    • S3 / MinIO

DR workflow

A standard path from continuous protection to repeatable drills

  1. 01Protect

    Protect

    Define application, resource, data, and synchronization policies.

  2. 02Sync

    Sync

    Continuously synchronize Kubernetes resources and persistent data.

  3. 03Failover

    Failover

    Run pre-check, final sync, and role switching as observable steps.

  4. 04Reprotect

    Reprotect

    Establish the new primary-standby relationship after failover.

  5. 05Drill

    Drill

    Validate recovery without affecting production workloads.

Key scenarios

Key scenarios spanning migration, protection, failover, and verification

01

Cross-cluster application migration

Move Kubernetes resources and persistent data to a target cluster with complete status and history.

02

Metro or remote disaster recovery

Continuously synchronize applications across source and target clusters and execute controlled failover.

03

Continuous recovery drills

Use standby resources or isolated drill namespaces to validate resource, data, and application recovery.

04

Layered multi-application recovery

Orchestrate multiple instances by business dependency with same-level parallelism and cross-level sequencing.

Open source

Testudo(玄龟阵)

Explore documentation, source code, and contribution paths. The Community Edition uses Apache 2.0 with project supplemental terms; review the full repository license before use, modification, or distribution.

Apache 2.0 and project supplemental terms

Implementation

Four stages from workload assessment to continuous operations

  1. 01Assess

    Inventory workloads, data dependencies, and recovery objectives.

  2. 02Protect

    Establish tiered policies and automated backups.

  3. 03Verify

    Validate data and application consistency through recovery drills.

  4. 04Operate

    Continuously audit, optimize, and update the continuity baseline.