Best Practices

Microsoft 365 Backup Operations Checklist

Published Jul 25, 20267 min readBy Security Engineering

Direct answer: Good Microsoft 365 backup operations are not defined by a green job status. They require an approved object-level scope, least-privilege authorization, a named monitor, a known usable recovery point, a tested restore, documented fidelity, and an exit procedure. Use this page as an operating checklist after business owners approve the backup and recovery policy.

Operating Standard

StageDecisionMinimum evidenceFailure condition
PlanWhich exact objects and scenarios require recovery?Approved scope register and acceptance criteria.Only workload names or generic "all data" language.
AuthorizeWhich application and human roles are required?Permission set, consent record, credential owner, access review.Unowned credentials or broad roles without review.
ProtectWhat establishes a usable recovery point?Completed job, item errors, snapshot identifier, retention state.Partial job recorded as fully successful.
MonitorWho acts on stale points, failures, and throttling?Named owner, review cadence, ticket and escalation trail.Console state is available but nobody is responsible for it.
RecoverCan the required object be restored to the required destination?Timed test with content and metadata acceptance.Backup readability assumed to equal write-back fidelity.
RetireWhat happens to access, retained data, billing, and evidence?Offboarding approval, deletion/retention decision, final record.Provider or tenant removed without a recovery-data decision.

1. Define Recoverable Scope

  1. [ ] List tenants and workloads: Include Exchange, OneDrive, SharePoint, and Teams only where there is an approved business scenario.
  2. [ ] List object types: Distinguish messages, mail folders, contacts, calendars, files, folders, lists, list items, pages, permissions, members, chats, and channels.
  3. [ ] Mark backup and restore separately: An object can be readable through an API but not faithfully restorable.
  4. [ ] Record exclusions: Assign an owner and accepted consequence to every unsupported object.
  5. [ ] Separate recovery from compliance: Purview retention supports preservation and discovery workflows; it is not automatically an operational restore. Check Microsoft's current retention documentation.

2. Verify Access Before Production

Background collection commonly uses application permissions and administrator consent, but the exact permission set is handler specific. Do not copy a static permission list from an undated article. Record the currently requested permissions, why each is needed, the granting administrator, credential owner, rotation or renewal process, and revocation path.

Application identity[Application ID, owning organization, tenant]
Permissions[Permission, workload/handler, reason, consent type]
Credential handling[Provider/customer owner, secret or certificate process, review date]
Preflight outcome[Handler passed/failed, error, remediation, approver]

3. Monitor Recovery Points, Not Schedules

A configured schedule expresses intent. Effective protection depends on authorization, service availability, pagination, throttling, storage writes, retention processing, and item-level outcomes. Microsoft states that Graph throttling varies by request type and scenario; applications should honor Retry-After and use documented backoff. Review the official Graph throttling guidance. No provider can guarantee that Microsoft Graph will never throttle a tenant.

For every operational review, capture:

4. Evaluate Storage and Deletion Controls

Do not use "immutable," "air-gapped," or "separate" as substitutes for a control review. Microsoft's native Backup product documents append-only storage but also a controlled offboarding path; see the current Microsoft 365 Backup overview. Managed and customer-controlled products have different administrators, deletion paths, regions, credentials, and exit risks. Record those facts with the storage decision guide.

5. Run a Repeatable Restore Test

  1. Select a supported object that represents the approved scenario. Do not substitute a simple email for a required site or Teams scenario.
  2. Record request, approver, source recovery point, destination, and conflict behavior before starting.
  3. Time authorization, initiation, provider completion, and business acceptance separately.
  4. Verify content, hierarchy, identifiers, timestamps, permissions, links, and any unsupported metadata.
  5. Record item errors and changed fields. A partial restore is not a pass unless the acceptance criteria permit those differences.
  6. Calculate effective objectives with the RTO/RPO worksheet.

Restore Evidence Record

FieldEntry
Scenario and requirement[Business event, object, approved RPO/RTO, owner]
Source[Tenant, workload, recovery point, object ID, timestamp]
Destination[Original/alternate, conflict mode, authorization]
Timing[Declared, authorized, started, completed, accepted]
Fidelity[Content, metadata, hierarchy, permissions, identifiers]
Result[Pass, partial, fail; defect/exception owner and due date]

6. Reconcile Change and Offboarding

Repeat scope and access checks after license changes, user departures, new sites, renamed objects, permission changes, provider releases, or retention updates. Before removing a tenant or provider, decide who authorizes stopping new backups, how long existing data remains, whether any export exists, when billing stops, and what evidence remains. Do not promise a portable historical archive unless the provider demonstrates it.

North Brook Vault Coverage and Limits

North Brook Vault is managed SaaS with provider-managed infrastructure and storage. It provides tenant-scoped credentials, jobs, snapshots, schedules, retention policies, service-side monitoring, full or delta-aware execution where supported, and selective restore handlers for supported object types. Customers own tenant consent, customer-side operators, recovery requirements, acceptance, and review of tenant-level outcomes.

The service does not offer customer-selected S3 or on-premises deployment, native Object Lock enforcement, OneDrive version-history capture, PST/PDF/ZIP export, compliance reports, or a zero-throttling guarantee. Teams messages and channel structures are not restorable. Do not define a full-workload restore, compliance production, or guaranteed RPO/RTO in a service commitment unless a current written agreement and successful representative test support it.

Operational Decision

Move a tenant to production only when scope, access, latest usable point, monitoring ownership, representative restore, known limitations, and offboarding ownership are recorded. Keep failed checklist items visible as accepted exceptions or remediation work.

Review managed service operations Discuss an operations consultation Price the approved scope