A working cluster is the start of a platform, not proof of production readiness. A production review should demonstrate that the team can deploy safely, restrict access, detect failures and recover the application and its data.

Establish ownership and access

Name the platform owner and the application owner. Separate routine deployment permissions from cluster administration. Connect access to the organisation’s identity process so departures and role changes are reflected promptly.

Namespaces help organise workloads but are not a complete security boundary by themselves. Review service accounts, role bindings, network policy enforcement and secrets handling. Test whether one workload can reach another workload’s data or credentials, rather than assuming separation from naming alone.

Make releases repeatable

Build images in a controlled pipeline and identify exactly which image is deployed. Keep configuration changes reviewable. Promote tested versions through environments and define what constitutes a failed rollout.

Health probes should reflect the application’s real condition without causing a restart loop during a temporary dependency problem. Agree when a release rolls back automatically and when a person investigates. Rehearse the rollback using the same process the team will use during an incident.

Treat state as a separate concern

A pod can be replaced without restoring a lost database. Document storage classes, volume behaviour and backup coverage for each stateful component. Include configuration and secrets required for recovery, with suitable protection.

Restore a representative application into a clean environment. Record how long it takes and what data is missing. Compare those results with the business recovery objectives. A backup job marked successful is not the same as a demonstrated restore.

Observe the user experience

Collect application errors, latency and dependency failures as well as cluster metrics. Alert on conditions that require action, with a named recipient and a runbook. A healthy node does not prove that customers can log in or complete a transaction.

Set requests, limits and capacity plans from workload measurements. Include expected growth and maintenance headroom. For GPU workloads, verify device support, scheduling and isolation with the actual inference runtime.

Plan the next year

  • Document supported versions and an upgrade rehearsal process.
  • Track certificates, image dependencies and security patches.
  • Test node loss, dependency failure and exhausted capacity.
  • Define support coverage and escalation responsibilities.
  • Review costs and unused resources with application owners.

The strongest readiness evidence is a short set of successful drills and clear operating records. A long checklist with every box marked “yes” is much less useful if nobody can show how the service is restored.

Sources & further reading

Primary references for the technical background and regional statements in this guide. Planning examples and checklists are Novacom’s practical guidance; examples are illustrative unless explicitly identified as project experience.

FROM UNDERSTANDING TO A WORKING SYSTEM

Apply this to your organisation.

Bring your workflow, data boundaries and existing systems. We’ll help define a useful scope and the evidence needed to evaluate it.