Everything here is advisory and read-only. Nothing in this guide changes your workloads (the sample incident happens in a scratch namespace you create and delete). ChangeGuard AI does not act on your cluster until you explicitly opt into remediation.
1
Install ChangeGuard AI
One command installs the operator and the read-only collector — the install page walks it through. Already installed? Carry straight on.
2
Connect your cluster
Confirm the operator and collector are healthy and your cluster appears in Fleet with green connection indicators (“Agent connected”, “Workloads detected”) — the checks are on Connect your first cluster. If anything doesn’t match, the fix table there covers it — it’s almost always egress or the API key.
3
See your first score
Your cluster’s CSC Score appears in Fleet within ~10 seconds of connecting — a deterministic 0–100 read of deploy readiness. Every point ties to a concrete signal you can inspect: Understanding your score.
4
Run a pre-flight check — and meet the Advisor
Open Safe to Ship? and run a pre-flight check. You get SHIP / HOLD / BLOCK with the score and reasons — advisory, never an enforced block. Where the Engineering Advisor is enabled (Early Access), its note appears alongside the verdict: the read a senior engineer would give.
Two things that are correct, not broken: if the Advisor has nothing material to add it stays silent — restraint is the feature. And if you don’t see the Advisor at all, that’s expected — it’s Early Access, off by default.
5
Trigger a sample incident, safely
Watch a failure get correlated to the change that caused it — in a scratch namespace that touches nothing real:The new pods can’t pull the image, the rollout starts failing, and ChangeGuard AI sees both the change and the failure. Within a few minutes, an incident appears for
sample-api — linked to the image change you just made. Cause, not just symptom.6
Understand the recommendation
Open the incident. The root cause comes with cited evidence — the events, the failing rollout, and the change that shipped it. ChangeGuard AI proposes the fix and states the criteria to verify it worked. At the default autonomy level it applies nothing — this is advice you act on. Where the Engineering Opinion is enabled (Early Access), you’ll also see an owned position: belief, confidence, tradeoffs, and what would change its mind.
7
View the Current Understanding
Where incident context is enabled, the incident opens with the Current Understanding — the maintained read of what’s happening right now: what changed, what it broke, and where things stand. It’s what a teammate joining mid-incident needs, without scrolling the history. Every incident also carries an append-only activity timeline — your audit trail of what was observed and decided.
8
See the loop close
Ship the fix and watch ChangeGuard AI observe the recovery:The rollout recovers, the incident reflects it, and the timeline records the whole arc — change, failure, fix, recovery. You applied this fix yourself; if you later opt into Autonomous Remediation, ChangeGuard AI can execute fixes within the policy you set, verify them against real workload health, and roll back once if verification fails.
Set your comfort level
ChangeGuard AI’s autonomy is a dial, not a switch: Observe → Advise → Approve → Auto, default Advise.You’ve now seen the loop
In 30 minutes: a verified install, a scored cluster, a pre-flight verdict, a failure correlated to its cause, an evidence-cited investigation, and the recovery on the timeline — plus, where enabled, the Understanding, Opinion, and Advisor that turn it into a teammate.Understand the concepts
What CSC Score, Change Intelligence, and the autonomy model really mean.
Operate it day-to-day
Health checks, upgrades, notifications, and keeping the collector happy.
Wire in your tools
GitOps, CI gates, and notifications — what each reads, writes, and needs.
Review permissions
Exactly what ChangeGuard AI can and cannot see or do in your cluster.