Quick answer
When an AI feature breaks production, root cause analysis must join application errors (GlitchTip/Sentry-class), deploy and CI state, customer tickets, latency and queue metrics, and feature-flag exposure—not only “the model got worse.” A spike in OAuth 500s can share a timeline with a failed integration test on main, a high-priority SSO ticket, and a flag flip that raised timeout rates. Cross-system RCA is either a disciplined playbook or a startup copilot with real connectors; tab hopping is not a strategy. Practice the narrative in CorpIM (Engineering loop + Copilot prompt style: “Root cause for OAuth 500 spike”) before the next real page.
Key takeaways
- Open one incident channel with a commander; a single timeline beats scattered threads.
- Correlate error fingerprints with deploy tags, prompt versions, and flag percentages.
- Read Outline-class runbooks before guessing that prompts need a rewrite.
- Publish status updates when user-visible API p95 breaches the SLA you sell.
- Postmortem evidence feeds SOC 2 change-management and availability controls.
Who this is for
- On-call engineers and eng leads for AI SaaS with async workers and external model APIs.
- Founders who currently RCA by opening Sentry, GitHub Actions, and Zendesk in parallel.
- SRE-minded PMs who need customer-impact context before approving hotfixes.
- Teams that ship agents with tools and need failure modes beyond HTTP status codes.
Who should skip
- Teams with no production users yet—invest in staging evals and load tests first.
- Pure research prototypes without customer SLAs or support queues.
- Readers who only need model/trace instrumentation—start with LLM observability.
- Readers debugging adoption softness without an outage—use AI feature adoption metrics.
What “AI feature breaks production” usually looks like
Customers rarely say “your LLM degraded on SWE-bench.” They say the product is slow, wrong, stuck, or broken. Desk synthesis of AI SaaS incidents clusters into a few shapes:
- Hard errors — 5xx, auth failures, tool exceptions, context overflow 400s.
- Soft failures — HTTP 200 with unusable outputs; success funnel drops while error dashboards stay quiet.
- Latency failures — p95 blows the SLA; users abandon; support hears “hung” before GlitchTip groups spike.
- Cascade failures — provider rate limits → queue depth critical → worker restarts → more retries → worse limits.
- Partial cohort failures — only flag-exposed tenants or one region hurt; global averages look fine.
RCA that only inspects model quality will miss OAuth regressions, bad deploys, stale embeddings, and permission bugs. RCA that only inspects app errors will miss “successful” bad answers. You need both planes, plus product outcomes when the outage is quiet.
First 60 minutes — decision table
| Minute | Action | Decision unlocked | Tool class |
|---|---|---|---|
| 0–5 | Declare incident; open IM channel; assign commander and scribe | Who owns rollback vs who digs | CorpIM / Slack |
| 5–15 | Triage error grouping—one route or systemic? | App bug vs platform-wide | GlitchTip |
| 15–25 | Deploy diff: last green CI, merges, flag %, prompt_version | Rollback candidate | Gitea/GitHub, Drone/CI, Unleash |
| 25–35 | Customer signal: tickets with same error string or “stuck generating” | Blast radius and severity | Chatwoot / Zendesk-class |
| 35–45 | Traces: DB, queue, upstream LLM timeout, tool calls | Which dependency failed | SigNoz / Grafana |
| 45–50 | Communicate: status page + support macro | Trust while you fix | Upptime, Outline |
| 50–60 | Fix path: flag rollback or revert → preview → staged re-roll | Restore SLA before elegance | Flags, CI, preview env |
If the “AI feature” is an agent with tools, keep agent failure modes and observability open beside this timeline—tool permission bugs look like “model success” with no side effect.
Declare severity with customer impact, not vibes
Before deep debugging, write three lines in the incident channel:
- Symptom — what users see (timeout, wrong auth, empty generation).
- Scope — which plans, regions, flag cohorts, surfaces (UI vs API).
- Severity — SEV based on paying impact and workaround availability, not on how embarrassed the team feels.
Example: “Summarize endpoint p95 > 45s for Growth tenants on flag model_v2 at 25%; paid API customers blocked; workaround = disable summarize.” That sentence prevents half the team from rewriting prompts while the other half ignores the only broken cohort.
AI-specific failure modes (not always “bad model”)
| Failure | Looks like | First check | Fast lever |
|---|---|---|---|
| Upstream model timeout | Hung UI, sparse app errors | Worker logs; provider status; queue age | Fail fast + retry policy; scale workers; temporary model pin |
| Context overflow | Sudden 400s after prompt/template change | Deploy diff on prompt templates; max token config | Roll back prompt_version |
| Rate limit cascade | Queue critical (demo-style OBS-441) | Queue depth; concurrency; shared API keys | Shed load; backoff; isolate tenants |
| Bad rollout | Errors or funnel drop at flag % | Unleash cohort vs error rate vs control | Decrease flag to 0%; pinned rollback |
| Tool permission bug | Agent “success” without side effect | Agent trace: tool name + status | Disable tool path via flag |
| Embedding index stale | Wrong retrieval, not timeout | RAG pipeline version tag; index build time | Pin prior index; pause reindex job |
| Auth/integration regression | OAuth 500s near AI features | CI auth tests; recent merges | Revert deploy; not model swap |
| Soft quality regression | 200 OK, success events down | Adoption funnel on exposed cohort | Flag rollback; eval offline after restore |
Notice the pattern: the fastest lever is usually a flag pin or deploy revert, not a mid-incident fine-tune. Training and prompt research belong after the SLA is restored.
Build one incident story (IDs that must match)
Cross-system RCA fails when every tool has a private vocabulary. Standardize a small set of IDs that appear everywhere:
incident_id— human channel name and status page IDrequest_id/trace_id— from edge to worker to model clienttenant_id— for blast radius and support macrosdeploy_tag/ git SHA — for rollback candidatesflag_key+ percentage + variantmodel_id+prompt_version- error fingerprint (GlitchTip group)
- ticket IDs that cite the same fingerprint or user-visible string
CorpIM-style demo narrative makes the point concrete: ERR-8821 (OAuth 500) ↔ CI-8821 (auth tests) ↔ CW-4092 (SSO ticket). Whether you use a copilot or a human commander, the job is the same—join those IDs into one paragraph a stranger can audit tomorrow.
Observability prerequisites (what you need before RCA works)
You cannot root-cause what you cannot see. Minimum planes:
- Errors — grouped exceptions with release tags
- Metrics/traces — p95, queue depth, upstream latency
- Product outcomes — success funnel for the impacted feature
- Change stream — deploys, flag changes, prompt publishes
How to wire those planes without proprietary lock-in is covered in open-source observability for LLM apps. This article assumes they exist. If they do not, your “RCA” will be opinion plus screenshots.
Logging discipline still matters under stress: keep tenant_id, model_id, tool status, and latency; redact raw prompts and PII unless policy explicitly allows short-retention incident captures.
Rollback decision table
| Evidence | Prefer | Avoid during active SEV |
|---|---|---|
| Errors track flag % only | Flag to previous model/prompt pin | Full redeploy of unrelated services |
| Errors track deploy SHA across features | Revert deploy / rollback release | Prompt fiddling in production console |
| Provider outage confirmed | Failover model or degrade feature; status update | Blaming your last prompt merge |
| Queue depth critical, app healthy | Scale workers; shed non-critical jobs | Rewriting UX copy |
| 200 OK + success funnel down on exposed cohort | Flag rollback; schedule eval after green | Retrain mid-incident |
| Auth/CI failures correlate | Revert auth-related merge | “Maybe the LLM hallucinated tokens” |
Rule of thumb: restore first, learn second. Elegant root cause writeups written while customers are blocked are vanity postmortems.
Rollbacks should leave an audit trail—who changed the flag, to what pin, at what time. That trail is change-management evidence for SOC 2 for AI SaaS, not paperwork theater.
Customer communication without making it worse
Status updates need a template, not poetry:
- What is impacted (feature + plans if known)
- What you know / do not know yet
- What customers can do (workaround)
- Next update time
Support macros should reuse the same language so Chatwoot replies do not contradict the status page. Engineers should not invent new theories in customer-visible channels. If you use Upptime-class pages, draft the incident note at minute 45 even if the fix is already in flight—silence reads as denial.
Production incident playbook (Startup OS shape)
In CorpIM Guide, a Production incident playbook typically sequences: acknowledge → triage → runbook → customer comms → fix → postmortem → GRC evidence. Map steps to tools you actually run:
- Acknowledge — IM channel + commander
- Triage — GlitchTip + metrics
- Runbook — Outline page for the feature (timeouts, flag pins, degrade switches)
- Comms — Upptime + support macro
- Fix — PR + preview + staged rollout (see gradual rollout)
- Postmortem — timeline, impact, root cause, contributing factors, owners
- Evidence — store for auditors reviewing CC7-class change and availability controls
If you lack a written runbook for “LLM timeout cascade,” write one after the incident while memory is fresh. The next page will use it.
Soft failures: when GlitchTip is green and the product is dead
This is the AI-specific trap. Model returns fluent nonsense; HTTP is 200; error rate is flat; D7 and success events fall for the exposed cohort. Treat that as a production incident of a different class:
- Confirm scope via adoption funnel on the flag cohort.
- Freeze further flag expansion immediately.
- Rollback to the last known good
model_id/prompt_version. - Only after restore, sample failures into an eval set and run offline comparison.
Do not “gather more production signal” by pushing to 100%. That is how you convert a canary lesson into a brand incident.
Copilot vs manual RCA
Manual tab hopping works for senior engineers who already know where every ID lives. It fails when the on-call is junior, when three systems disagree on timestamps, or when the incident spans auth + queue + model. A startup copilot helps only if connectors are real: same incident ID across errors, deploys, and tickets. Without connectors, copilots invent tidy stories. For the product distinction—and when chatbot UX is not enough—see copilot vs chatbot for startups.
Whether human or assisted, require the output format: timeline → evidence links → leading hypothesis → rollback action → verification metric. Reject narratives that skip verification (“we think it was the model”).
Post-incident: prove the fix and watch adoption
After the SEV channel closes:
- Verify error rate and p95 returned to baseline on the impacted cohort.
- Check success funnel and short-window retention for exposed users—incidents can be “over” in ops while habit is broken.
- File action items with owners and due dates; separate “prevent recurrence” from “nice refactors.”
- If the trigger was a model/prompt change, update the rollout ladder gates so the same change cannot skip canary metrics.
- Attach the postmortem to compliance evidence if you sell enterprise trust.
Re-entry to production should use gradual exposure again—not a pride-driven 100% flip because “we fixed it.”
Common mistakes (failure table)
| Mistake | Why it hurts | Fix |
|---|---|---|
| Debugging without a commander | Duplicate work and contradictory rollbacks | Assign commander + scribe in first five minutes |
| Blaming the model first | Misses auth, deploy, queue, flag bugs | Check change stream and infra before prompt surgery |
| No flag/deploy on error events | Cannot correlate | Require release + flag properties on errors |
| Global averages during partial rollout | Hides cohort outage | Slice by flag exposure and plan |
| Silent status page | Support and Twitter invent the narrative | Update on a clock even with incomplete cause |
| Fix without verification metric | Declare victory while users still fail | Define “green” as error + p95 + success funnel |
| Retrain / huge prompt edit mid-SEV | Adds variables under fire | Pin known good; experiment after restore |
| Postmortem theater | No owners, no evidence, repeat outage | Action items + SOC-oriented evidence pack |
| Ignoring soft failures | Brand damage without red dashboards | Watch adoption success on exposed cohorts |
| Skipping agent tool traces | False “model OK” stories | Log tool name + status in traces |
Worked narrative (desk synthesis)
Paid tenants report “Generate brief” hangs. Error rate is only mildly up. On-call opens five tabs and argues about prompts. Commander forces the timeline: flag brief_model_b moved from 5% to 25% ninety minutes ago; queue depth for brief workers climbed; provider rate-limit headers appear in worker logs; control cohort at 0% flag is healthy.
Action: flag → 0%, status note published, support macro sent. p95 recovers in twelve minutes. Postmortem root cause: concurrency assumed provider headroom that disappeared after a quiet provider-side limit change; contributing factor: no alert on queue age for the brief worker. Action items: alert on queue age, lower default concurrency, keep model B behind 5% until headroom proven. Adoption check next Monday: exposed cohort D7 dipped slightly; email to affected tenants with credit policy. No one fine-tuned anything during the SEV.
That is successful RCA: restore, explain, prevent—not a model bake-off under fire.
How this ties to rollout and adoption
RCA is the emergency twin of gradual rollout. The rollout ladder defines gates; RCA fires when a gate was skipped or a signal was ignored. After incidents, feed lessons back into flag thresholds and Wednesday release reviews described in gradual rollout for models and prompts.
Adoption metrics tell you whether the outage scarred habit. If success events fell without a hard error spike, your observability planes were incomplete—or you were not watching the product plane at all.
Roles on the incident channel
Small teams skip role labels and pay for it with contradictory actions. Even a five-person startup should name:
- Commander — decides rollback vs dig; only voice that authorizes customer-visible status text.
- Ops digger — owns errors, traces, queues, provider status.
- Change digger — owns deploy diff, flag history, prompt publish log.
- Customer voice — owns tickets, macros, and “who is hurt” counts.
- Scribe — keeps the timeline; becomes the postmortem skeleton.
One person may wear two hats; nobody should wear all five while also rewriting prompts live. If a founder jumps in with a new hypothesis every three minutes, the commander parks it on a “parking lot” list until the SLA is green.
Evidence pack for the postmortem
A usable postmortem is a folder of proof, not a blog post. Capture:
- Channel export or timeline notes with UTC timestamps
- Screenshots or links to error group, queue chart, flag history
- Deploy SHAs and who merged them
- Status page revisions and support macro versions
- Verification charts showing return to baseline
- Action items with owners and due dates
Store that pack where GRC/Probo-class evidence already lives if you sell enterprise. Auditors do not need your Slack jokes; they need controlled change and availability narrative. If you are early-stage, a markdown file in the repo still beats tribal memory.
Severity examples (tune to your SLA)
Write severity definitions before the next page so arguments do not start from zero:
- SEV-1 — paying customers cannot complete a core AI job; no workaround; widespread.
- SEV-2 — core AI job degraded (latency/quality) for a material paid cohort; workaround exists.
- SEV-3 — non-core AI surface broken; or only free/trial cohorts hurt.
- SEV-4 — cosmetic or internal-only; track as a defect, not an incident channel.
Soft failures can still be SEV-1 if enterprise workflows depend on correct outputs. Do not require red HTTP charts to take severity seriously.
Provider and dependency coordination
When the upstream model API is degraded, your RCA still has work: prove it, communicate it, decide degrade vs failover, and protect your queues from retry storms. Open a vendor status link in the channel, but do not stop correlating—your SDK timeout settings and concurrency caps are still your code. If you multi-home models, document which pin is the failover and test it in staging; an untested failover is fiction.
Dependency failures also include embedding stores, object storage for attachments, and OAuth IdPs. AI features sit on classical SaaS dependencies; treat them as first-class in the first-hour table.
CorpIM CTA
Practice the Engineering loop story line in CorpIM: error groups, CI failures, support tickets, and copilot RCA prompts in one IM surface. Rehearse before production teaches you the expensive way.
Open CorpIM → https://www.romewayai.com/corp-im/
Decision depth: isolate plane before you blame the model
Root cause collapses when every engineer opens a different pane. Force a single isolation order in the first hour—write it on the incident channel so people stop parallel guessing:
- Edge / gateway — Auth failures, rate limits, and request size rejects show up as customer “AI broken” while the model never ran. Check gateway 4xx/5xx and auth error codes before touching prompts.
- Queue / worker — Backlog growth with healthy HTTP on the API means users see stale or empty jobs. If queue depth and age are the smoking gun, scale or shed load; do not swap model pins.
- Retrieval / tools — Empty retrieval, tool timeout, or ACL denials produce confident-looking failures. Compare retrieval-empty rate and tool-error rate on the failing cohort versus a healthy window.
- Model / prompt pin — Only after the three planes above are clean (or ruled out with evidence) do you compare flag cohorts and pin versions. This is where gradual rollout history matters: if the symptom tracks flag %, pin the previous revision.
- Product / UX contract — UI that swallows errors into blank states will look like “soft AI failure.” Confirm the client surfaces error codes before you declare a quality SEV.
If two planes are red, fix the one that restores SLA fastest, then continue RCA—do not wait for a perfect postmortem diagram while customers are blocked. Pair this order with your three-plane observability map so the channel links to the same IDs every time.
Worked example: green errors, red adoption
Illustrative desk narrative: Tuesday 14:10 UTC, enterprise tenant reports “assistant returns nothing useful.” GlitchTip is quiet. SigNoz-class p95 on the chat route is normal. PostHog shows ai_task_success down ~30% for that tenant since 13:40, while control tenants are flat.
First-hour findings: a retrieval worker deploy at 13:38 truncated vector search results to zero for one namespace after a migration flag flipped only for “beta embeddings.” HTTP stayed 200 because the chat handler returned an empty citations array and a generic “I don’t know” completion. Model pin never changed.
Correct sequence: (1) SEV-2 for paid cohort with workaround (manual search); (2) roll back the embeddings flag, not the LLM pin; (3) add an alert on retrieval-empty rate joined to tenant; (4) postmortem evidence pack includes migration flag ID, empty-retrieval spike, and the adoption drop—so next time nobody opens the model playground first. Practice the same Engineering story line in CorpIM before the next real page.
FAQ
Should we rollback the model or the deploy first?
Rollback the fastest lever that restores SLA. If symptoms track flag percentage, pin the previous model/prompt. If symptoms track a SHA across routes, revert the deploy. Do not retrain or redesign prompts during an active SEV.
When is model quality actually the root cause?
When HTTP is healthy, infra and deploys are clean, and the exposed cohort’s success funnel drops while control stays flat. Restore via flag first, then run offline eval—not live prompt roulette.
How does RCA tie to adoption metrics?
After the incident, check AI cohort retention and success for exposed users. Latency spikes can kill D7 even when GlitchTip looks calm.
What evidence do enterprise buyers want?
Incident timeline, customer comms log, postmortem, and change control on the fix PR—overlapping the SOC 2 checklist for AI SaaS.
Do we need a copilot to do cross-system RCA?
No. You need shared IDs and a playbook. A copilot accelerates correlation when connectors are wired; without them it invents stories. See copilot vs chatbot.
Continue the semantic path
Instrument the planes with open-source LLM observability, prevent repeats with gradual rollout, and watch scars in adoption metrics.
LLM observability · Copilot vs chatbot · Gradual rollout · SOC 2 checklist · Agent failure modes · Measure adoption