Quick answer
A practical SOC 2 checklist for AI SaaS is not a separate “AI SOC” framework. It is Trust Services Criteria evidence plus AI-specific artifacts buyers now ask for: LLM subprocessors, training-use claims, prompt and agent logs, human review for high-risk outputs, and retention rules for customer content sent to inference APIs. Seed teams should start a lightweight control map and subprocessor list before the first enterprise RFP; Series A diligence often expects Type II in progress or a dated timeline. Automate evidence from CI (PR merges, Trivy/Semgrep-class scans) into a Probo-class GRC tool instead of Notion screenshots. CorpIM Compliance demos cap table + SOC gaps + IR in one loop—open the CorpIM Compliance tab.
Key takeaways
- Inventory every place customer content can go for inference, embeddings, support AI, and analytics—then keep DPAs and a public subprocessor list current.
- CC6.x access: SSO, MFA, least privilege on prod, separate model API keys per environment, and offboarding that revokes vector-index admin access.
- CC7.2 change management: required PR review + CI on
main; attach security scan artifacts as living evidence. - CC8.1 vendor reviews: treat model API providers, embedding hosts, and vector DBs like any other critical vendor—calendar annual reviews.
- Privacy (P1.x): decide prompt logging, redaction, retention, and deletion before auditors invent the question for you.
Who this is for
- Seed–Series A AI SaaS founders who just received a security questionnaire or will within two quarters.
- CTOs and eng leads who must map model vendors, embedding hosts, and agent tooling into subprocessors and access reviews.
- Ops or security owners automating evidence from Git/CI instead of assembling PDF scrapbooks before diligence.
- Product leads shipping copilots who need a clear line between “feature risk” and “company control evidence.”
- Founders aligning SOC narrative with monthly investor updates—see AI-assisted investor updates.
Who should skip
- Pre-revenue consumer apps with no B2B pipeline—do access hygiene and secrets management; full Type II theater is premature.
- SMB-only sellers whose buyers never ask for SOC—run a short security FAQ; revisit when ACV or verticals change.
- Readers who only need investor narrative structure—start with investor updates, not this checklist.
- Teams hunting a guaranteed “pass SOC in 30 days” vendor claim—no checklist replaces observation period and auditor judgment.
- Enterprises with mandated GRC programs—adapt AI nuances into existing controls; do not rip out governance for a blog stack.
What SOC 2 is (and is not) for AI products
SOC 2 is a report on controls relevant to Trust Services Criteria—commonly Security, and often Availability, Confidentiality, Processing Integrity, and Privacy depending on scope. The public overview lives at the AICPA SOC for Service Organizations page. Auditors do not “certify your model accuracy.” They evaluate whether you designed and operated controls that match your system description.
For AI SaaS, the system description gets longer: inference paths, third-party model APIs, retrieval stores, agent tool executors, eval environments, and logging that may contain customer content. Buyers treat that length as risk surface. Your job is to make the surface documented and controlled—not to pretend it is small.
What SOC 2 is not:
- Not a model quality score or safety eval leaderboard.
- Not legal advice, ISO 27001, HIPAA, or FedRAMP by itself.
- Not a substitute for product safety engineering—pair with practical AI safety for builders and jailbreak and prompt-injection risks.
- Not something you “finish” once; Type II is about operating effectiveness over time.
Desk synthesis pattern: teams that treat SOC as a PDF project fail questionnaire follow-ups. Teams that treat SOC as a control map wired to Git, CI, vendor calendar, and incident RCA survive diligence with less founder panic.
Decision table: when to start which depth
| Stage / trigger | Minimum depth | Why now | Defer |
|---|---|---|---|
| Pre-revenue, no B2B pilots | SSO/MFA habit, secrets hygiene, incident runbook stub, no shared prod keys in chat | Habits are cheaper than retrofit | Formal Type II observation, heavy GRC spend |
| First enterprise pilot / security questionnaire | Subprocessor list, DPA template, security FAQ, access offboarding checklist | RFPs arrive before you feel “ready” | Full auditor engagement if timeline allows lightweight answers first |
| Seed with ~$500k+ ACV pipeline or regulated vertical | Control map, evidence collection, gap remediation owners, CI→GRC attachments | Deals stall on “Type II timeline?” | ISO expansion until buyers ask |
| Series A diligence / multi-enterprise renewals | Type II in progress or report in data room; vendor reviews current | Investors mirror enterprise buyer questions | Vanity compliance badges without evidence links |
Use the table as a spending filter. Buying a GRC seat before you can produce a subprocessor list is theater. Waiting until a Fortune 500 legal team is on the email thread is how you lose a quarter.
Trust Services Criteria — AI SaaS reading guide
You do not need to memorize every criterion letter. You need a mapping from criteria families to AI-specific proof.
Security (CC series) — always in play
Logical access, change management, risk assessment, and vendor management dominate early questionnaires. For AI SaaS, “logical access” includes who can rotate production model keys, who can query production vector indexes, and who can approve agent tools that write to customer systems.
Availability (A series) — if you sell uptime
Model provider outages, queue backlogs, and GPU worker death are availability events. Buyers want runbooks and status communication, not “the model was slow.” Tie availability narrative to open-source observability for LLM apps and postmortems from production root-cause practice.
Confidentiality / Privacy — if you process sensitive content
Customer prompts, retrieved chunks, and tool outputs can be confidential. Retention, redaction, and deletion become concrete control statements. If you claim “we do not train on your data,” that claim must match vendor contracts and your own fine-tune pipelines.
Processing Integrity — when outputs drive decisions
If your product’s AI output triggers billing, medical-adjacent advice, or automated actions, buyers ask how you detect bad processing. Human review gates, eval gates before rollout, and audit logs for agent actions are the practical answers—not a promise of perfect models.
Core control checklist (AI SaaS starter map)
| Control area | What to prove | AI SaaS nuance | Typical evidence |
|---|---|---|---|
| CC6.1 Logical access | SSO, MFA, role-based prod access | Separate prod/staging model API keys; no shared “company OpenAI key” in a wiki | IdP config export, access review CSV |
| CC6.6 Offboarding | Revoke Git, cloud, SaaS within policy window | Revoke embedding index admin, agent tool credentials, eval dataset vaults | Closed offboarding ticket + IdP deprovision log |
| CC7.2 Change management | PR review + CI required on main | Prompt/version and model flag changes treated as releases | Branch protection settings, merge logs, CI green artifacts |
| CC7.3 Vuln management | Container/dependency scans on deploy path | Scan inference worker images and agent runners, not only the web API | Trivy/Semgrep-class JSON attached to control |
| CC8.1 Vendor management | Inventory + risk reviews on cadence | Model APIs, embedding hosts, vector DB, eval SaaS on the list | Vendor register, DPA folder, review dates |
| A1.x Availability | Incident response, status, backups as scoped | Model timeout / provider outage playbooks with customer messaging | Runbook links, status page history, postmortems |
| P1.x Privacy (if in scope) | Notices, retention, deletion, DPIA where required | Prompt log policy, training opt-out, RAG chunk retention | Policy docs, deletion tickets, DPIA for AI features |
Map rows into whatever GRC you use. The point is ownership and refresh, not perfect taxonomy on day one.
Subprocessor inventory — the AI SaaS failure point
Enterprise questionnaires often open with “list all subprocessors that may process personal data or customer content.” AI startups under-list because inference feels “just an API.” Treat every path as a candidate:
- Primary chat/completion model API
- Fallback or router providers
- Embedding API or self-hosted embedding workers (hosting provider still matters)
- Vector database host
- Speech-to-text / OCR vendors if in product path
- Support tools that summarize tickets with customer text
- Analytics that store event properties with free-text prompts
- Error trackers that may capture request payloads
For each entry record: purpose, data categories, region, DPA status, training-use stance, and owner. Publish a customer-facing list you can update without rewriting the whole security whitepaper.
Desk synthesis failure mode: marketing says “data never leaves our VPC” while the product calls a hosted model API. Auditors and sophisticated buyers reconcile claims against architecture diagrams. Align claims with reality before diligence week.
Access control for AI stacks (CC6.x in practice)
Classic SaaS access reviews miss AI-specific privilege:
- Key isolation. Production inference keys must not live in developer laptops or shared chat. Prefer secret managers and short-lived credentials where vendors support them.
- Environment separation. Staging should not share production vector indexes that contain customer documents.
- Tool ACLs for agents. An agent that can refund, email customers, or mutate CRM needs human approval gates and role checks—see operating patterns in AI-native Startup OS.
- Eval data vaults. Golden datasets often contain sensitive customer-like text. Restrict who can export them.
- Break-glass. Document who can temporarily elevate access during incidents and how it is reviewed afterward.
Offboarding is where AI stacks leak. A departed engineer with lingering admin on the vector DB or agent runner is a CC6.6 finding waiting to happen. Put embedding and agent credentials on the same offboarding checklist as Git and cloud.
Change management — prompts and models are releases
Many AI teams ship prompt edits from an admin UI with no PR. Buyers who care about change control will ask how you prevent silent regressions. Practical pattern:
- Version prompts and model IDs as config reviewed like code, or as flag changes with owners.
- Require CI checks for services that execute tools or touch billing.
- Attach container and SAST scan results to the release evidence pack.
- Correlate deploy tags with error fingerprints—observability is evidence when availability is in scope.
Gradual rollout is both product risk management and change hygiene: see gradual rollout for models and prompts. A flagged canary with monitors is easier to defend than “we hotfixed the system prompt on Friday night.”
Vendor management for model providers (CC8.1)
Annual vendor reviews feel bureaucratic until a provider changes retention defaults or a region outage cascades. For each critical AI vendor, keep:
- Contract / DPA location
- Security questionnaire or SOC report receipt date (if provided)
- Data residency options you actually use
- Training / logging stance you rely on in customer contracts
- Exit notes: how you migrate if pricing or policy shifts
Do not invent vendor “scores.” Record what you reviewed and when. If a vendor does not share a report, document compensating controls (encryption in transit, minimization, contractual clauses) and your residual risk acceptance owner.
Privacy and logging — prompts are data
Decide these before the questionnaire arrives:
- Do you store full prompts in application logs?
- Do error trackers capture request bodies?
- Do product analytics store free-text AI inputs as event properties?
- How long are traces retained, and can customers request deletion that covers retrieved chunks?
- Is there a training opt-out path that matches vendor settings and your own fine-tunes?
Recommended desk pattern for early B2B: default to metadata logging (request_id, tenant_id, model_id, latency, token counts, tool name/status) and keep full content behind explicit policy + retention. Detail the observability side in open-source observability for LLM apps.
If Privacy criteria are in scope, DPIA-style reviews for new AI features that process customer content are the concrete artifact—not a blog post titled “we take privacy seriously.”
Agent actions need audit trails
Buyers ask: “What did the AI do in our tenant?” If you run agents with tools, you need durable logs of tool name, parameters classification (not necessarily raw secrets), status, actor (user/system), and timestamp. That overlaps product observability and compliance evidence.
Without tool-call traces, incident RCA becomes speculation—and speculation is a diligence smell. Pair agent logging with production RCA discipline and the safety themes in prompt injection and agent risks.
AI-specific questionnaire themes (prepare answers, not slogans)
Expect variations of:
- Where is customer data sent for inference, and for how long is it retained by you and by vendors?
- Can customers opt out of model training on their data—and how is that enforced technically?
- How do you handle prompt injection in support tickets, uploaded docs, or retrieved content?
- Is there human review for high-risk outputs or irreversible actions?
- Do you maintain audit logs for agent tool calls?
- Is the subprocessor list complete for model API and vector DB hosts?
- How do you test for jailbreaks or policy bypass before major releases?
Answer with architecture + policy + evidence links. “We use a leading model provider” is not an answer. Cross-link product safety work to practical AI safety so security and product teams do not invent conflicting stories.
Evidence automation — CI into GRC
Auditors prefer fresh artifacts. Humans prefer not to screenshot every week. Automate where the path is clear:
- PR merge log → change management evidence
- CI scan JSON → vulnerability / malware control attachments
- Postmortem PDF from incident RCA → availability narrative
- Closed offboarding tickets → access removal proof
- Vendor review calendar tasks → CC8.1 cadence proof
CorpIM-class demos show attaching Trivy HIGH findings to a Probo-class control such as CC7.2. Whether you use that stack or another, the pattern is identical: systems produce evidence; GRC stores pointers; humans own exceptions.
Do not automate fake completeness. If a scan is skipped for an emergency hotfix, record the exception and the follow-up. Silent skips become Type II findings.
Type I vs Type II — seed-stage decision
| Report | What it shows | When it helps | Trap |
|---|---|---|---|
| Type I | Design of controls at a point in time | Early enterprise conversations while observation period starts | Treating Type I as “done forever” |
| Type II | Operating effectiveness over a period | Serious enterprise and many Series A data rooms | Starting observation with empty evidence folders |
Start the control map early so the Type II window is not three months of “we will document it later.” Buyers increasingly ask for Type II or a dated path to it. Honest timelines beat inflated claims in investor updates—see AI-assisted investor updates.
Decision table: GRC tooling vs spreadsheets
| Situation | Prefer | Avoid |
|---|---|---|
| First questionnaire, <10 employees | Structured spreadsheet + policy pack + owner column | Six-figure GRC before evidence exists |
| Repeated RFPs, CI already mature | GRC with CI attachments and task cadence | Notion screenshot folders as sole system of record |
| Multi-product, multi-region AI paths | GRC + clear system description diagrams | One vague “AI policy” PDF for everything |
| Investor diligence in <60 days | Gap list with owners + Type II timeline | Buying a badge without remediation plan |
90-day starter plan (desk pattern)
- Days 1–14: Draft system description (inference paths, data stores). Build subprocessor inventory. Turn on MFA everywhere that touches prod.
- Days 15–30: Write access offboarding checklist including AI privileges. Enforce PR + CI on production paths. Define prompt/log retention policy.
- Days 31–60: Attach CI scan artifacts to vulnerability controls. Schedule vendor reviews for model and vector hosts. Stand up incident runbooks for model outages.
- Days 61–90: Run a mock questionnaire. Close top gaps. Decide Type I/II timing with counsel/auditor. Mirror status into investor risk section if fundraising.
Parallel track: keep weekly operating rhythm so compliance work does not become a once-a-year panic—see AI startup weekly operating rhythm and the broader OS pattern in AI-native Startup OS.
Common SOC 2 mistakes for AI startups
| Mistake | Buyer / auditor question | Why it fails | Fix |
|---|---|---|---|
| No subprocessor list | “Who sees our data for inference?” | Trust collapses at first follow-up | Publish list; DPAs on file; owner for updates |
| Prompt logging unclear | “Do you store prompts?” | Conflicting answers across eng and sales | Written retention + redaction policy |
| Manual prod deploys / prompt hotfixes | “Change control?” | No auditable release path | CI-only deploy; version prompts/flags |
| No agent audit log | “What did the AI do?” | Cannot reconstruct actions | Tool-call traces with tenant IDs |
| Stale vendor reviews | “CC8.1 evidence?” | Calendar empty during observation | Annual tasks in GRC with artifacts |
| Claim “SOC complete” without Type II | “May we see the report?” | Diligence mismatch; investor risk | Accurate status + dated timeline |
| Safety theater only in marketing | “How do you handle injection?” | No eng evidence | Link to product controls + tests |
SOC 2 vs ISO 27001 (short decision)
US enterprise SaaS buyers often ask SOC 2 first. ISO 27001 may matter for EU-heavy pipelines or specific industries. Many seed AI SaaS teams start with SOC 2 Security (plus Availability/Privacy as deals require), then expand frameworks when pipeline demand is real. Do not run dual frameworks as status symbols while basic access reviews are still broken.
CorpIM demo path
- Compliance — Open the SOC 2 checklist view; note gaps on CC6.1, CC7.2, CC8.1.
- Guide — Select an enterprise-ready stage; review north-star pipeline and compliance todos.
- Startup Copilot — Ask: “What SOC 2 evidence is still missing?” and require cited control links—not vibes.
Live demo: https://www.romewayai.com/corp-im/. Soft next step: walk Compliance once with your real subprocessor list and CI evidence paths in mind, then assign owners for the top five gaps.
What we did not test
This article is desk synthesis. We did not conduct an audit, publish pass rates, invent control scores, or claim any GRC vendor is “best.” AICPA materials define the framework; your auditor defines report scope. Tool names (Probo-class GRC, Trivy/Semgrep-class scanners) are illustrative of common evidence pipelines, not endorsements or benchmarks.
FAQ
Is this a SOC 2 checklist for AI SaaS that replaces an auditor?
No. It is an operating checklist so seed–Series A teams collect the right evidence early. Auditors still define scope, sampling, and opinion. Use this to stop scrambling; do not treat it as a certificate.
SOC 2 Type I vs Type II for seed stage?
Type I is point-in-time design; Type II is operating effectiveness over months. Enterprise buyers increasingly want Type II or a clear timeline. Start the control map early so the observation period is not empty.
Does using a hosted model API require a separate “AI policy”?
You need subprocessor disclosure, data processing terms, and an internal policy on what customer content may be sent to inference APIs. Policy plus technical controls (minimization, redaction, region routing) together—not policy alone.
How should investor updates mention SOC status?
Disclose honestly in risks: “Type II observation started {month}; CC8.1 vendor reviews scheduled” beats silence or “SOC done” without a report. Details in AI-assisted investor updates.
SOC 2 vs ISO 27001 for AI SaaS?
US enterprise SaaS buyers often ask SOC 2 first. ISO may matter for EU or specific industries. Start where your pipeline asks; expand when demand is documented, not aspirational.
Continue the semantic path
Next: draft diligence-safe narrative with AI-assisted investor updates, wire company loops via AI-native Startup OS, and harden runtime proof with open-source observability for LLM apps. For incident evidence quality, continue to AI feature production root cause.