Quick answer
To bill for AI API usage, separate three layers: (1) meter inference or agent actions at the edge, (2) aggregate usage in an open metering plane (e.g. Lago), (3) charge via Stripe with plan tiers and overage rules. Seat-only pricing breaks when one power user burns GPU budget; pure pay-as-you-go scares finance buyers. Hybrid plans with included quotas plus metered overage are the 2026 default for AI SaaS. See the Revenue loop in CorpIM Studio (Lago + Stripe Bridge connectors).
Key takeaways
- Meter at the request boundary—tokens, calls, or “agent steps”—not only monthly invoices from the model provider.
- Open-source metering (Lago) pairs with Stripe for cards, invoices, and dunning.
- Expose usage in-product; trials with high API use and no card are margin risks, not leads.
- Align product analytics events with billing events so PostHog and finance tell the same story.
- Model cost changes require a repricing playbook—do not silently eat margin.
Who this is for
- AI SaaS founders pricing API or copilot features with variable inference cost.
- FinOps leads comparing proprietary billing suites vs a composable open stack.
- Engineers implementing usage events alongside product analytics.
- RevOps owners who need churn signals from quota and payment failures.
- Teams preparing enterprise procurement that asks for exportable usage logs.
Who should skip
- Pure seat-based B2B with no variable inference cost per customer.
- Teams already locked into Zuora/Recurly with working metered lines—migrate only if lock-in or cost hurts.
- Pre-revenue prototypes with fewer than ~10 design partners—manual invoices can suffice until usage patterns stabilize.
- Readers needing only TCO of self-host vs API inference—see self-host vs API TCO.
Prerequisites
- Stable
tenant_id(or account ID) shared across API gateway, analytics, support, and billing. - A place to emit events at request completion (gateway, worker, or agent runtime).
- A payment provider account (this guide uses Stripe’s public usage-based billing docs as the charge layer).
- Willingness to show customers a usage page—enterprise buyers expect it.
- An owner for failed payments (RevOps or founder)—webhooks without owners are noise.
If you lack shared IDs, fix identity before meters. Meters without join keys produce invoices nobody can explain.
Pricing models (pick one primary)
| Model | Pros | Risks | When to pick |
|---|---|---|---|
| Per seat | Simple sales | Heavy users destroy margin | Low variance in AI usage per seat |
| Included quota + overage | Predictable for buyers | Needs clear usage UI | Default for AI SaaS in 2026 desk synthesis |
| Pure usage | Aligns with cost | Revenue volatility, CFO fear | Developer APIs, infra products |
| Outcome-based | High value capture | Attribution disputes | Clear measurable outcomes only |
| Credits wallet | Flexible packaging | FX of credit→token confusion | Multi-model products with changing costs |
Desk recommendation: start with included quota + overage for B2B SaaS. Pure usage works for developer platforms with sophisticated buyers. Outcome-based pricing is a later SKU when attribution is dispute-proof.
Seat-only remains fine when AI is a thin assist and cost variance per seat is negligible. The moment one customer’s agent runs dominate GPU spend, hybrid metering becomes mandatory.
Billable units: design for humans and machines
Choose units customers understand and engineers can emit without debate:
- Tokens / 1k tokens — aligns with provider cost; harder for non-technical buyers.
- AI credits — package tokens + tool calls; requires a published conversion table.
- API calls — simple; unfair if call cost variance is huge.
- Agent runs / steps — matches product value for agentic features; define “run” tightly.
- Completed reports / jobs — outcome-adjacent; need idempotency so retries do not double bill.
Publish the unit definition in docs and in-product. Ambiguity becomes support tickets and chargebacks.
Align engineering metrics with billing events so product and finance see the same numbers—see AI feature adoption metrics. PostHog can track product funnels; billing must use a durable event stream you can replay. Do not use analytics as the only billing ledger.
Implementation checklist
Step 1: Define billable units and plan matrix
Write a one-pager: unit definition, included quotas per plan, overage rate, free-tier cap, and grandfathering policy for price changes. Get founder + engineering + sales sign-off before coding. Changing units after customers pay is painful.
Step 2: Emit usage events at the edge
At API gateway or worker completion, emit structured events: tenant_id, feature, units, model_id, timestamp, optional request_id. Buffer and retry; billing loss is worse than analytics loss.
Idempotency keys matter. Retries and duplicate workers should not double-count. Store raw events for dispute windows even if you aggregate for invoices.
Include model_id even if you do not bill by model yet—repricing and margin analysis need it later.
Step 3: Aggregate in Lago (or equivalent open metering)
Map plans, thresholds, and coupons. Lago’s open model fits teams that want to avoid heavy proprietary lock-in at seed stage. Per Lago documentation, define billable metrics, attach them to plans, and record usage events against subscriptions.
Test: inject synthetic usage for a test tenant, generate an invoice preview, and reconcile event sum to billed units. If you cannot reconcile in staging, do not go live.
Step 4: Stripe for money movement
Use Stripe for subscriptions, metered line items, taxes where applicable, and failed payment retries. Follow Stripe’s public guidance on usage-based subscriptions.
Bridge webhooks back to IM or RevOps todos when cards fail. CorpIM’s Revenue loop demo pattern (for example a failed payment todo) is the operating shape—even if your ticket IDs differ.
Decide whether Lago (or your meter) is the source of truth for quantities and Stripe is the charge executor, or whether Stripe Billing alone meters. Many teams keep metering logic in Lago and charge in Stripe to preserve plan flexibility.
Step 5: Customer-facing usage dashboard
Buyers expect a usage page before enterprise procurement. Show period usage, quota remaining, projected overage, and exportable logs. Security reviews often ask who can access usage metadata and how long you retain it—tie this to your privacy controls and SOC 2 checklist for AI SaaS.
Surface soft alerts at 50/80/100% quota. Silent 100% cutoffs create rage churn; silent overages create invoice shock.
Step 6: Wire product analytics (not as ledger)
Emit parallel PostHog events for adoption funnels (feature opened → successful outcome). Finance uses billing events; product uses analytics. Reconcile weekly in the Friday metrics ritual—see weekly operating rhythm.
Step 7: Repricing when model costs shift
When provider list prices change or you swap models, run a playbook: margin model → customer comms → grandfathering policy → plan update in Lago → Stripe price objects. Do not absorb large cost spikes silently and hope usage stays flat.
Document effective dates. Enterprise contracts may need amendment windows; self-serve plans need in-app notices.
Step 8: Verify end-to-end
- Create test tenant on each plan.
- Generate known unit counts (e.g. 1,000 credits).
- Confirm Lago aggregation matches.
- Confirm Stripe invoice line items match.
- Fail a card and confirm a RevOps todo appears.
- Export usage CSV and spot-check request IDs.
Verification is part of the how-to, not optional QA theater.
Open-source stack map (reference)
| Layer | Example | Job |
|---|---|---|
| Edge meter | Your gateway / worker | Emit units with idempotency |
| Aggregation | Lago | Plans, metrics, coupons, invoices logic |
| Payments | Stripe | Cards, dunning, tax hooks, payouts |
| Product analytics | PostHog | Adoption funnels (not billing ledger) |
| Support | Chatwoot-class | “Unexpected usage” tickets |
| Ops surface | IM + todos / CorpIM Revenue | Owners for churn and failed payments |
This stack sits inside a broader AI-native Startup OS: revenue loop events should appear beside support and product signals, not in a finance silo.
Self-hosting Lago vs buying a SaaS billing suite is a TCO and compliance choice. Pair with self-host vs API TCO for inference hosting decisions—those costs feed your margin model even when customers pay you in credits.
Margin model (simple, non-invented)
At minimum track per tenant:
- Billed revenue (period)
- Provider inference cost attributed to tenant (from your edge meter × provider rates)
- Gross margin after inference (before other COGS)
Do not invent industry-average margins here. Build your sheet from your provider invoices and your meter. When provider prices change, update the sheet the same day you update plans.
If you cannot attribute cost per tenant, you cannot price overage safely. That is the real reason edge metering exists.
Event schema and idempotency (engineer-facing)
A practical usage event payload:
{
"tenant_id": "cus_123",
"subscription_id": "sub_456",
"feature": "report_agent",
"billable_metric_code": "ai_credits",
"units": 3,
"model_id": "provider-model-id",
"request_id": "req_789",
"idempotency_key": "req_789:ai_credits",
"timestamp": "2026-09-03T12:00:00Z"
}
Rules that prevent finance pain:
- Emit only after the billable work is durable (job completed or streamed tokens finalized).
- Use
idempotency_keyso retries do not double bill. - Keep raw events for a dispute window even after aggregation.
- Never let the LLM invent units—metering code owns the number.
Map billable_metric_code to Lago billable metrics per Lago docs. Keep a parallel analytics event for product funnels in PostHog, but do not invoice from PostHog.
Hybrid plan design worked example (illustrative structure)
The numbers below are structural examples, not market benchmarks or recommended prices:
- Starter: included N credits / month; overage per credit; soft alert at 80%.
- Growth: higher included credits; lower overage; usage export enabled.
- Enterprise: custom included pool; committed invoice; usage API for procurement.
Sales needs a one-pager that explains what a credit buys (e.g. “1 credit ≈ X tokens or 1 agent step on model class Y”). When you change models, update the credit definition with an effective date—not a silent shift that makes historical invoices incomparable.
Trial policy should state: free credits amount, hard stop vs soft stop, and whether a card is required before overage. Soft stops without RevOps alerts recreate the “lead burns inference” failure mode.
Webhook and dunning operations
Stripe will emit payment lifecycle events. Your job is to turn them into owned work:
- Failed payment → IM card + RevOps todo with tenant deeplink.
- Successful retry → close todo; restore soft-locked features if you locked them.
- Subscription canceled → product offboarding checklist (data retention, API keys).
- Dispute / chargeback → pause aggressive collection automation; involve human.
Document the mapping in your weekly rhythm so it is not tribal knowledge—see AI startup weekly operating rhythm. Billing without operators is how SaaS discovers silent churn at month-end.
Testing matrix before production
| Case | Expected |
|---|---|
| Exact included quota used | Invoice shows zero overage |
| Quota + 1 unit | Overage line item appears |
| Duplicate event same idempotency key | Billed once |
| Worker retry after success | Billed once |
| Card decline | Todo created; access policy applied as designed |
| Plan upgrade mid-cycle | Proration matches your published policy |
| Model_id change, same units | Margin sheet updates; customer unit definition unchanged unless announced |
Automate as many of these as you can in CI against test Stripe/Lago projects. Manual-only verification decays the first time you change plan catalogs.
Credits vs tokens: packaging without lying
Credits help sales tell a story. Tokens help FinOps track provider cost. You need both layers:
- Customer layer: credits, runs, or seats+quota—stable language in contracts and UI.
- Cost layer: tokens, tool calls, GPU-seconds—internal only unless you sell raw API access.
Publish a conversion table and version it. When a new model is 2× costlier per token, either raise credit price, change credits-per-action, or restrict that model to higher plans. Doing none of the three is a margin leak.
Avoid “unlimited AI” marketing while your provider invoice is linear with usage. Unlimited promises without hard engineering caps are how startups learn the word “adverse selection.”
Accounting sync (light)
Finance still needs recognizable artifacts: invoices, tax fields, revenue recognition notes your accountant accepts. Stripe is usually the customer-facing invoice source; export or sync into your bookkeeping tool on a schedule. Do not make Lago your general ledger.
For AI SaaS, call out deferred revenue vs usage overage clearly with your accountant—especially when customers prepay credit packs. This article does not prescribe GAAP treatment; it only flags that prepaid credits create operational and accounting complexity you should design for early.
Rollout sequence for an existing seat-priced product
- Add edge metering with no customer charge—shadow mode for 2–4 weeks.
- Build internal margin by tenant; identify outliers.
- Ship usage dashboard read-only to customers.
- Announce hybrid plans with grandfathering for current seats.
- Enable overage billing; monitor support ticket themes for two billing cycles.
- Only then remove pure seat SKUs that destroy margin.
Shadow metering first reduces political fights. You will learn which customers would have blown overage before you surprise them with invoices.
Support playbooks tied to billing
Three tickets every AI SaaS eventually sees:
- “Why did my bill jump?” — show usage breakdown by feature and day; offer plan coaching, not only refunds.
- “I was charged for failed jobs.” — check idempotency and whether failed work was marked billable; fix meter if needed.
- “Turn off AI for my account.” — product flag + stop emitting billable events; confirm in Lago that usage halts.
Put these playbooks next to your Revenue loop so CS is not inventing policy in the inbox. That is Startup OS discipline applied to billing, not a separate finance island.
Bottom-line decision rule
If AI cost varies materially by tenant, you must bill for AI API usage with edge meters, not seat hope. Prefer hybrid quota + overage, open metering (Lago) plus Stripe for money movement, and a usage UI customers can trust. Wire failed payments and quota risk into the same weekly ops rhythm as product and engineering—or margin and churn will surprise you on the same Friday.
Common billing mistakes
| Mistake | Symptom | Fix |
|---|---|---|
| Meter only provider invoice | Per-customer margin unknown | Edge metering per tenant |
| No card on high trial usage | “Lead” burns large inference | Quota cap + usage alerts |
| Analytics ≠ billing events | Product and finance argue in board meeting | Single event schema; dual sinks |
| Seat plan for agent product | One customer, huge agent run volume | Hybrid quota + overage |
| Ignore failed payment webhooks | Churn surprise at renewal | RevOps todo on Stripe dunning |
| Bill retries as new jobs | Customer disputes | Idempotent billable units |
| Hide usage UI | Invoice shock / trust loss | In-product quota dashboard |
| Silent cost absorption | Margin collapse after model price change | Repricing playbook with dates |
Churn signals in usage data
Patterns worth automating into todos:
- Trial tenant above ~80% quota with no payment method.
- Declining weekly API calls after an onboarding spike.
- Support tickets about “unexpected usage” before a downgrade.
- Repeated soft declines on card with continued product login.
- Enterprise account flat usage while seats expand (seat-only leakage or dead AI feature).
CorpIM Copilot demo answers “which trial accounts are at churn risk?” by joining Lago + support context. Wire the same join in production RevOps playbooks under weekly operating rhythm.
Enterprise and trust requirements
Usage logs may contain customer metadata (prompts hashes, feature names, volumes). Treat retention, access control, and subprocessors as part of privacy and security reviews. The AICPA SOC for Service Organizations overview gives vocabulary buyers use; map your metering and Stripe subprocessors into vendor inventories.
Exportable usage and clear unit definitions reduce procurement friction more than marketing claims about “transparent AI pricing.”
Decision table: when to add metering
| Signal | Action |
|---|---|
| AI COGS variance by customer is material | Ship edge meter + hybrid plan now |
| Sales cannot explain overages | Pause discounts; ship usage UI first |
| <10 design partners, flat usage | Manual invoice + soft caps OK |
| Provider price change > your buffer | Run repricing playbook this week |
| Enterprise asks for usage export | Prioritize CSV/API export before vanity charts |
CorpIM demo path
- Studio → Revenue loop → subscriptions table.
- Workbench todo for churn risk on a trial account.
- Run playbook Churn risk save under Guide.
- Ask Startup Copilot which trials are at risk and check citations.
Explore the pattern live: https://www.romewayai.com/corp-im/. Soft next step: list your billable unit and whether Stripe or Lago will own quantity truth before you write code.
Failure modes in metered AI billing
Usage billing fails in predictable ways. Design for these before you announce overage to customers:
- Double-count on retries. Provider retries and client retries both emit “usage” if the meter key is request-id instead of idempotent
usage_event_id. Customers dispute invoices; finance loses trust. Require idempotency at the edge and in Lago-class aggregation. - Clock skew across regions. Billing period close that uses local worker time will strand events in the wrong month. Standardize on UTC event time from the meter, not invoice-job local time.
- Seat plan + uncapped agents. One “seat” that spawns unbounded agent loops will blow COGS while MRR looks healthy. Soft caps and kill switches belong in the product, not only in the invoice.
- Credits that do not map to COGS. Marketing “10k credits” without a published conversion to tokens (or a fixed COGS buffer) creates support theater when providers reprice. Publish the mapping or stop calling them credits.
- Webhook lag as “churn.” Treating delayed Stripe events as cancelations creates false save plays. Deduplicate by subscription id and event type before RevOps todos fire—same shape as the CorpIM Revenue loop pattern.
When a dispute arrives, your evidence pack should show: raw meter event, aggregated quantity, plan rule applied, and Stripe invoice line—four IDs, not a screenshot of a dashboard. Tie disputes into the same weekly review as operating rhythm Friday metrics so billing debt does not wait for quarterly panic.
What we did not test
This is not tax, accounting, or legal advice. We did not publish vendor price comparisons, claim ROI percentages, or certify Lago/Stripe configurations for your jurisdiction. Read current Lago docs and Stripe usage-based billing docs for API details that change over time.
FAQ
Should we bill tokens or outcomes?
Tokens align with provider cost and are easy to meter. Outcomes align with customer value but need dispute-proof definitions. Most AI SaaS ship token/credit metering first, then add outcome tiers for premium SKUs.
Lago vs Stripe Billing alone?
Stripe handles payments and can meter usage, but Lago adds plan complexity, coupons, and multi-product metering with an open core. Many teams use Lago for metering logic + Stripe for money movement. Choose based on plan complexity and lock-in tolerance.
How do usage metrics tie to SOC 2?
Usage logs may contain customer content metadata—treat retention and access controls as part of privacy controls. See SOC 2 checklist for AI SaaS and AICPA’s SOC overview for buyer vocabulary.
When should free tier include AI credits?
Enough for activation (first successful outcome), not enough for production workloads. Cap free inference and alert RevOps when trials approach quota without a card.
What if we self-host models?
You still meter tenant usage for fairness and packaging. Cost basis shifts from provider invoices to GPU TCO—use self-host vs API TCO—but the customer-facing billable unit can remain credits or runs.
Sources and further reading
- Lago documentation — billable metrics, plans, usage events
- Stripe usage-based billing — metered subscriptions
- PostHog documentation — product analytics alongside billing
- AICPA SOC for Service Organizations — trust services context
- Internal: AI-native Startup OS, adoption metrics, TCO calculator
Continue the semantic path
Self-host vs API TCO · AI-native Startup OS · Measure AI feature adoption · Weekly operating rhythm