What is Cohort Analysis?
Cohort analysis compares metrics for groups of users who share a start event and the time of that event. A cohort is not “everyone in March” and not a segment such as “Android from Germany.” It is the set of people who first clicked, signed up, paid, or installed in the same date (or week) — then watched along lifetime, not in a calendar slice.
A “March revenue” slice mixes people acquired in January with people who arrived yesterday. If March revenue rose, you cannot say the channel improved: leftover rebills can explain it. A cohort answers differently: people whose first purchase was in week 12 had brought $Y by day 30 of life — and that figure is compared with week 8. In performance, cohorts exist mainly for LTV and retention. Without them, “average LTV of all time” lags current traffic and breaks unit economics.
How to read a cohort table
The classic grid: rows are start date (week), columns are cohort age (day 0, 7, 30…), cells are return rate or cumulative revenue per user.
Example (cumulative LTV per payer, $). Cohort W1: 1,000 users, D0 = 12, D7 = 18, D30 = 24. Cohort W4: the same 1,000 and the same D0 = 12, but D7 = 14 and D30 = 16. At the start the funnels look identical. By D30 the second lags: the same day-0 CPA, a worse tail. Scaling W4 “because yesterday was green” repeats the worse economics. A blended calendar-month report hides this if W1 is still dripping rebills.
Read the grid on two axes. Along a row — how one cohort ages: how fast LTV accrues, where the curve flattens (further retention spend is wasted). Down a column — whether fresh traffic is worse on the same day of life (D7 of W4 below D7 of W1). Cohort size matters: forty conversions at D30 is noise, not a law. Holds and late approvals shift cells right: an “empty” D7 may fill two weeks later — payout delay, not churn.
LTV cohorts and retention cohorts
LTV cohort. The start axis is the date of first value (payment, FTD, approved lead); cells are cumulative margin. That is how you tell whether CAC pays back and in how many days. Gambling uses NGR of the cohort; nutra uses trial plus the rebill chain; apps use subscription and IAP. LTV formulas are the same as in the LTV article; the cohort decides whom you average.
Retention cohort. A cell is the share of the cohort that did a repeat event on Dn (session, payment, deposit). D1 / D7 / D30 are conventional checkpoints, not separate glossary entities. Retention explains the shape of LTV: if D7 falls, the LTV tail shortens even at the same average check.
The two reports are not interchangeable. High D1 with low LTV happens on cheap return visits that never pay. High LTV with middling retention happens when a minority has a fat check. Margin needs revenue, not only “came back to the app.”
Other cohort keys besides date: offer, GEO, source, creative, funnel (with / without pre-land). The point is the same — do not average unlike things. An “all Facebook” cohort hides that one ad set feeds LTV and another feeds only day-0 leads with zero approval.
GA4 as an example, not the whole article
In Google Analytics 4 the cohort report is one Explorations template (Cohort exploration): cohort by first-visit or first-event date, return or value as the metric. That is a convenient example of an interface, not the definition of the method.
Limits of GA4 as the only cohort source: the event model and identity thresholds empty small cells; value in GA4 is whatever you sent as purchase/value, not affiliate payout; attribution on the GA4 property does not match CPA-network last click. For buyer money, cohorts usually sit in a tracker or BI on click date / postback date and approved status. GA4 is useful on your own landing: did the visit return, did it reach the form. It does not know payout or hold.
Affiliate practice and limits
A practical minimum: cohort by week of first approved and cumulative payout at D0 / D7 / D30, plus approve rate on the same row. If the network sends rebills as a separate goal, use separate columns — do not sum “all conversions.” Reconciliation with the ad account can use UTMs or sub IDs, but the cohort key should stay server-side (click ID + conversion time); otherwise returns without tags smear the row.
Small n, payout or cap changes mid-week, fraud the network cuts 20 days later — all of that breaks W1 vs W4. Privacy (ATT, ITP) cuts observed return in web analytics harder than in S2S postbacks. A cohort does not prove causation: a D30 drop after a creative change may be GEO mix, not “the creative kills LTV.” Causation needs a control and enough volume, not one grid.
See also: LTV, CAC, unit economics, GA4, tracker.