# Analysis — balance, segmentation, pilot comparison, measurement plan

Generated by `python -m pso build` from `data/raw/pso_funnel_events.csv` (sha256 `0d7f8ca86978e06923e8e659d4841a7aca909f92227e094144505964a12ba5e7`, 37,899,350 bytes) with polars 2.0.0 on CPython 3.12.3. Do not edit by hand: every number here is regenerated by the command, and the output is deterministic, so `git diff` after regenerating shows exactly what changed.

## Definitions (A20–A24) — thresholds fixed before any number was computed; items marked † were added after the first numbers existed (see the assumption list, A20 and A24)

- **Intervals (A20):** a rate gets a *Wilson score interval* — the range of true rates under
  which the observed count would not be surprising at 95%; a difference between two rates gets
  *Newcombe's hybrid score interval*, built from the two Wilson intervals; the pooled early −
  late difference across projects is the *Mantel–Haenszel-weighted risk difference* (weights
  n₁n₂/N per project) with the Greenland–Robins (1985) variance in its large-strata form. All
  closed-form. † The one exception is a difference of medians (time to finish), which has no
  closed form and gets a percentile bootstrap: each arm's finish times resampled with
  replacement 2,000 times under a fixed seed, middle 95% of the differences. The H2 time guard
  reads the point estimate, not the interval.
- **Balance (A21):** per pilot project and dimension, the largest absolute early − late share
  difference, on mature pairs; flagged at ≥ 5.0 pp. Descriptive, no test statistic.
- **Segmentation (A22):** for each lens's top-ranked stage, completion per level of each declared
  dimension on mature pairs; levels under 200 pairs are pooled into "other";
  *strength* = range of completion across the remaining levels; the two strongest dimensions get
  a two-way table whose cells are flagged, not pooled, under 200. Acceptance source
  cannot segment the Invitation stage (it exists only after acceptance). Estimated assessment
  minutes is a project attribute and is read through *project*.
- **Hypotheses (A23):** verdicts are three-valued from an interval — *supported* when it lies
  at or above 0, *not supported* when it lies below 0, *inconclusive* when it straddles 0. H1 is
  *supported* when early − late finished is at or above 0 and no post-slot stage (Assessment,
  Contract) has the early arm significantly below the late arm; *not supported* when the
  finished difference or any post-slot stage lies below 0; *inconclusive* otherwise. H2 is
  *supported* when late − early finished lies at or above 0 and the late arm's median
  allocation → finished is less than 12.0 h longer; *not supported* when the
  rate difference lies below 0 or the time guard fails with the rate supported; *inconclusive*
  otherwise. Consent rate at the slot is shown but survival-conditioned.
- **Measurement plan (A24):** two-proportion z-test, 80% power, 5% two-sided; horizons
  2, 4, 8 weeks; effects 1, 2, 3 pp. † Computed for each
  population a recommendation would run on — all projects; projects with assessments of
  40 minutes or more (a split read off §2.2, not pre-registered; no project sits
  between 30 and 40 minutes); one pilot-type project — with each population's weekly volume
  as the mean over its own full weeks.

## 1. Balance of the pilot arms (A10)

| project | dimension | largest gap at | early_% | late_% | gap_pp | flag |
|---|---|---|---|---|---|---|
| Lumen | cohort_week | 2026-07-13 | 14.8 | 12.4 | 2.4 | no |
| Lumen | device | unknown | 41.4 | 40.1 | 1.3 | no |
| Lumen | tenure | new | 53.1 | 50.5 | 2.6 | no |
| Lumen | prior_projects | 0 | 53.1 | 50.5 | 2.6 | no |
| Lumen | country | US | 43 | 40.7 | 2.3 | no |
| Lumen | concurrent_psos | 1 | 70.5 | 71.1 | 0.6 | no |
| Lumen | accept_source | email | 34.8 | 36.1 | 1.3 | no |
| Orchid | cohort_week | 2026-07-13 | 12.5 | 15.6 | 3 | no |
| Orchid | device | unknown | 41.4 | 39.3 | 2 | no |
| Orchid | tenure | new | 53.6 | 52.1 | 1.5 | no |
| Orchid | prior_projects | 1–2 | 25.4 | 27 | 1.6 | no |
| Orchid | country | CA | 8.1 | 6.6 | 1.5 | no |
| Orchid | concurrent_psos | 1 | 68.8 | 71.3 | 2.5 | no |
| Orchid | accept_source | unknown | 61.2 | 59.5 | 1.7 | no |
| Vesper | cohort_week | 2026-07-20 | 13.4 | 11.7 | 1.7 | no |
| Vesper | device | desktop | 15.5 | 17.2 | 1.7 | no |
| Vesper | tenure | new | 53.1 | 52.4 | 0.7 | no |
| Vesper | prior_projects | 1–2 | 25 | 26.4 | 1.4 | no |
| Vesper | country | US | 41.6 | 44.2 | 2.6 | no |
| Vesper | concurrent_psos | 1 | 69.7 | 71.3 | 1.5 | no |
| Vesper | accept_source | unknown | 62.2 | 60.7 | 1.6 | no |

Flags: 0 of 21 project × dimension checks at or above 5.0 pp.

## 2.1 Segmentation — lens `all`, stage `invitation` (mature pairs)

Among the 26,351 mature allocations that never accepted: 672 were never sent an email, 1,587 were sent one that was never delivered, 15,175 were delivered with no open recorded, and 8,917 opened and did not accept (four exclusive categories). These are recorded signals, not causes: an open can go unrecorded (R7, A14).

Device coverage on this scope — a device is recorded only on events the tasker performs, so "unknown device" is a property of what happened, not of the tasker, and one project records no device at all:

| project | pairs | with a device | unknown device | unknown device and accepted |
|---|---|---|---|---|
| Beacon | 2,044 | 1,243 | 801 | 0 |
| Cobalt | 4,635 | 2,821 | 1,814 | 0 |
| Granite | 2,421 | 1,454 | 967 | 0 |
| Halcyon | 3,880 | 2,332 | 1,548 | 0 |
| Juniper | 5,502 | 3,348 | 2,154 | 0 |
| Lumen | 3,228 | 1,913 | 1,315 | 0 |
| Meridian | 4,188 | 2,532 | 1,656 | 1 |
| Orchid | 3,095 | 1,847 | 1,248 | 0 |
| Quartz | 4,719 | 0 | 4,719 | 1,950 |
| Sequoia | 3,648 | 2,166 | 1,482 | 0 |
| Tundra | 3,163 | 1,865 | 1,298 | 0 |
| Vesper | 3,123 | 1,830 | 1,293 | 0 |

Strength ranking:

| dimension | strength_pp |
|---|---|
| device | 57.6 |
| country | 3.4 |
| project | 2.8 |
| cohort_week | 2.7 |
| prior_projects | 1.1 |
| concurrent_psos | 0.4 |
| tenure | 0.2 |

#### device — strength 57.6 pp

| level | reached | completed | completion_% | 95% interval |
|---|---|---|---|---|
| desktop | 6,474 | 4,035 | 62.3 | 61.1–63.5 |
| mobile | 16,078 | 10,800 | 67.2 | 66.4–67.9 |
| tablet | 799 | 509 | 63.7 | 60.3–67.0 |
| unknown | 20,295 | 1,951 | 9.6 | 9.2–10.0 |

#### country — strength 3.4 pp

| level | reached | completed | completion_% | 95% interval |
|---|---|---|---|---|
| BR | 3,176 | 1,259 | 39.6 | 38.0–41.4 |
| CA | 3,322 | 1,344 | 40.5 | 38.8–42.1 |
| DE | 3,183 | 1,252 | 39.3 | 37.7–41.0 |
| GB | 3,109 | 1,250 | 40.2 | 38.5–41.9 |
| IN | 3,135 | 1,259 | 40.2 | 38.5–41.9 |
| KE | 2,963 | 1,155 | 39 | 37.2–40.8 |
| NG | 2,966 | 1,126 | 38 | 36.2–39.7 |
| PH | 3,240 | 1,340 | 41.4 | 39.7–43.1 |
| US | 18,552 | 7,310 | 39.4 | 38.7–40.1 |

#### project — strength 2.8 pp

| level | reached | completed | completion_% | 95% interval |
|---|---|---|---|---|
| Beacon | 2,044 | 803 | 39.3 | 37.2–41.4 |
| Cobalt | 4,635 | 1,852 | 40 | 38.6–41.4 |
| Granite | 2,421 | 950 | 39.2 | 37.3–41.2 |
| Halcyon | 3,880 | 1,519 | 39.1 | 37.6–40.7 |
| Juniper | 5,502 | 2,225 | 40.4 | 39.2–41.7 |
| Lumen | 3,228 | 1,258 | 39 | 37.3–40.7 |
| Meridian | 4,188 | 1,649 | 39.4 | 37.9–40.9 |
| Orchid | 3,095 | 1,229 | 39.7 | 38.0–41.4 |
| Quartz | 4,719 | 1,950 | 41.3 | 39.9–42.7 |
| Sequoia | 3,648 | 1,434 | 39.3 | 37.7–40.9 |
| Tundra | 3,163 | 1,222 | 38.6 | 37.0–40.3 |
| Vesper | 3,123 | 1,204 | 38.6 | 36.9–40.3 |

#### cohort_week — strength 2.7 pp

| level | reached | completed | completion_% | 95% interval |
|---|---|---|---|---|
| 2026-06-01 | 1,771 | 707 | 39.9 | 37.7–42.2 |
| 2026-06-08 | 2,530 | 980 | 38.7 | 36.9–40.6 |
| 2026-06-15 | 3,071 | 1,227 | 40 | 38.2–41.7 |
| 2026-06-22 | 3,405 | 1,378 | 40.5 | 38.8–42.1 |
| 2026-06-29 | 3,758 | 1,489 | 39.6 | 38.1–41.2 |
| 2026-07-06 | 4,142 | 1,638 | 39.5 | 38.1–41.0 |
| 2026-07-13 | 4,366 | 1,723 | 39.5 | 38.0–40.9 |
| 2026-07-20 | 4,620 | 1,839 | 39.8 | 38.4–41.2 |
| 2026-07-27 | 4,844 | 1,946 | 40.2 | 38.8–41.6 |
| 2026-08-03 | 5,009 | 1,982 | 39.6 | 38.2–40.9 |
| 2026-08-10 | 5,363 | 2,096 | 39.1 | 37.8–40.4 |
| 2026-08-17 | 767 | 290 | 37.8 | 34.4–41.3 |

#### prior_projects — strength 1.1 pp

| level | reached | completed | completion_% | 95% interval |
|---|---|---|---|---|
| 0 | 22,854 | 9,082 | 39.7 | 39.1–40.4 |
| 1–2 | 11,112 | 4,333 | 39 | 38.1–39.9 |
| 3+ | 9,680 | 3,880 | 40.1 | 39.1–41.1 |

#### concurrent_psos — strength 0.4 pp

| level | reached | completed | completion_% | 95% interval |
|---|---|---|---|---|
| 1 | 31,962 | 12,631 | 39.5 | 39.0–40.1 |
| 2+ | 11,684 | 4,664 | 39.9 | 39.0–40.8 |

#### tenure — strength 0.2 pp

| level | reached | completed | completion_% | 95% interval |
|---|---|---|---|---|
| new | 22,854 | 9,082 | 39.7 | 39.1–40.4 |
| returning | 20,792 | 8,213 | 39.5 | 38.8–40.2 |

#### Two-way: device × country

| device | country | reached | completed | completion_% | 95% interval | under 200 |
|---|---|---|---|---|---|---|
| unknown | BR | 1,483 | 144 | 9.7 | 8.3–11.3 | no |
| unknown | CA | 1,517 | 139 | 9.2 | 7.8–10.7 | no |
| unknown | DE | 1,473 | 146 | 9.9 | 8.5–11.5 | no |
| unknown | GB | 1,461 | 152 | 10.4 | 8.9–12.1 | no |
| unknown | IN | 1,467 | 162 | 11 | 9.5–12.7 | no |
| unknown | KE | 1,406 | 122 | 8.7 | 7.3–10.3 | no |
| unknown | NG | 1,405 | 109 | 7.8 | 6.5–9.3 | no |
| unknown | PH | 1,454 | 148 | 10.2 | 8.7–11.8 | no |
| unknown | US | 8,629 | 829 | 9.6 | 9.0–10.2 | no |
| desktop | BR | 490 | 297 | 60.6 | 56.2–64.8 | no |
| desktop | CA | 499 | 317 | 63.5 | 59.2–67.6 | no |
| desktop | DE | 468 | 276 | 59 | 54.5–63.3 | no |
| desktop | GB | 411 | 254 | 61.8 | 57.0–66.4 | no |
| desktop | IN | 500 | 321 | 64.2 | 59.9–68.3 | no |
| desktop | KE | 412 | 255 | 61.9 | 57.1–66.5 | no |
| desktop | NG | 422 | 264 | 62.6 | 57.8–67.0 | no |
| desktop | PH | 484 | 311 | 64.3 | 59.9–68.4 | no |
| desktop | US | 2,788 | 1,740 | 62.4 | 60.6–64.2 | no |
| mobile | BR | 1,149 | 784 | 68.2 | 65.5–70.9 | no |
| mobile | CA | 1,239 | 844 | 68.1 | 65.5–70.7 | no |
| mobile | DE | 1,180 | 793 | 67.2 | 64.5–69.8 | no |
| mobile | GB | 1,183 | 813 | 68.7 | 66.0–71.3 | no |
| mobile | IN | 1,109 | 738 | 66.5 | 63.7–69.3 | no |
| mobile | KE | 1,085 | 744 | 68.6 | 65.7–71.3 | no |
| mobile | NG | 1,090 | 719 | 66 | 63.1–68.7 | no |
| mobile | PH | 1,234 | 835 | 67.7 | 65.0–70.2 | no |
| mobile | US | 6,809 | 4,530 | 66.5 | 65.4–67.6 | no |
| tablet | BR | 54 | 34 | 63 | 49.6–74.6 | yes |
| tablet | CA | 67 | 44 | 65.7 | 53.7–75.9 | yes |
| tablet | DE | 62 | 37 | 59.7 | 47.3–71.0 | yes |
| tablet | GB | 54 | 31 | 57.4 | 44.2–69.7 | yes |
| tablet | IN | 59 | 38 | 64.4 | 51.7–75.4 | yes |
| tablet | KE | 60 | 34 | 56.7 | 44.1–68.4 | yes |
| tablet | NG | 49 | 34 | 69.4 | 55.5–80.5 | yes |
| tablet | PH | 68 | 46 | 67.6 | 55.8–77.6 | yes |
| tablet | US | 326 | 211 | 64.7 | 59.4–69.7 | no |

## 2.2 Segmentation — lens `accepted`, stage `assessment` (mature pairs)

Strength ranking:

| dimension | strength_pp |
|---|---|
| project | 9.9 |
| device | 7.8 |
| cohort_week | 7 |
| country | 3.6 |
| prior_projects | 1.7 |
| tenure | 0.7 |
| concurrent_psos | 0.2 |
| accept_source | 0.1 |

#### project — strength 9.9 pp

| level | reached | completed | completion_% | 95% interval |
|---|---|---|---|---|
| Beacon | 742 | 222 | 29.9 | 26.7–33.3 |
| Cobalt | 1,720 | 516 | 30 | 27.9–32.2 |
| Granite | 874 | 328 | 37.5 | 34.4–40.8 |
| Halcyon | 1,411 | 432 | 30.6 | 28.3–33.1 |
| Juniper | 2,094 | 756 | 36.1 | 34.1–38.2 |
| Lumen | 1,100 | 316 | 28.7 | 26.1–31.5 |
| Meridian | 1,526 | 589 | 38.6 | 36.2–41.1 |
| Orchid | 1,083 | 325 | 30 | 27.4–32.8 |
| Quartz | 1,810 | 654 | 36.1 | 34.0–38.4 |
| Sequoia | 1,327 | 513 | 38.7 | 36.1–41.3 |
| Tundra | 1,127 | 347 | 30.8 | 28.2–33.5 |
| Vesper | 1,034 | 351 | 33.9 | 31.1–36.9 |

#### device — strength 7.8 pp

| level | reached | completed | completion_% | 95% interval |
|---|---|---|---|---|
| desktop | 3,714 | 1,446 | 38.9 | 37.4–40.5 |
| mobile | 9,865 | 3,071 | 31.1 | 30.2–32.1 |
| tablet | 459 | 178 | 38.8 | 34.4–43.3 |
| unknown | 1,810 | 654 | 36.1 | 34.0–38.4 |

#### cohort_week — strength 7.0 pp

| level | reached | completed | completion_% | 95% interval |
|---|---|---|---|---|
| 2026-06-01 | 667 | 233 | 34.9 | 31.4–38.6 |
| 2026-06-08 | 917 | 323 | 35.2 | 32.2–38.4 |
| 2026-06-15 | 1,139 | 392 | 34.4 | 31.7–37.2 |
| 2026-06-22 | 1,284 | 475 | 37 | 34.4–39.7 |
| 2026-06-29 | 1,355 | 462 | 34.1 | 31.6–36.7 |
| 2026-07-06 | 1,492 | 482 | 32.3 | 30.0–34.7 |
| 2026-07-13 | 1,562 | 532 | 34.1 | 31.8–36.4 |
| 2026-07-20 | 1,657 | 578 | 34.9 | 32.6–37.2 |
| 2026-07-27 | 1,787 | 590 | 33 | 30.9–35.2 |
| 2026-08-03 | 1,809 | 578 | 32 | 29.8–34.1 |
| 2026-08-10 | 1,916 | 625 | 32.6 | 30.6–34.8 |
| 2026-08-17 | 263 | 79 | 30 | 24.8–35.8 |

#### country — strength 3.6 pp

| level | reached | completed | completion_% | 95% interval |
|---|---|---|---|---|
| BR | 1,156 | 368 | 31.8 | 29.2–34.6 |
| CA | 1,225 | 395 | 32.2 | 29.7–34.9 |
| DE | 1,135 | 367 | 32.3 | 29.7–35.1 |
| GB | 1,164 | 412 | 35.4 | 32.7–38.2 |
| IN | 1,155 | 394 | 34.1 | 31.4–36.9 |
| KE | 1,046 | 349 | 33.4 | 30.6–36.3 |
| NG | 1,049 | 358 | 34.1 | 31.3–37.1 |
| PH | 1,240 | 419 | 33.8 | 31.2–36.5 |
| US | 6,678 | 2,287 | 34.2 | 33.1–35.4 |

#### prior_projects — strength 1.7 pp

| level | reached | completed | completion_% | 95% interval |
|---|---|---|---|---|
| 0 | 8,320 | 2,781 | 33.4 | 32.4–34.4 |
| 1–2 | 3,981 | 1,327 | 33.3 | 31.9–34.8 |
| 3+ | 3,547 | 1,241 | 35 | 33.4–36.6 |

#### tenure — strength 0.7 pp

| level | reached | completed | completion_% | 95% interval |
|---|---|---|---|---|
| new | 8,320 | 2,781 | 33.4 | 32.4–34.4 |
| returning | 7,528 | 2,568 | 34.1 | 33.1–35.2 |

#### concurrent_psos — strength 0.2 pp

| level | reached | completed | completion_% | 95% interval |
|---|---|---|---|---|
| 1 | 11,581 | 3,904 | 33.7 | 32.9–34.6 |
| 2+ | 4,267 | 1,445 | 33.9 | 32.5–35.3 |

#### accept_source — strength 0.1 pp

| level | reached | completed | completion_% | 95% interval |
|---|---|---|---|---|
| email | 14,560 | 4,915 | 33.8 | 33.0–34.5 |
| in_app | 1,288 | 434 | 33.7 | 31.2–36.3 |

#### Two-way: project × device

| project | device | reached | completed | completion_% | 95% interval | under 200 |
|---|---|---|---|---|---|---|
| Beacon | desktop | 182 | 72 | 39.6 | 32.7–46.8 | yes |
| Beacon | mobile | 537 | 139 | 25.9 | 22.4–29.8 | no |
| Beacon | tablet | 23 | 11 | 47.8 | 29.2–67.0 | yes |
| Cobalt | desktop | 433 | 155 | 35.8 | 31.4–40.4 | no |
| Cobalt | mobile | 1,224 | 337 | 27.5 | 25.1–30.1 | no |
| Cobalt | tablet | 63 | 24 | 38.1 | 27.1–50.4 | yes |
| Granite | desktop | 244 | 97 | 39.8 | 33.8–46.0 | no |
| Granite | mobile | 595 | 218 | 36.6 | 32.9–40.6 | no |
| Granite | tablet | 35 | 13 | 37.1 | 23.2–53.7 | yes |
| Halcyon | desktop | 387 | 139 | 35.9 | 31.3–40.8 | no |
| Halcyon | mobile | 973 | 275 | 28.3 | 25.5–31.2 | no |
| Halcyon | tablet | 51 | 18 | 35.3 | 23.6–49.0 | yes |
| Juniper | desktop | 540 | 214 | 39.6 | 35.6–43.8 | no |
| Juniper | mobile | 1,490 | 517 | 34.7 | 32.3–37.2 | no |
| Juniper | tablet | 64 | 25 | 39.1 | 28.1–51.3 | yes |
| Lumen | desktop | 305 | 105 | 34.4 | 29.3–39.9 | no |
| Lumen | mobile | 762 | 195 | 25.6 | 22.6–28.8 | no |
| Lumen | tablet | 33 | 16 | 48.5 | 32.5–64.8 | yes |
| Meridian | desktop | 401 | 186 | 46.4 | 41.6–51.3 | no |
| Meridian | mobile | 1,077 | 381 | 35.4 | 32.6–38.3 | no |
| Meridian | tablet | 48 | 22 | 45.8 | 32.6–59.7 | yes |
| Orchid | desktop | 285 | 111 | 38.9 | 33.5–44.7 | no |
| Orchid | mobile | 765 | 202 | 26.4 | 23.4–29.6 | no |
| Orchid | tablet | 33 | 12 | 36.4 | 22.2–53.4 | yes |
| Quartz | unknown | 1,810 | 654 | 36.1 | 34.0–38.4 | no |
| Sequoia | desktop | 371 | 157 | 42.3 | 37.4–47.4 | no |
| Sequoia | mobile | 902 | 338 | 37.5 | 34.4–40.7 | no |
| Sequoia | tablet | 54 | 18 | 33.3 | 22.2–46.6 | yes |
| Tundra | desktop | 308 | 124 | 40.3 | 34.9–45.8 | no |
| Tundra | mobile | 789 | 211 | 26.7 | 23.8–29.9 | no |
| Tundra | tablet | 30 | 12 | 40 | 24.6–57.7 | yes |
| Vesper | desktop | 258 | 86 | 33.3 | 27.9–39.3 | no |
| Vesper | mobile | 751 | 258 | 34.4 | 31.0–37.8 | no |
| Vesper | tablet | 25 | 7 | 28 | 14.3–47.6 | yes |

## 3. Likeness pilot — early vs late, mature pairs

| project | early allocated | early finished | early_% | early 95% | late allocated | late finished | late_% | late 95% | early − late (pp) | 95% interval (pp) |
|---|---|---|---|---|---|---|---|---|---|---|
| Lumen | 1,601 | 105 | 6.56 | 5.45–7.88 | 1,627 | 155 | 9.53 | 8.19–11.05 | -2.97 | -4.85 to -1.09 |
| Orchid | 1,501 | 145 | 9.66 | 8.27–11.26 | 1,594 | 132 | 8.28 | 7.03–9.74 | 1.38 | -0.64 to 3.41 |
| Vesper | 1,570 | 143 | 9.11 | 7.78–10.63 | 1,553 | 160 | 10.3 | 8.89–11.91 | -1.19 | -3.28 to 0.89 |
| pooled — rates are simple pools, the difference is MH-weighted | 4,672 | 393 | 8.41 | 7.65–9.24 | 4,774 | 447 | 9.36 | 8.57–10.22 | -0.96 | -2.11 to 0.19 |

### Stages by arm, hypotheses per project

##### Lumen

| stage | early reached | early completion_% | late reached | late completion_% | early − late (pp) | 95% interval (pp) | verdict |
|---|---|---|---|---|---|---|---|
| Invitation | 1,601 | 38.4 | 1,627 | 39.6 |  |  |  |
| Confidentiality agreement | 614 | 81.6 | 644 | 93 |  |  |  |
| Qualification assessment | 501 | 25 | 599 | 31.9 | -6.9 | -12.2 to -1.6 | not supported |
| Work contract | 125 | 84 | 191 | 81.2 | 2.8 | -6.1 to 11.0 | inconclusive |

Consent at the slot (survival-conditioned, not evidence on its own): early 538 signed / 76 declined (87.6%), late 173 / 18 (90.6%). Median allocation → finished among finishers: early 70.851 h, late 85.191 h; late − early = 14.341 h, 95% bootstrap interval 0.076 to 32.138 h (2,000 resamples, seed 20261008); threshold: less than 12.0 h → exceeded.

**H1 not supported** (early − late finished at or above 0: not supported; post-slot stages not significantly below: assessment not supported, contract inconclusive). **H2 not supported** (late − early finished at or above 0: supported; time not ok).

##### Orchid

| stage | early reached | early completion_% | late reached | late completion_% | early − late (pp) | 95% interval (pp) | verdict |
|---|---|---|---|---|---|---|---|
| Invitation | 1,501 | 38.8 | 1,594 | 40.5 |  |  |  |
| Confidentiality agreement | 583 | 82 | 646 | 93.7 |  |  |  |
| Qualification assessment | 478 | 34.1 | 605 | 26.8 | 7.3 | 1.8 to 12.8 | supported |
| Work contract | 163 | 89 | 162 | 81.5 | 7.5 | -0.3 to 15.2 | inconclusive |

Consent at the slot (survival-conditioned, not evidence on its own): early 511 signed / 72 declined (87.7%), late 152 / 10 (93.8%). Median allocation → finished among finishers: early 63.13 h, late 81.411 h; late − early = 18.281 h, 95% bootstrap interval 6.646 to 32.983 h (2,000 resamples, seed 20261008); threshold: less than 12.0 h → exceeded.

**H1 inconclusive** (early − late finished at or above 0: inconclusive; post-slot stages not significantly below: assessment supported, contract inconclusive). **H2 inconclusive** (late − early finished at or above 0: inconclusive; time not ok).

##### Vesper

| stage | early reached | early completion_% | late reached | late completion_% | early − late (pp) | 95% interval (pp) | verdict |
|---|---|---|---|---|---|---|---|
| Invitation | 1,570 | 37.8 | 1,553 | 39.3 |  |  |  |
| Confidentiality agreement | 593 | 80.8 | 611 | 90.8 |  |  |  |
| Qualification assessment | 479 | 34.2 | 555 | 33.7 | 0.5 | -5.2 to 6.3 | inconclusive |
| Work contract | 164 | 87.2 | 187 | 85.6 | 1.6 | -5.7 to 8.8 | inconclusive |

Consent at the slot (survival-conditioned, not evidence on its own): early 509 signed / 84 declined (85.8%), late 179 / 8 (95.7%). Median allocation → finished among finishers: early 63.133 h, late 74.867 h; late − early = 11.734 h, 95% bootstrap interval 4.584 to 23.631 h (2,000 resamples, seed 20261008); threshold: less than 12.0 h → within.

**H1 inconclusive** (early − late finished at or above 0: inconclusive; post-slot stages not significantly below: assessment inconclusive, contract inconclusive). **H2 inconclusive** (late − early finished at or above 0: inconclusive; time ok).

##### pooled (finished difference MH-weighted; stage rows and medians are simple pools)

| stage | early reached | early completion_% | late reached | late completion_% | early − late (pp) | 95% interval (pp) | verdict |
|---|---|---|---|---|---|---|---|
| Invitation | 4,672 | 38.3 | 4,774 | 39.8 |  |  |  |
| Confidentiality agreement | 1,790 | 81.5 | 1,901 | 92.5 |  |  |  |
| Qualification assessment | 1,458 | 31 | 1,759 | 30.7 | 0.3 | -2.9 to 3.5 | inconclusive |
| Work contract | 452 | 86.9 | 540 | 82.8 | 4.2 | -0.3 to 8.6 | inconclusive |

Consent at the slot (survival-conditioned, not evidence on its own): early 1,558 signed / 232 declined (87.0%), late 504 / 36 (93.3%). Median allocation → finished among finishers: early 64.032 h, late 81.1 h; late − early = 17.068 h, 95% bootstrap interval 10.062 to 23.591 h (2,000 resamples, seed 20261008); threshold: less than 12.0 h → exceeded.

**H1 inconclusive** (early − late finished at or above 0: inconclusive; post-slot stages not significantly below: assessment inconclusive, contract inconclusive). **H2 inconclusive** (late − early finished at or above 0: inconclusive; time not ok).

## 4. Measurement plan — how we would know

- **Metric:** finished-PSO rate per allocated pair, mature at W = 14 days.
- **Guardrails:** median allocation → finished hours; first-task-submitted rate.
- **Design:** randomized at allocation within project, as the pilot was.
- Results are read 14 days after the last allocation. Each recommendation runs on its own population, so the volume and baseline below are per population.

#### all projects

Baseline 10.78% finished (4,707 of 43,646 mature pairs); 4,097 allocations per week on average over the 13 full weeks from 2026-06-01 to 2026-08-24 (a mean, not the current rate: projects ramp up, so recent weeks run higher and these test lengths are conservative).

What a test of a given length can detect (two arms, 80% power, 5% two-sided):

| weeks | pairs_per_arm | detectable_change_pp |
|---|---|---|
| 2 | 4,097 | 2 |
| 4 | 8,194 | 1.4 |
| 8 | 16,388 | 0.98 |

How long a test needs to detect a given change:

| change_to_detect_pp | pairs_per_arm | weeks_of_allocation |
|---|---|---|
| 1 | 15,714 | 7.7 |
| 2 | 4,079 | 2 |
| 3 | 1,879 | 0.9 |

#### projects with assessments of 40 minutes or more

Baseline 9.43% finished (1,890 of 20,045 mature pairs); 1,889 allocations per week on average over the 13 full weeks from 2026-06-01 to 2026-08-24 (a mean, not the current rate: projects ramp up, so recent weeks run higher and these test lengths are conservative).

What a test of a given length can detect (two arms, 80% power, 5% two-sided):

| weeks | pairs_per_arm | detectable_change_pp |
|---|---|---|
| 2 | 1,889 | 2.83 |
| 4 | 3,778 | 1.97 |
| 8 | 7,556 | 1.37 |

How long a test needs to detect a given change:

| change_to_detect_pp | pairs_per_arm | weeks_of_allocation |
|---|---|---|
| 1 | 14,038 | 14.9 |
| 2 | 3,665 | 3.9 |
| 3 | 1,697 | 1.8 |

#### one pilot-type (video) project — mean of the pilot projects

Baseline 8.89% finished (840 of 9,446 mature pairs); 465 allocations per week per project on average over the 8 full weeks from 2026-07-06 to 2026-08-24 (a mean, not the current rate: projects ramp up, so recent weeks run higher and these test lengths are conservative).

What a test of a given length can detect (two arms, 80% power, 5% two-sided):

| weeks | pairs_per_arm | detectable_change_pp |
|---|---|---|
| 2 | 465 | 5.93 |
| 4 | 930 | 4.05 |
| 8 | 1,861 | 2.79 |

How long a test needs to detect a given change:

| change_to_detect_pp | pairs_per_arm | weeks_of_allocation |
|---|---|---|
| 1 | 13,359 | 57.4 |
| 2 | 3,498 | 15 |
| 3 | 1,624 | 7 |

## 5. What each lever is worth — what-if arithmetic, not causal estimates

Arithmetic on the tables above — a group's pass rate moved to another group's — not causal estimates: assessment length is confounded with project (§2.2, twelve projects with one length each) and the device gap is observational. Extra passes are carried to finished at the contract stage's completion rate, 88.0%, and compared with the 4,707 pairs that finish today. The two assessment levers overlap (mobile pairs sit in long projects too), so they do not add.

| lever | pairs it touches | rate now | rate if | extra passes | extra finished | % of finished now | worth in acceptance points |
|---|---|---|---|---|---|---|---|
| the 6 projects at 40 minutes or more pass at the short group's rate | 7,183 | 30 | 36.8 | 487 | 429 | 9.1 | 3.6 |
| mobile passes at desktop's rate | 9,865 | 31.1 | 38.9 | 770 | 677 | 14.4 | 5.7 |

One point of acceptance at the invitation is 436 more accepted pairs and, carried through every later stage at its observed rate, about 119 more finished, 2.5% of today's. So an invitation experiment would have to lift acceptance by 5.7 points to match the mobile lever alone; who is invited moves it by 3.4 points at most (country, §2.1; device is recorded only once the tasker acts, so it is not known at allocation), and nothing in these tables measures what a change of channel would do.
