# Findings — where we lose people, why, and where the likeness agreement should go

*Hand-written. Every number below appears in a generated document (`05`–`08`); `python -m pso
cite` checks that, and its limits are stated in the docstring of `analysis/pso/cite.py` and in the README. The rules behind the numbers are in
`06`, the assumptions in `ASSUMPTIONS.md`. The unit throughout is an **allocation** — one tasker invited to
one project — not a person: 30,025 taskers account for 53,935 allocations.*

## The answer in four sentences

1. Of the 43,646 allocations made early enough to be judged (14 days before the export
   closed), 4,707 finished onboarding — 10.78%. Two stages account for almost all of the loss:
   the **invitation**, which 26,351 allocations never accept, and the **qualification
   assessment**, which 10,499 of the 15,848 who sign the confidentiality agreement never pass.
2. The invitation loss is uniform — no project, country, tenure, week or prior-work segment
   moves acceptance by more than 3.4 points — and most of it is associated with email that has
   no recorded open. The assessment loss is not uniform: the six projects whose assessments
   are estimated at 40 minutes or more pass 28.7–30.8% of candidates, the six at 30 minutes
   or less pass 33.9–38.7%, and mobile candidates pass 31.1% against 38.9% on desktop.
3. **Recommendation:** fix the assessment first — it is the loss with a lever we control.
   Bring long assessments down into the shorter group (the data can only place the line
   between 30 and 40 minutes: no project sits in between) and make them work on a phone. Run
   the invitation as a cheaper second experiment: in-app invitations with a reminder,
   measured on acceptance.
4. **Likeness agreement:** place it **early, right after acceptance**. End-to-end completion
   is not distinguishable between the two piloted positions (pooled −0.96 points, interval
   −2.11 to 0.19); the late position makes finishing slower in every project (pooled
   17.068 hours longer at the median, interval 10.062 to 23.591); and the early position
   surfaces refusals before any assessment is spent on them (232 declines at the early slot
   against 36 after passing the assessment).

## Q2 — where are we losing people, and why

### The two lenses give two different top stages, on purpose

The brief counts the invitation as part of onboarding, so the first lens is every allocation:
the Invitation stage loses 26,351 of 43,646 (39.6% accept). The second lens is allocations that
were accepted and therefore showed intent: there the Qualification assessment loses 10,499 of
15,848 (33.8% pass), with a median 44.467 hours from signing the agreement to passing.
Confidentiality (91.6% complete) and the contract (88% complete) are not where the problem is.

### Invitation: large, uniform, and about the channel

Of 43,646 allocations, 42,974 were emailed, 41,387 delivered, 24,096 opened, 18,381 clicked
and 17,295 accepted. Among the 26,351 that never accepted, 672 were never sent an email,
1,587 were sent one that was never delivered, 15,175 were delivered with no open recorded, and
8,917 opened and still did not accept.
Opens are a tracking signal, not a step — 1,272 allocations clicked an email with no open
recorded (R7) — so "no open recorded" is where most of the non-acceptance sits, not a cause
we have measured.

Segmenting acceptance by every attribute known at allocation shows almost nothing: country
ranges 38–41.4%, project 38.6–41.3%, allocation week 37.8–40.5%, prior projects 39–40.1%,
tenure 39.5–39.7%. The one dimension that appears to matter — device, at 57.6 points — is an
artefact and must not be read as a driver, for two reasons. A device is recorded only on
events the tasker performs, so for most projects "unknown device" (20,295 allocations, 9.6%
acceptance) means "never engaged". And one project, Quartz, records no device on any event:
its 4,719 allocations are all "unknown", and 1,950 of the 1,951 unknown-device acceptances are
Quartz. Among allocations with a recorded device, mobile accepts 67.2% and desktop 62.3%.

What this supports: acceptance does not vary with who the tasker is. What it does not support:
any claim about *why* an email is not opened — the data has no reminder, no bounce reason and
no expiry. The in-app channel is small: 1,288 of the 15,848 who signed the agreement had
accepted in-app, and all 1,701 in-app acceptances in the full data never click the invitation
email.

### Assessment: smaller, concentrated, and about the assessment

Of 15,848 who signed the agreement, 13,623 started the assessment, 8,061 submitted an attempt
and 5,349 passed. Two things move this stage:

- **Project, and project here tracks assessment length.** Pass rates run from Lumen 28.7% to
  Sequoia 38.7%, and they split in two by the project's estimated assessment time. Thirty
  minutes or less: Meridian (15 min) 38.6%, Granite (20) 37.5%, Juniper (20) 36.1%, Sequoia
  (25) 38.7%, Quartz (30) 36.1%, Vesper (30) 33.9%. Forty minutes or more: Beacon (40) 29.9%,
  Cobalt (45) 30%, Lumen (45) 28.7%, Tundra (50) 30.8%, Orchid (55) 30%, Halcyon (60) 30.6%.
  It is a step, not a smooth gradient — Sequoia at 25 minutes beats Meridian at 15, Halcyon at
  60 beats Beacon at 40 — but the lowest short-assessment project (Vesper, 33.9%) sits above
  the highest long-assessment project (Tundra, 30.8%), and with the three pilot projects set
  aside the comparison is 36.1% against 30.8%.
- **Device.** Desktop passes 38.9%, mobile 31.1%, tablet 38.8%. Within projects the gap holds
  in most of them where both cells have 200 allocations — Orchid 26.4% mobile against 38.9%
  desktop, Tundra 26.7% against 40.3%, Cobalt 27.5% against 35.8% — and reverses only in
  Vesper (34.4% against 33.3%). It is not the length effect in disguise.

Weeks and countries move it by 7 points or less, tenure and prior work by under 2, and the
acceptance channel not at all (33.8% vs 33.7%).

What this supports: the assessment's length and its usability on a phone are the levers. What
it does not support: a causal size. Estimated minutes is a project attribute, so with 12
projects the length effect is confounded with everything else that differs by project; the data
has no assessment-failed event, so "attempted but not passed" mixes failures with pending
grades (grading takes up to 54.572 hours at the 90th percentile); and the pass rate drifts
down from 34.9% in the first full week (2026-06-01) to 32.6% in the last (2026-08-10), which is
consistent with residual censoring inside this two-day stage for the latest cohorts, though
the data cannot separate that from a real change over the summer.

### Recommendation and how we would know

Shorten or split the long assessments and make them pass on mobile. The metric is the
finished-onboarding rate per allocation, read 14 days after allocation; the guardrails are
median allocation-to-finish time and the first-task-submitted rate, so a shorter assessment
that admits weaker taskers shows up. Randomize at allocation within project, exactly as the
likeness pilot was run. The change only applies to the six long-assessment projects, which
made 1,889 allocations per week on average over the summer (recent weeks run higher, so these
lengths are conservative) at a 9.43% baseline: two weeks detect a 2.83-point change in the
finished rate, four weeks 1.97, eight weeks 1.37; a two-point change needs 3,665 allocations
per arm, about 3.9 weeks.

What the levers are worth, as arithmetic on the tables rather than causal estimates (`08` §5):
if the six long-assessment projects passed at the short group's rate, 36.8% instead of 30%,
about 487 more allocations would pass and, at the contract stage's 88% completion, about 429
more would finish — 9.1% more than today's 4,707. If mobile passed at desktop's rate, 38.9%
instead of 31.1%, about 770 more would pass and 677 more would finish, 14.4%. The two overlap
(mobile allocations sit in the long projects too), so they do not add. One point of acceptance
at the invitation is 436 more accepted allocations and, carried through every later stage,
about 119 more finished, 2.5% — so an invitation experiment would have to lift acceptance by
5.7 points to match the mobile lever alone. Who is invited moves it by 3.4 points at most
(country); what a change of channel would do, nothing in the tables measures.

For the invitation: in-app invitations plus one reminder, randomized the same way across all
projects (4,097 allocations per week on average), measured on acceptance first — the effect,
if any, is large and fast to see — and on the finished rate second, where two weeks detect a
2-point change.

## Q3 — where should the likeness agreement go

The pilot ran on three projects from July, each split between an **early** slot (after
acceptance, before the confidentiality agreement) and a **late** slot (after passing the
assessment, before the contract). The arms look balanced: 0 of 21 checks of tenure, country,
device, prior work, concurrency, acceptance channel and week differ by 5.0 points or more; the
largest gap is 3 points. Randomization itself is assumed (A10); balance is consistent with it.

**Completion.** Early finishes 8.41% of mature allocations, late 9.36%; the pooled difference
is −0.96 points with a 95% interval of −2.11 to 0.19. By project: Lumen early 6.56% against
late 9.53% (−2.97, interval −4.85 to −1.09 — the one clear result); Orchid 9.66% against 8.28%
(1.38, −0.64 to 3.41); Vesper 9.11% against 10.3% (−1.19, −3.28 to 0.89). Under the
pre-registered reading, H1 (early finishes at least as many, and no worse after the slot) is
inconclusive pooled and not supported in Lumen; H2 (late finishes more, without being
materially slower) is inconclusive pooled — and in Lumen the late arm's completion advantage is
real, but H2 still fails there on time.

**Where the losses fall.** The early slot costs at the confidentiality stage: 81.5% complete it
in the early arm against 92.5% in the late arm, and 232 of the early arm's allocations decline
the agreement there. After the slot the arms are not distinguishable: assessment pass rates
31% against 30.7% (0.3 points, interval −2.9 to 3.5), contract 86.9% against 82.8% (4.2
points, −0.3 to 8.6; the late arm's 36 declines land in that stage). The early slot therefore
does what it should — it removes allocations that will not consent before an assessment is
spent on them — but it does not raise the pass rate of those who stay. Consent at the slot
(87.0% early, 93.3% late) is conditioned on who reaches it and is not evidence either way.

**Time.** The late slot makes finishing slower. Among allocations that finished, median
allocation-to-finish is 64.032 hours in the early arm and 81.1 in the late arm — 17.068 hours
more, bootstrap interval 10.062 to 23.591 hours. The direction holds in every project (Lumen
14.341 hours, interval 0.076 to 32.138; Orchid 18.281, 6.646 to 32.983; Vesper 11.734, 4.584 to
23.631); two of the three exceed the pre-registered 12.0-hour threshold and Vesper sits just
inside it. That is the brief's own cost — capacity a customer is waiting on — and it is the
one effect in the pilot whose interval excludes zero in all three projects.

**Recommendation: early.** On the pooled interval the early slot costs between 2.11 points and
nothing (it may even gain 0.19); the late slot costs 17.068 hours of median time to finish and
an assessment for every allocation that then declines. If Lumen's −2.97 is the true effect on
video projects and the other two are noise, the early slot would cost about three points —
so run the measurement plan on the next video project, knowing what it can see: a single
pilot-type project made about 465 allocations per week on average over its full weeks, so two
weeks detect only a 5.93-point change and eight weeks 2.79; a 2-point difference needs 3,498
allocations per arm, about 15 weeks of one project, or proportionally less across several.

**What the pilot cannot say.** Only two positions were tried; a slot after the confidentiality
agreement but before the assessment, or consent folded into the contract itself, was never
run. The three projects are all July allocations with 45-, 55- and 30-minute assessments; a
15-minute project might behave differently. And the mixed per-project results mean the pooled
estimate is an average of effects that may genuinely differ by project.

## What this data cannot tell us, in one place

- Why an invitation email is not opened (no reminder, bounce or expiry events; open tracking
  is itself blind for 1,272 clickers), why 672 of the mature allocations were never sent an
  email at all, or whether their 1,587 undelivered emails are an address problem or a sending
  one.
- Whether an attempted assessment failed or is still being graded (no failed event).
- The size of the length effect free of project confounding (12 projects, one length each),
  or where between 30 and 40 minutes the step actually is.
- Device for Quartz (none recorded) and for anyone who never acted.
- Anything after August 31: the export stops at the window's end, so every rate is on
  allocations made by August 17, and the latest of those are still slightly censored in the
  assessment stage.
- Cost: there is no price of an assessment, a grader hour or a day of customer waiting, so
  "matters most" is ranked by people and time, not money.
