This summer, 30,025 people were invited to 53,935 project onboardings.
Of the 43,646 old enough to judge, 4,707 finished.
Two doors.
The invitation, which 26,351 never answer. The assessment, which 10,499 of the 15,848 who reach it never pass.
Fix the second door first. It is the one we control.
What follows is how the data got there.
315,269 rows. Seventeen event types. One summer.
A loader that asserts the schema, the timestamp format and the event vocabulary did not raise once.
the data profile3,121 rows were exact duplicates of an earlier row.
Set aside under their reason and kept in the file. The alternative, keeping them, is re-run in the sensitivity set and moves no headline figure.
the cleaning rulesThe device column arrived in nine spellings.
Folded into three classes: mobile, desktop, tablet. A bare “web” counts as desktop, the historical meaning of the label; the case-fold-only alternative is in the sensitivity set.
Nothing was deleted. Every rule has a count; every choice has a counted alternative.
One person can be invited to several projects: 53,935 pairs, 30,025 people.
The unit of every number is the pair of tasker and project. Attributes are constant within a pair, so the funnel is counted on pairs.
the data profilePairs invited in the last week of August finished 5.5%. Pairs invited in June, 11.6%.
Mostly not because August’s taskers differ: they accept at 37.3% against June’s 39.9%. The export stops on August 31, and the gap opens in the stages the late pairs have not had time to reach.
So the rates are read only on pairs allocated at least 14 days before the export closed.
43,646 of 53,935. Re-run at 7 and 21 days, the top stage under each lens is the same and the finished rate moves by about a tenth of a point.
the funnel tablesCensoring is not a drop-off. It is the calendar.
Lens two, the accepted: the assessment loses 10,499 of the 15,848 who sign the confidentiality agreement.
33.8% pass, a median 44.467 hours from signing to passing.
Confidentiality completes 91.6%. The contract, 88%.
Not where the problem is.
Rank under both lenses, name the top of each, then ask which one we can act on.
No country, project, week, prior-work or tenure segment moves acceptance by more than 3.4 points.
Country 38–41.4%. Project 38.6–41.3%. Week 37.8–40.5%.
the analysisDevice appears to: 57.6 points between mobile and “unknown”.
It is an artefact.
A device is recorded only on events the tasker performs, so “unknown” mostly means “never engaged”: 20,295 pairs, accepting 9.6%. One project, Quartz, records no device at all, and 1,950 of the 1,951 unknown-device acceptances are Quartz.
Of the 26,351 who never accepted, 15,175 were delivered an email with no open recorded.
An open is a tracking signal, not a step: 1,272 pairs clicked without one. Where the loss sits, not a cause we have measured.
Acceptance does not vary with who the tasker is. The loss sits in the channel.
The six projects whose assessments are estimated at 30 minutes or less pass 33.9–38.7% of the pairs that reach the assessment.
the analysisThe six at 40 minutes or more pass 28.7–30.8%.
Every short project beats every long one: the lowest short, Vesper at 33.9%, sits above the highest long, Tundra at 30.8%.
No project sits between 30 and 40 minutes, so the data cannot place the line more finely.
And estimated minutes is a project attribute: with twelve projects, length is confounded with everything else that differs by project. A direction, not a causal size.
Mobile passes 31.1%. Desktop, 38.9%.
Within projects the gap holds wherever both cells have 200 allocations, and reverses only in Vesper.
There is no assessment-failed event.
So “attempted, not passed” mixes failures with grades still pending: grading takes up to 54.572 hours at the 90th percentile.
the funnel tablesAnd the pass rate drifts over the summer.
From 34.9% in the week of June 1 to 32.6% in the week of August 10: consistent with censoring inside this slow stage for the latest pairs, though the data cannot separate that from a real change.
the analysisLength and the phone are the levers. Both are ours to pull.
If the six long assessments passed at the short group’s rate, 36.8% instead of 30%, about 487 more pairs would pass and about 429 more would finish: 9.1% more than today.
Arithmetic on the tables, not causal estimates: length is confounded with project, and the device gap is observational. The extra passes are carried to finished at the contract’s 88%.
the analysisIf mobile passed at desktop’s rate, 38.9% instead of 31.1%: about 770 more passes, 677 more finished, 14.4%.
The device gap is observational, and the two levers overlap: mobile pairs sit in the long projects too, so they do not add.
Each point of acceptance at the invitation is worth about 119 finished, 2.5%.
So an invitation experiment would have to lift acceptance by 5.7 points to match the mobile lever alone. Who is invited moves it by 3.4 points at most; what a change of channel would do, nothing here measures.
That is the order. The assessment first; the invitation as the cheaper second experiment.
Three projects piloted two positions, each split roughly in half.
Hypotheses and verdict rules were written before the comparison was read. Balance: 0 of 21 checks differ by 5 points or more. Randomization is assumed; balance is consistent with it.
the analysis · assumptionsCompletion is not distinguishable: pooled −0.96 points, interval −2.11 to 0.19.
Lumen alone shows a clear cost to the early slot: −2.97, interval −4.85 to −1.09.
Time is: the late slot finishes 17.068 hours slower at the median, interval 10.062 to 23.591. In every project.
Capacity a customer is waiting on.
232 declines at the early slot. 36 after an assessment was already passed.
The early slot surfaces refusals before the work they would waste.
Place it early. Right after acceptance.
The metric: finished onboarding per allocated pair, read 14 days after allocation.
Guardrails: median time to finish, and the first-task-submitted rate, so a shorter assessment that admits weaker taskers shows up.
the analysisRandomize at allocation inside the six long-assessment projects: 1,889 allocations a week at a 9.43% baseline.
Two weeks detect a 2.83-point change; four weeks 1.97; eight weeks 1.37. A two-point change needs 3,665 per arm, about 3.9 weeks.
Every number here regenerates from the raw file.
Three commands rebuild the profile, the rules and the funnel. Goldens pin every count, verdict and interval. The dashboard recomputes the headline in the browser and matches the pipeline exactly. Eight sensitivity runs each swap one rule; none moves the answer.
the funnel tables · the findings · the dashboardThresholds were fixed before the numbers. The numbers were allowed to disagree.
The invitation loses more people.
The assessment is the door we can change.
Fix the assessment first.
Each number above appears, as a whole number, in one of the documents below, which were regenerated from the raw export and checked on 8 October 2026. The links open the documents beside this page.
What this data cannot say, in one place. Why an invitation email is not opened: there are no reminder, bounce or expiry events, and open tracking is itself blind for the 1,272 pairs that clicked without an open; nor why 672 of the judged pairs were never sent an email at all, or whether their 1,587 undelivered emails are an address problem or a sending one. Whether an attempted assessment failed or is still being graded: there is no failed event, so attempted-not-passed mixes the two. The size of the length effect free of project confounding: twelve projects, one length each, and no project between 30 and 40 minutes; the 40-minute split the measurement plan runs on was read off the segmentation, not fixed beforehand. Device for Quartz, which records none, and for anyone who never acted. The pilot: randomization is assumed and balance is consistent with it; only two positions were tried, on three July projects with 45-, 55- and 30-minute assessments, so a 15-minute project might behave differently, and the pooled estimate averages effects that may genuinely differ by project. If Lumen’s −2.97 is the true effect on video projects and the other two are noise, the early slot would cost about three points, which is why the next video project should run the measurement plan before the placement is treated as settled. Anything after August 31: every rate is read on pairs allocated at least 14 days before the export closed, and the latest of those are still slightly censored inside the assessment stage. Cost: there is no price of an assessment, a grader hour or a day of customer waiting, so “matters most” is ranked by people and time, not money.