# Assumption log

The brief: "If something is ambiguous, make a reasonable assumption, note it, and keep going.
We'll ask about your assumptions." This is that list.

Numbering is permanent — never renumber. Status is one of **open** (a claim about the data not
yet checked), **verified** (checked against the data, with the evidence cited), **adopted** (a
definition or policy I chose; it cannot be verified, only kept consistent and given an
alternative), **revised** (link to the superseding entry), or **withdrawn**. Every cleaning rule
and every definition in `analysis/` has an entry here before it is applied. Section numbers (§)
refer to the generated `05-data-profile.md`; rule numbers (R) to the generated
`06-cleaning-rules.md`.

## Assumptions

| # | Assumption | Basis | Status | Checked how |
|---|---|---|---|---|
| A1 | The PSO steps occur in the order the brief lists them: invitation → confidentiality agreement → qualification assessment → work contract. The brief lists them; it does not state an order. | Brief, "Background". | **verified** 2026-10-08 | §10: for every consecutive pair of events, 0 later-before-earlier and 0 later-without-earlier except the optional email steps (A14). The data has 17 events; the brief's four map onto them as in the sub-step tables of `07`. The likeness event's position is per pilot arm (§13). |
| A2 | "Finishing PSO" means completing the last PSO step present for that project — the work contract, or the likeness agreement where a pilot project placed it after the contract. Paid work beginning is the outcome, not a PSO step. Depends on A1. | Brief, "Background": PSO happens "before a tasker can start paid work". | **verified** 2026-10-08 | §10–11: `pso_completed` is an explicit finish marker and marks exactly the `contract_finished` pairs (5,503). Paid-work events (`first_task_claimed`, `first_task_submitted`) exist and follow it. No pilot arm placed the likeness step after the contract (§13), so that clause is moot. |
| A3 | Taskers invited near the end of the window (late August) are **right-censored** — the export on 2026-09-01 may predate their later steps — so late cohorts' completion is understated unless a maturity window is applied. | Brief, "The data": invites to Aug 31, extracted Sep 1. | **verified, severe** 2026-10-08 | §4, §14: the last event is 2026-08-31 23:59:48 and nothing from September is present; pairs allocated Aug 25–31 reach finished PSO at 5.5% vs 11.6% for June. Handled by R3. |
| A4 | Every anomaly in the export is treated as real pipeline behaviour that needs an explicit, counted handling rule. Whether any anomaly was planted is not knowable and is not assumed. | Brief, "The data": "as it came out of the pipeline"; "Logistics" 3: the data is synthetic. | **adopted** (policy) | §3, §5 list the literal dirt — 3,121 exact duplicate rows (R1) and `device` casing (R2). The semantic anomalies each have a rule with a count in `06`: censoring (R3), never-emailed pairs (R4), the pilot (R5), multi-session assessments (R6), missing email-tracking events (R7, A14). |
| A5 | A tasker can be invited to more than one project; the unit of the funnel is the **(tasker, project)** pair, not the tasker. | Brief: onboarding is "Project-Specific"; taskers are "invited to projects". | **verified** 2026-10-08 | §8: 53,935 pairs, 30,025 taskers, 14,344 in more than one project. Attributes are constant within a pair (`tests/test_funnel.py`). |
| A6 | Time spent in a step is a loss alongside drop-off, because the cost the brief names is "capacity a customer is waiting on", and a slow tasker is capacity the customer is still waiting on. | Brief, "Background". | **adopted** (criterion) | Feasible on this export — §12: per-step latencies from first occurrences; 0 zero-or-negative deltas. |
| A7 | "The drop-off that matters most" is ranked by **people lost at the step** (reached × loss rate), with median **time-in-step** as the second criterion. No cost weighting. | Brief, Q2 and "Background". | **revised → A13** | Superseded by the two-lens criterion in A13. |
| A8 | A step is **reached** when the previous step was completed; the invitation is reached by being invited. A (tasker, project) pair that reached a step and never completed it is **lost at that step**. There is no state between steps. | My definition; needed so that every pair falls on exactly one edge of the funnel. | **adopted** (definition) | Consistent with §10: the only steps with separate started/finished events are the assessment (started → attempted → passed) and the contract (started → finished), and both are modelled as sub-steps. For the optional email steps "reached" means "has the event" (A14). |
| A9 | The evaluator opens `index.html` straight from disk (`file://`), possibly offline. | My conservative reading of "runs locally in a browser" and "share the files needed to run it locally" — the brief does not say offline. | **adopted** (design) | Not checkable. It is why the data ships as `dashboard/data.js` and not a fetched JSON: a `file://` page cannot `fetch` a sibling file. |
| A10 | The pilot's `early` / `late` arms were assigned at random within each project. | §13: both arms are present in each of the three pilot projects at roughly 50/50; the arm is constant per pair, not per project. That is consistent with randomization and does not prove it. | **open** — balance checked 2026-10-08 | `08` §1: 0 of 21 project × dimension checks differ by 5 pp or more (largest gap 3.0 pp). Balance is consistent with randomization and is not proof of it, so every doc keeps saying "assumed randomized". |
| A11 | "Invitation" = `pso_allocated`, the event every pair has first; the Invitation stage's denominator is every allocated pair. | §10–11: present for 100% of pairs, always first. | **adopted** (definition, R4) | R4's alternative — `email_sent` as anchor; the 853 never-emailed pairs treated as a system failure — runs in the sensitivity set. |
| A12 | `concurrent_psos` is an as-shipped attribute whose derivation is unknown; it is used only as a categorical and never recomputed. | During exploration no overlap-based definition reproduced it exactly; that exploration is not regenerated by `05`, so no figure from it is cited. | **adopted** (policy) | Not checkable from this file. |
| A13 | Supersedes A7. "The drop-off that matters most" is ranked under **two lenses** — people lost at each stage over **all allocated pairs** (the brief's own definition of PSO, which includes the invitation) and over **accepted pairs** (taskers who engaged) — with median time-in-step as the second criterion in each. One recommendation names the top of each lens and argues which is actionable. No cost weighting: neither the brief nor the data carries cost. | Brief Q2 and "Background". | **adopted** (criterion) | Stated before any stage count is computed. |
| A14 | `email_opened` and `email_clicked` are **tracking signals, not gates**. Opened can be missing even when the email was clicked (consistent with a blocked tracking pixel; the data cannot say why); clicked is bypassed entirely by in-app acceptance. Both are reported as "has the event" and never have losses attributed to them. | §10: 1,272 pairs clicked without an opened event. R7 (`06`): 1,113 of them accepted by email; the pairs that accepted without a click are exactly the in-app acceptors (set difference 0). | **adopted** (definition, R7) | R7's alternative imputes `opened` for pairs that clicked or accepted by email; the sensitivity table in `07` prints the opened-pair count and the opened → clicked latency it moves, next to its headline verdict. |
| A15 | `device` is classified as mobile = {mobile, ios, android}, desktop = {desktop, web}, tablet = {tablet, ipad}; a pair's device is the first non-empty class in time order (system events carry none). | §5: nine raw spellings; `web` could be either class — desktop was chosen because it is the historical meaning of the label. | **adopted** (definition, R2b) | R2a (case-fold only, seven values) runs in the sensitivity set. |
| A16 | The three pilot projects are part of the aggregate funnel, with the arm available as a filter and likeness losses labelled. | They are a quarter of all pairs (12,347 of 53,935 — R5 in `06`; the six arm sizes in §13 sum to it) and the brief's "where do we lose people" asks about the whole funnel. | **adopted** (definition, R5) | R5's alternative (nine non-pilot projects only) runs in the sensitivity set. |
| A17 | Assessment reach uses the first `assessment_started`; assessment latencies are measured from the first session. | R6 (`06`): 1,479 pairs have more than one session, with the first-to-last session gap printed there. | **adopted** (definition, R6) | R6's alternative (latency from the last session) runs in the sensitivity set; `07` prints the started → attempted median and p90 it moves, next to its headline verdict. |
| A18 | A sensitivity run leaves the headline **robust** when, under each lens, the top-ranked stage is the primary run's and the finished-PSO rate on mature pairs moves by less than one percentage point; **moved** otherwise. Two qualifications: **headline identical** when nothing in the headline differs (observed), and **robust — cannot move the headline answer** for a run whose swapped rule cannot reach either verdict input by construction (only R3, R4 and R5 change which pairs are counted). The verdict judges the headline answer only; the "what it moves" column beside it shows the figures each alternative does change. Every quantile reported is the **lower** quantile, `sorted[floor(q·(n−1))]`, taken on whole seconds and turned into hours by one scalar division; the dashboard's JS implements the same rule; the node test (`tools/tests/engine.test.cjs`) demands equality (maximum relative difference 0), and the badge on the page passes at a 1e-9 relative tolerance and reports the measured maximum. | A verdict rule stated before the runs; the dashboard self-test needs a quantile both implementations compute identically. Polars' column division by 3600 is a reciprocal multiply and is not identical in the last digit, which is why the division moved out of the expression. | **adopted** (criterion) | `07` prints the verdict and moves table; `tests/test_sensitivity.py` pins the threshold at exactly 1.00 pp and the by-construction runs. |
| A19 | `dashboard/data.js` encodes each event as whole seconds after the pair's allocation (`−1` = absent), categoricals as lookup indexes, and `score` as `score × 100` as an integer. Lossless for this export: timestamps are whole seconds (§4) and every score has at most two decimals — enforced by `load.validate()` (`score:more_than_two_decimals` must be 0) and re-checked by the exporter, which raises otherwise. | `05` §4; `load.validate()`; `tests/test_load.py::test_raw_export_scores_have_at_most_two_decimals`. | **adopted** (encoding) | `tests/test_export.py` round-trips a payload through node and refuses a three-decimal score; the committed `data.js` carries the same headline as `out/headline.json`. |
| A20 | **Intervals** are closed-form score intervals: Wilson for a rate; Newcombe's hybrid score interval (method 10) for a difference of two rates; a Mantel–Haenszel-weighted risk difference (weights n₁n₂/N per project) with the Greenland–Robins (1985) variance in its large-strata form for the pooled early − late comparison. Deterministic, no random seed, reproducible in JS. **One exception, added after the first numbers existed (see the note at the end):** a difference of medians (time to finish) has no closed form and gets a percentile bootstrap — each arm's finish times resampled with replacement 2,000 times under a fixed seed (20261008), middle 95% of the differences — deterministic given the seed on one CPython minor version (the project pins 3.12; the generated headers and `data.js` meta record the exact version), and the only random draw in the pipeline; the H2 time guard reads the point estimate, not this interval, so no verdict depends on it. | One method I can explain in a sentence; the dashboard must be able to show the same intervals. | **adopted** (method; bootstrap clause added 2026-10-08, after the first build) | `tests/test_intervals.py` pins known values (Newcombe 1998 worked example; Wilson at 0/10; an unequal-strata MH hand computation; 3,841 per arm) and the bootstrap's determinism. |
| A21 | **Balance check** (A10): per pilot project and dimension (tenure, country, device, prior projects banded 0 / 1–2 / 3+, concurrent PSOs banded 1 / 2+, acceptance source, allocation week), the largest absolute early − late share difference on mature pairs; flagged at ≥ 5 pp. Descriptive — no test statistic. | A 5 pp gap is the smallest a reader would notice in a share table; smaller gaps cannot move a 10% completion rate by a percentage point. | **adopted** (threshold) | `08` §1 prints every check and its flag. |
| A22 | **Segmentation**: declared dimensions are project, allocation cohort week, device class, tenure, prior projects (banded), country, concurrent PSOs (banded), acceptance source; estimated assessment minutes is a project attribute read through *project*; acceptance source cannot segment the Invitation stage. Levels under 200 mature pairs are pooled into "other". *Strength* = range of completion across the remaining levels; the two strongest dimensions get one two-way table whose small cells are flagged, not pooled. | 200 pairs gives a Wilson interval narrower than ±7 pp at the rates seen. | **adopted** (threshold) | `08` §2 lists every slice examined, including the empty ones. |
| A23 | **Hypothesis verdicts** are three-valued from an interval: *supported* when it lies at or above 0, *not supported* when it lies below 0, *inconclusive* when it straddles 0. H1 is *supported* when early − late finished is at or above 0 and no post-slot stage (Assessment, Contract) has the early arm significantly below the late arm (a non-inferiority reading: equal rates straddle 0 and pass); *not supported* when the finished difference or any post-slot stage lies below 0; *inconclusive* otherwise. H2 is *supported* when late − early finished lies at or above 0 **and** the late arm's median allocation → finished is **less than 12 h** longer than the early arm's; *not supported* when the rate difference lies below 0 or the time guard fails with the rate supported; *inconclusive* otherwise. | The hypotheses in this file, operationalised before the comparison was read (one wording change during implementation; see the note at the end); 12 h is half a day of customer waiting against a 64 h median. | **adopted** (threshold) | `08` §3 prints the verdict and the intervals behind it, per project and pooled; `tests/test_analysis.py` pins the boundaries (interval at 0; time at exactly 12 h). |
| A24 | **Measurement plan**: a two-proportion z-test at 80% power and 5% two-sided significance, on the mature finished-PSO rate per allocated pair; horizons of 2, 4 and 8 weeks of the observed weekly allocation volume; effects of 1, 2 and 3 percentage points. **Added after the first numbers existed (see the note at the end):** the plan is computed per population a recommendation would run on — all projects; projects with assessments of 40 minutes or more, a split read off `08` §2.2 (not pre-registered; no project sits between 30 and 40 minutes); one pilot-type project — with each population's weekly volume as the mean over its own full weeks (from its first full Monday-week to 2026-08-24), stated as a mean because projects ramp up and recent weeks run higher, which makes the test lengths conservative. | The textbook design a data scientist states from memory. | **adopted** (settings; population clause added 2026-10-08, after the first build) | `08` §4 prints, per population, the detectable change per horizon and the weeks per effect. |

## Hypotheses

Mechanisms I want to test, not assumptions I rely on. They were written down before the data
was opened; what the data turned out to contain is noted without changing the tests.

**What the pilot actually is (§13).** Three projects — Lumen, Vesper, Orchid, all allocating
from 2026-07-01 — each split roughly 50/50 into an `early` arm (likeness step after acceptance,
before the confidentiality agreement) and a `late` arm (after passing the assessment, before the
contract). Exactly two positions were piloted. Because both arms sit inside every pilot project,
placement is not confounded with project; whether the arms are otherwise comparable is **assumed
until the balance check** (A10).

**The metric the placement decision rests on** is end-to-end PSO completion per allocated
(tasker, project) pair, and time to complete, by arm, compared per project. Consent rate *at the
slot* is conditioned on having survived to that slot: the late arm's slot sees a population that
already passed the earlier steps, so a higher consent rate there is expected under either
hypothesis and is not evidence on its own.

| # | Hypothesis | What would support it (observable) | What would count against it (observable) |
|---|---|---|---|
| H1 | An **early** likeness slot removes taskers who will not consent before the later steps are spent on them. | Within each pilot project: the early arm's end-to-end completion per allocated pair is no lower than the late arm's, and its losses are concentrated at the slot rather than spread over later steps. | The early arm's end-to-end completion is lower than the late arm's by more than the loss at the slot itself — the slot is costing people who would otherwise have finished. |
| H2 | A **late** likeness slot loses fewer people overall, because taskers who have already invested in the earlier steps are less likely to drop at the consent step. | Within each pilot project: the late arm's end-to-end completion per allocated pair is higher than the early arm's, and time to complete is not materially longer. | The late arm's end-to-end completion is no higher, or late placement adds enough time-in-funnel to offset it. |

The comparison is made per project first and then
pooled across the three with project as a stratum, after the balance check (A10).

## Two changes after the first numbers existed

The thresholds and definitions above were written down before the numbers they judge. Two
things changed after a build had already printed numbers; I record them here because
"pre-registered" is only honest with them.

1. **H1's post-slot clause (A23).** As first coded, H1 was *supported* only if every post-slot
   stage was itself *supported* — its interval at or above zero. A fixture with equal post-slot
   rates showed that reading could never pass, because equal rates straddle zero, so it became
   the non-inferiority reading A23 now states: H1 fails only where a post-slot stage lies below
   zero. The first build that printed pilot verdicts ran under the original wording, in the same
   command as the failing test; I changed the wording before reading those verdicts. The verdicts
   are identical under both readings on this export.
2. **The bootstrap interval and the per-population plans (A20, A24).** The seeded percentile
   bootstrap on the difference of median finish times, and the measurement plans computed per
   population — including the "40 minutes or more" population, which is read off the
   segmentation rather than fixed in advance — were added after the first numbers existed.
   Neither changes a verdict: the H2 time guard reads the point estimate, and the plans size
   future tests. Both are marked † in `08`.
