Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Question versions, order effects and comparable trends

Last updated: 5 Oct 20265 min read
tutorial
IntermediateBy AITrove Editorial

Changing an item, its position or the survey mode can change answers without changing the underlying experience.

Treat the form as an instrument

A handoff-clarity item may follow a complaint question in one version and a factual case-date question in another. Earlier questions can frame what respondents consider when answering. A mobile form may also show response options differently from a phone script. Store item text, option labels, order, device or mode, and release version for each response. Instrumentation changes can break a trend just as event-schema changes can.

Design a split before release

When a revision is needed, randomly assign eligible invitations within the same field period to form A or B. Keep the case-selection process, reminder policy and outcome window aligned. For an example pilot, 16 responses to A include six top-two answers and 15 responses to B include three. These are observed shares of 6/16 and 3/15 among respondents, not proof of a wording effect or a population trend. Assignment integrity and response patterns must be checked first.

Check exposure and response

Report invitations by assigned form, delivered invitations, starts, completed responses and substantive item answers. Random assignment at invitation does not guarantee that respondent subsets remain comparable if one version discourages completion. Compare response rates and known pre-invitation case attributes by form. Follow-up analysis addresses who did not answer; the split alone does not remove response selection.

Avoid a naive splice

Do not join old-form and new-form percentages into one time series and call the difference a service change. If a randomized overlap period supports a defensible bridge, show both raw series and the adjustment with uncertainty. If no overlap exists, mark a measurement break. A correlation between forms does not show that individual answers are interchangeable, especially when response categories changed.

Preserve the comparison packet

Save randomization unit, assignment seed or system log, field dates, question versions, codebooks, denominator counts and analysis code. Predeclare the primary outcome and what difference would trigger keeping the old item, revising the new one or treating the series as broken. A reproducible snapshot lets an analyst inspect the exact version each respondent saw.

Implementation

python
def form_response_funnel(invited_by_form, completed_by_form):
    report = {}
    for form_version, invited in invited_by_form.items():
        completed = completed_by_form.get(form_version, 0)
        if invited <= 0 or not 0 <= completed <= invited:
            raise ValueError("invalid invitation or completion count")
        report[form_version] = {"invited": invited, "completed": completed,
                                "response_share": completed / invited}
    if set(completed_by_form) - set(invited_by_form):
        raise ValueError("completion without a known form assignment")
    return report

funnel = form_response_funnel({"A": 27, "B": 26}, {"A": 16, "B": 15})
assert funnel["A"]["completed"] + funnel["B"]["completed"] == 31

Performance and operating cost

For F form versions, count reconciliation is O(F) time and output space. Assignment and outcome analysis over N invitees is O(N). A credible comparison depends more on exposure and response logs than on computing two percentages.

Common Mistakes

  • Do not infer a wording effect from respondent-only percentages after unequal nonresponse.
  • Do not splice item versions into one trend without an overlap study or a marked break.
  • Do not randomize form versions after seeing a respondent’s answers.

Read next

ai-data
data-science
Storage details