Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Response options and ordinal coding without false precision

Last updated: 5 Oct 20265 min read
tutorial
IntermediateBy AITrove Editorial

Response categories define what a survey result can mean; preserve their order, labels and special states before analysis.

Match categories to the question

For handoff clarity, a five-category scale can run from “not at all clear” through “completely clear.” Each label should refer to clarity, not switch midway to satisfaction or trust. Offer “I do not remember” and “no handoff occurred” only when those states are possible and meaningful. They are not the midpoint. A respondent who never experienced a handoff may be ineligible for the item rather than a low-clarity answer.

Keep non-substantive states distinct

A blank response, a deliberate skip, a displayed “do not remember” choice and a branch that never displayed the item are four separate data states. Encoding all as zero creates a fake low category. Build a codebook with stable string IDs and display labels. If labels are translated, retain the source version and language so category mapping can be audited. Missingness mechanisms still apply to survey items.

Respect ordinal data

The order from unclear to clear is meaningful, but the distance between adjacent labels is not guaranteed equal. Report category shares and a predeclared top-two share before treating numeric scores as interval measurements. If a mean is used for operational convenience, say that equal spacing is an assumption and show whether the decision changes under category shares. A shift from mostly clear to completely clear need not have the same magnitude as a shift from unclear to somewhat clear.

Preserve the denominator

Suppose 31 respondents completed the survey, but only 27 saw the handoff item and 23 gave a substantive clarity answer. A top-two count of 17 is 17/23 among substantive answers, not 17/31 among all respondents or 17/53 among invitees. Show each denominator and why four item displays lack a substantive answer. Denominator contracts keep the comparison reviewable.

Version category changes

Adding a neutral midpoint, changing its label or moving a “do not know” option can alter the response distribution. Do not silently merge old and new data. Retain raw category IDs, make a documented crosswalk if defensible, and display distributions by form version. Version checks separate a measurement change from a service change.

Implementation

python
from collections import Counter

def clarity_distribution(answer_ids):
    substantive = {"not_clear", "slightly_clear", "moderately_clear",
                   "mostly_clear", "completely_clear"}
    special = {"do_not_remember", "not_applicable", "skipped", "not_displayed"}
    counts = Counter(answer_ids)
    unknown = set(counts) - substantive - special
    if unknown:
        raise ValueError(f"unrecognized answers: {sorted(unknown)}")
    answered = sum(counts[answer] for answer in substantive)
    top_two = counts["mostly_clear"] + counts["completely_clear"]
    return {"answered": answered, "top_two": top_two,
            "top_two_share": top_two / answered if answered else None,
            "special_counts": {answer: counts[answer] for answer in special}}

result = clarity_distribution(["mostly_clear", "completely_clear",
                               "do_not_remember", "not_displayed"])
assert result["answered"] == 2 and result["top_two_share"] == 1.0

Performance and operating cost

Counting N answers uses O(N) time and O(K) space for K categories. The fixed category set makes K bounded. Keep raw records and version metadata; an aggregate alone cannot reveal whether a branch or label changed.

Common Mistakes

  • Do not code “do not remember” as a neutral clarity score.
  • Do not divide a substantive-answer count by all invitations without labeling the result.
  • Do not assume equal distances between ordinal response categories.

Read next

ai-data
data-science
Storage details