Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Vision dataset contract: image unit, label definition and consent

Last updated: 5 Oct 20265 min read
tutorial
IntermediateBy AITrove Editorial

A vision task begins by defining what one image represents, what a label means and whether that image may be retained for the task.

Name the operational decision

A receipt-quality classifier could flag images that require manual recapture because text is unreadable. That is different from predicting fraud or extracting line items. Define the capture channel, image format, label unit and action before collecting examples. A multi-page receipt may be one submission with several images; splitting pages across train and test would leak the same submission. Dataset grain must match the decision.

Write a label rubric

“Unreadable” needs observable criteria: perhaps the merchant and total cannot both be verified by a trained reviewer at normal zoom. Borderline glare, partial crops and unfamiliar scripts need examples in the rubric. Record an uncertainty or adjudication state instead of forcing every case into a confident binary label. Annotation quality is measurable only after the rubric is stable.

Constrain image use

A receipt image can contain names, addresses and payment fragments. Store the minimum necessary image and metadata, define access and retention, and prevent a training export from bypassing deletion requests. A person’s consent to process a receipt for reimbursement does not automatically authorize every unrelated model experiment. Keep raw-image and derived-feature lifecycles linked.

Test the manifest

Create two images from one receipt, one blurred image and one image whose label is unknown. Verify that the two pages share a submission group, the unknown label does not become a negative, and deleted images cannot reappear through a cached training set. Count eligible submissions as well as image files.

Implementation

python
def validate_image_record(record):
    required = {"image_id", "submission_id", "capture_time", "consent_scope", "label"}
    if required - record.keys():
        raise ValueError("missing image fields")
    if record["label"] not in {"readable", "unreadable", "uncertain"}:
        raise ValueError("unknown label state")
    if "receipt_quality" not in record["consent_scope"]:
        raise ValueError("image is outside consent scope")
    return record["submission_id"]

Performance and operating cost

Validation is O(1) per record and O(N) over N images; retaining group IDs for a split costs O(G) memory for G submissions. Image storage and annotation review dominate the arithmetic cost, so record licensing, retention and label decisions before scaling capture.

Common Mistakes

  • Do not mix pages from one submission across evaluation partitions.
  • Do not convert an uncertain label into a negative.
  • Do not assume operational upload consent covers unrelated training.

Read next

Continue the workflow: Multimodal data contract: paired records, identity and consent.

Continue the workflow: Pretrained encoder and target-task contract.

ai-data
computer-vision
Storage details