Separate newly authored text from copied history so summaries and classifiers do not assign an old statement to the current sender.
Email threads: quote boundaries and authored-text attribution
A thread is a set of messages, not one document
An email reply often embeds earlier messages. If the whole body is treated as the current sender’s words, a classifier can count the same complaint several times or attribute a previous promise to the wrong person. Preserve message ID, parent ID, sender, recipient, timestamp and MIME part before text analysis. Quote depth and forwarded-message markers are evidence about provenance, not the business content itself. Mention identity helps when replies use “that change” or “it” without repeating the subject.
Separate new text conservatively
Plain-text replies may prefix copied lines with a greater-than marker; HTML mail can use blockquote markup, tables or client-specific wrappers. Some users answer inline inside an old quote. A parser should mark spans as authored, quoted, signature or uncertain rather than delete all text after the first quote. Keep source offsets and a link to the original message. The small example below handles only plain-text line prefixes and intentionally leaves other formats for review. Section boundaries provide a related document-segmentation pattern.
Attribute claims to original senders
A copied sentence remains evidence of the earlier sender’s statement. The current sender may endorse it, reject it or merely include it for context. Those are separate speech acts. Resolve attribution using message lineage, not position in a flattened string. Quoted text can be useful for retrieval, but it should not count as a fresh commitment. Commitment boundaries distinguish a request from a promise.
Test messy mail
Evaluate top-posted, bottom-posted and inline replies; forwarded messages; signatures; mobile footers; HTML-only bodies; and broken parent headers. Measure authored-span precision, quote recall, speaker attribution and downstream decision accuracy. A high quote-removal score can still fail if the one fresh line inside an old block is erased. The correction lesson consumes these attributed messages, and the project audits the final state.
Implementation
def plain_text_authorship(body):
segments = []
for line_number, line in enumerate(body.splitlines(), start=1):
stripped = line.lstrip()
if not stripped:
role = "blank"
elif stripped.startswith(">"):
role = "quoted"
else:
role = "authored-or-uncertain"
segments.append({"line": line_number, "role": role, "text": line})
return segments
reply = "The window is now 03:47 UTC.\n> The window is 02:47 UTC."
parts = plain_text_authorship(reply)
assert [part["role"] for part in parts] == [
"authored-or-uncertain", "quoted"]
assert parts[1]["text"].startswith(">")
Performance and operating cost
Scanning n characters takes O(n) time and O(n) output space because the code retains each line. It is a deliberately limited plain-text heuristic. MIME parsing, HTML quote wrappers, inline replies and signatures require separate rules or review; a non-quoted line is not automatically authenticated as a new statement.
Common Mistakes
- Assigning all copied text to the latest sender.
- Deleting everything after the first quote marker.
- Ignoring MIME parts and message IDs.
- Treating a forwarded claim as a fresh commitment.
