Research protocols · Practical guide
Reply classification study: study protocol
Can two reviewers apply the same positive, negative, referral, automated and ambiguous labels? Use this proposed protocol to collect and interpret evidence.
Reviewed · Examples are illustrative
Who this helps: Teams collecting evidence. Limited software observations are identified separately from unperformed campaign experiments and participant studies.
Define the decision
This page is a proposed research protocol, not a completed study or a report of findings. The decision is: Can two reviewers apply the same positive, negative, referral, automated and ambiguous labels? The observation unit is one first human reply per contact thread. Define the population and owner before collecting records; do not substitute an available convenience dataset without documenting the change.
Work through the procedure
- Draft category definitions, boundary examples and an ambiguous option. Remove unnecessary personal data before annotation.
- Have two reviewers label independently without campaign performance totals. Adjudicate disagreements only after preserving both original labels.
- Before collection, write the primary outcome, observation window, exclusion rules and stopping conditions. Preserve excluded observations with a reason rather than quietly removing them.
- Pilot the procedure with fictional or owned test data, resolve ambiguous fields, and freeze a dated protocol version before the main run.
Worked example
The following is a synthetic example for this procedure, not a customer result or performance benchmark.
reply: ask our operations manager; reviewer A referral; reviewer B positive; resolve the rubric before counting it as sales interest
Suggested record fields: observation_id, condition, evidence_reference, outcome, exclusion_reason, reviewer, protocol_version
Status: illustrative record only; no study has been run for this page.Read the result
Publish the disagreement matrix and adjudication rules. Agreement measures consistency, not whether a reply will become revenue. Keep the numerator, denominator and missing evidence visible. If the available observations cannot answer the registered question, report that limitation rather than selecting a more favorable metric after collection.
Collect the evidence
Observation unit: first-human-reply.
Redacted reply excerpts and two independent human reviewers.
Hide campaign totals and the other reviewer label; retain ambiguous and automated categories.
Analysis: Label disagreement matrix before adjudication; this is not model accuracy.
Download the JSON collection template under Source references. Set an owner, eligibility rules, outcome definition, observation window and sample justification before collecting records. Templates contain no participant data or results; keep original private records outside the public website.
Check before moving on
- Name the person responsible for collection and review.
- Check that the evidence can be inspected without exposing private messages or credentials.
- Record deviations from the protocol and analyze their possible effect.
- Retain a dated, redacted evidence worksheet with the final interpretation.
Limits and next action
Do not force automated replies into a positive or negative category. Keep opt-outs distinct from ordinary declines. Publish results only after the evidence, method and limitations have been reviewed. This protocol provides no benchmark, expected lift or completed-study claim.
Source: Method or workflow reference
Source references
Worked examples are illustrative. Editorial procedures are suggested methods, not measured performance claims or promises of additional product features.
Related guides
- DNS diagnostics: controlled error checks and authentication study protocol →
- Open tracking noise study: study protocol →