Marking quality controls¶
Quality controls keep marking consistent across a project, surface problems before they become incidents, and produce the audit trail you need when a result is queried.
Insights supports three layered controls — validation, double-marking, and adjudication — that you mix to suit the stakes of the project.
Validation¶
Validation responses are pre-marked items with a known, agreed-on score. Insights inserts them into each marker's queue at the rate you configure. The marker can't tell which responses are validation and which are real candidate work.
Validation does three things:
- Calibrates new markers before they affect real results.
- Detects drift — markers becoming more lenient or more harsh over time.
- Triggers retraining — markers who fall outside the agreed tolerance are paused automatically.
Configuring validation¶
In the project's Quality tab:
- Validation rate — what proportion of each marker's queue is validation. 5–15% is typical. Higher rates catch problems faster but slow throughput.
- Tolerance — how far a marker's score can deviate from the validation score before they're flagged. Usually expressed as a number of rubric points (e.g. ±1).
- Action on breach — flag for review, pause marker, or re-train. Pause marker is the safest for high-stakes work.
Where do validation responses come from?
Insights ships a pre-marking workflow — see Author → Pre-marking sets — that lets a panel of expert markers agree on scores for a small set of responses before the main marking begins. Those become the validation pool.
Double-marking¶
Every response is marked twice, by two different markers. If the marks agree (within tolerance), the response is final. If they disagree, it goes to adjudication.
When to enable double-marking¶
- High-stakes items. Constructed-response items worth four marks or more.
- High-stakes sittings. End-of-cycle and certification assessments.
- Quality assurance batches. Sometimes double-marking is enabled for a small sample only, to validate the overall marking process.
Configuring double-marking¶
In the project's Quality tab:
- Double-mark scope — all items, items above N marks, or sample only.
- Sample size — for the sample option, how many responses per marker.
- Tolerance — the maximum acceptable difference between Marker 1 and Marker 2 before adjudication is triggered.
Adjudication¶
When two markers disagree beyond tolerance, the response is routed to an adjudicator — a senior marker with the Adjudicator role.
The adjudicator sees:
- The candidate response.
- The rubric.
- Both marker scores and any annotations they left.
They then either confirm one of the existing marks or enter a new one. The adjudicator's mark is final.
Adjudicator independence
By default, an adjudicator cannot see which markers gave which scores — only the scores themselves. This is to remove any bias toward more experienced or senior colleagues. The setting can be relaxed for training contexts, but should remain on for live marking.
Reading the quality dashboard¶
The Quality tab in any active project shows:
- Marker accuracy — how each marker scores against validation, with a trend line.
- Inter-rater reliability — agreement between markers on the same items.
- Adjudication rate — what proportion of double-marked responses are going to adjudication. A rising rate often indicates a rubric ambiguity that needs clarification.
- Throughput — responses per hour per marker, by item and overall.
| Metric | Healthy range | When to act |
|---|---|---|
| Marker accuracy | 95%+ | If <90%, pause and retrain |
| Adjudication rate | 5–15% | If >20%, re-examine rubric |
| Throughput | Project-specific | Flag markers >30% off the median |
After the project closes¶
Quality data is preserved for two years by default and is included in the audit export for the project. See Report → Quality and audit for the standard quality report.