
EBT practice guide
Instructor concordance: measuring and closing the grading gap
Concordance is the degree to which two instructors observing the same crew reach the same assessment. Without measurement it decays quietly, and the competency data your programme depends on becomes partly noise. This guide covers how to measure it, what the numbers mean, and how to run a calibration cycle that holds.
Contextual Composition Protocol
Concordance built in before the session, not only corrected after it
Most concordance work happens after the fact: measure the drift, then calibrate. The Contextual Composition Protocol, the SkyDynamics technology behind every AeroEBT application, also works upstream. It holds all the data relevant to the task in memory, so when an instructor records and assesses in the ORCA cycle, the instructor app already knows the scenario phase, the event or failure in play, the observable behaviours it is designed to elicit, and your competency framework and grading scale. What the instructor records at that moment is correlated with that context, with high specificity. Every instructor observes and assesses against the same correlated reference, so there is less drift to correct, and measurement and calibration work on a narrower gap. That is how a training department moves towards world-class concordance and instructor standardisation.
- The scenario, the event or failure in play and its expected behaviours held behind every observation
- Each record correlated with the exact moment of the session it belongs to
- The same context for every instructor at the same point of the session
- Every grade carries the fullest possible justification: the moment, the event and the expected behaviours it was assessed against
- Instructor standardisation supported where required, behind the scenes: the protocol never intervenes, and no instructor has to learn how to use it
- Drift still measured against the fleet baseline, so the effect is evidenced rather than assumed
Definition
Three words that get used interchangeably, and should not be
Concordance
Do two instructors observing the same performance agree on what they saw and how they assessed it?
Inter-rater reliability
The statistical expression of that agreement across a group of instructors and a body of sessions.
Referent rater
A reference assessment of a recorded session that instructors are compared against during calibration.
Causes
Why grading drifts apart
None of these are instructor failings. They are programme design problems that show up in the data.
The scale means different things to different people
Without worked examples, "average" is an instructor's personal midpoint rather than a shared standard.
Observation is harder than grading
Instructors consistently rank competency observation and facilitation as the skills they most want to strengthen. The grade is the easy part.
Severity and leniency are stable traits
Individual instructors tend to be consistently harsher or softer than the group. That is measurable and correctable.
Procedural competencies absorb everything else
A low grade for procedures often really reflects workload management or decision-making. The number hides the cause.
Calibration is annual, drift is continuous
A yearly standardisation day cannot hold a standard that moves every month.
Nobody sees the aggregate
Instructors see their own sessions. Without a fleet view, no one can see that two bases have separated.
Method
How to measure concordance honestly
1. Normalise the data
Compare like with like: same competency, comparable scenario phase, comparable crew experience.
2. Establish the baseline
The fleet distribution per competency is the reference, not an abstract ideal.
3. Position each instructor
Show each instructor's distribution against that baseline, including how often they use the extremes of the scale.
4. Add a referent check
A recorded session graded by everyone gives a direct comparison that live sessions cannot.
5. Read the narrative
Where grades agree but notes diverge, the observation is drifting even though the numbers look healthy.
6. Re-measure after calibration
A calibration session that does not move the distribution did not work. Measure it.
Operating rhythm
A calibration cycle that holds
In AeroEBT
Concordance as part of the workflow, not a project
AeroEBT measures inter-rater reliability from the grades your instructors already enter, shows each instructor against the fleet baseline, and keeps calibration sessions and their outcomes in the same record as the training itself.
- Distribution per instructor, per competency, per base
- Calibration sessions planned, recorded and measured for effect
- Observations extracted from instructor notes so narrative drift is visible too
- Evidence for your authority produced from the same data
Questions training standards managers ask
Is an instructor who grades lower than the fleet wrong?
Not necessarily. Severity is a trait, not an error, and a strict instructor may be observing more carefully. What matters is that the difference is known, discussed and stable rather than invisible.
How many sessions do we need before the numbers mean anything?
Enough that an instructor's distribution is not dominated by a handful of crews. In practice that means reading trends over a training cycle rather than reacting to single sessions.
Does measuring concordance turn into instructor scoring?
It should not. The purpose is a shared standard. Programmes that publish league tables get defensive grading, which is worse than drift.
How does the Contextual Composition Protocol improve concordance?
Much grading variation starts before the grade: instructors look for different things because they carry different context in their heads. The Contextual Composition Protocol, SkyDynamics technology running behind every AeroEBT application, holds that context in memory instead. When an instructor records and assesses in the ORCA cycle, the instructor app already knows the scenario phase, the event or failure in play, the observable behaviours it is designed to elicit and the competency framework, and it correlates what the instructor records with that moment. Observations start from a shared reference, and measurement and calibration then close the gap that remains.
Can AI grade instead of the instructor?
No. Assessment of human performance stays with the instructor. AI is useful for extracting observations from notes and showing patterns across a population that no individual can see.
What does an authority expect to see?
Evidence that standardisation is managed: how it is measured, what was found, what was done about it and whether it worked.
See where your instructors actually stand.
In a 30-minute demo we look at how your grading is distributed today and what a calibration cycle would change.
We respond within one business day.
