Skip to main content
SkyDynamics - home
Instructor grading comparison across a fleet in AeroEBT

EBT practice guide

Instructor concordance: measuring and closing the grading gap

Concordance is the degree to which two instructors observing the same crew reach the same assessment. Without measurement it decays quietly, and the competency data your programme depends on becomes partly noise. This guide covers how to measure it, what the numbers mean, and how to run a calibration cycle that holds.

Contextual Composition Protocol

Concordance built in before the session, not only corrected after it

Most concordance work happens after the fact: measure the drift, then calibrate. The Contextual Composition Protocol, the SkyDynamics technology behind every AeroEBT application, also works upstream. It holds all the data relevant to the task in memory, so when an instructor records and assesses in the ORCA cycle, the instructor app already knows the scenario phase, the event or failure in play, the observable behaviours it is designed to elicit, and your competency framework and grading scale. What the instructor records at that moment is correlated with that context, with high specificity. Every instructor observes and assesses against the same correlated reference, so there is less drift to correct, and measurement and calibration work on a narrower gap. That is how a training department moves towards world-class concordance and instructor standardisation.

  • The scenario, the event or failure in play and its expected behaviours held behind every observation
  • Each record correlated with the exact moment of the session it belongs to
  • The same context for every instructor at the same point of the session
  • Every grade carries the fullest possible justification: the moment, the event and the expected behaviours it was assessed against
  • Instructor standardisation supported where required, behind the scenes: the protocol never intervenes, and no instructor has to learn how to use it
  • Drift still measured against the fleet baseline, so the effect is evidenced rather than assumed
ScenarioSETINSERTOBSERVEINSERTPublishPhase · ClimbSAWWLMOBSERVESAW12345OfflineSame scenarioAI-assisted · instructor in control
Every observation correlated with the moment it belongs to.Illustrative

Definition

Three words that get used interchangeably, and should not be

  • Concordance

    Do two instructors observing the same performance agree on what they saw and how they assessed it?

  • Inter-rater reliability

    The statistical expression of that agreement across a group of instructors and a body of sessions.

  • Referent rater

    A reference assessment of a recorded session that instructors are compared against during calibration.

Causes

Why grading drifts apart

None of these are instructor failings. They are programme design problems that show up in the data.

  • The scale means different things to different people

    Without worked examples, "average" is an instructor's personal midpoint rather than a shared standard.

  • Observation is harder than grading

    Instructors consistently rank competency observation and facilitation as the skills they most want to strengthen. The grade is the easy part.

  • Severity and leniency are stable traits

    Individual instructors tend to be consistently harsher or softer than the group. That is measurable and correctable.

  • Procedural competencies absorb everything else

    A low grade for procedures often really reflects workload management or decision-making. The number hides the cause.

  • Calibration is annual, drift is continuous

    A yearly standardisation day cannot hold a standard that moves every month.

  • Nobody sees the aggregate

    Instructors see their own sessions. Without a fleet view, no one can see that two bases have separated.

Method

How to measure concordance honestly

  1. 1. Normalise the data

    Compare like with like: same competency, comparable scenario phase, comparable crew experience.

  2. 2. Establish the baseline

    The fleet distribution per competency is the reference, not an abstract ideal.

  3. 3. Position each instructor

    Show each instructor's distribution against that baseline, including how often they use the extremes of the scale.

  4. 4. Add a referent check

    A recorded session graded by everyone gives a direct comparison that live sessions cannot.

  5. 5. Read the narrative

    Where grades agree but notes diverge, the observation is drifting even though the numbers look healthy.

  6. 6. Re-measure after calibration

    A calibration session that does not move the distribution did not work. Measure it.

Operating rhythm

A calibration cycle that holds

In AeroEBT

Concordance as part of the workflow, not a project

AeroEBT measures inter-rater reliability from the grades your instructors already enter, shows each instructor against the fleet baseline, and keeps calibration sessions and their outcomes in the same record as the training itself.

  • Distribution per instructor, per competency, per base
  • Calibration sessions planned, recorded and measured for effect
  • Observations extracted from instructor notes so narrative drift is visible too
  • Evidence for your authority produced from the same data
Before12345CalibrationAfter · re-measured12345Fleet baselineInstructorMeasured, not steered
A calibration session, measured for effect.Illustrative

Questions training standards managers ask

Not necessarily. Severity is a trait, not an error, and a strict instructor may be observing more carefully. What matters is that the difference is known, discussed and stable rather than invisible.

Enough that an instructor's distribution is not dominated by a handful of crews. In practice that means reading trends over a training cycle rather than reacting to single sessions.

It should not. The purpose is a shared standard. Programmes that publish league tables get defensive grading, which is worse than drift.

Much grading variation starts before the grade: instructors look for different things because they carry different context in their heads. The Contextual Composition Protocol, SkyDynamics technology running behind every AeroEBT application, holds that context in memory instead. When an instructor records and assesses in the ORCA cycle, the instructor app already knows the scenario phase, the event or failure in play, the observable behaviours it is designed to elicit and the competency framework, and it correlates what the instructor records with that moment. Observations start from a shared reference, and measurement and calibration then close the gap that remains.

No. Assessment of human performance stays with the instructor. AI is useful for extracting observations from notes and showing patterns across a population that no individual can see.

Evidence that standardisation is managed: how it is measured, what was found, what was done about it and whether it worked.

See where your instructors actually stand.

In a 30-minute demo we look at how your grading is distributed today and what a calibration cycle would change.

We respond within one business day.