Every structural member in an airliner carries a set of numbers: the load it is allowed to take, the fatigue life it is certified for, the interval at which somebody must look for a crack in it, and the path the load will travel if it fails anyway. Nothing flies until all four exist. That is what certification means.
Now open the crew room door. Everyone inside holds a licence, a rating, a medical and a check record. Every one of those documents certifies the same thing — that this person is permitted to do the job. Not one of them says what load they carry best, where their margin is thin, or what the operation does when that margin runs out.
This is not because the industry is careless about people. The opposite: read the human-factors literature and you find something close to an engineering marvel. A set of distinct instruments, each specified by a named document, most descended from a specific accident, all aimed at the same target — the fallibility of the person doing the work.
Read them again and you notice what nine of them have in common. Safety management systems, crew resource management, evidence-based training, line operations safety audits, flight data monitoring, fatigue risk management, just culture, maintenance human factors, peer support programmes — every one acts on a person who is already employed, already licensed, already rostered. They train that person, observe them, protect them, and limit their duty hours. Not one of them has anything to say about which person is put in the seat.
That is not an oversight, and it is not negligence. It is a reasonable historical outcome, and there are good arguments for it, which this article will make. But it is a shape worth naming, because it defines exactly where a behavioural benchmark can be useful and — more importantly — where it cannot.
What follows is an account of building one: a reference of 305 civil aviation roles with behavioural ranges across eight parameters and five organisational life-cycle stages, published as a free research preview. The method, the reasoning, and a candid list of what is wrong with it.
1What the reports actually say
The number everyone quotes — that human error causes 70 to 80 per cent of aviation accidents — is real, but it has a specific and rarely stated provenance. It comes from Shappell and Wiegmann's 2000 FAA technical report on HFACS, whose actual wording is more careful than the version in circulation: "estimates in the literature indicate that between 70 and 80 percent of aviation accidents can be attributed, at least in part, to human error." The citation it gives is to the authors' own 1996 study of US naval aviation mishaps between 1977 and 1992.
Three things get lost in the retelling. It is an estimate drawn from the literature, not a finding. It says "at least in part," which is a much weaker claim than causation. And its cited evidence base is naval aviation data from 1977 to 1992. We have not found the figure restated in the current ICAO, IATA or EASA safety publications we checked, and the honest picture underneath it is more interesting than the slogan.
IATA's 2025 Annual Safety Report classifies contributing factors using the Threat and Error Management framework: its analysts "analyzed the underlying contributors to the accidents using the Threat and Error Management (TEM) framework." The categories mix crew competencies, environmental threats and organisational conditions, and an accident can carry several — so they do not sum to 100 per cent and cannot be added into a "human factors percentage." The full published set:
Read it without cherry-picking. The top of the list is mixed: two crew competencies at 29 per cent, weather at 29 per cent, and abnormal runway contact — an outcome rather than a cause — tying crew response at 25 per cent. What is fair to say is that crew-behavioural categories occupy four of the top six positions and cluster between 24 and 29 per cent, while the largest organisational category — the failure of the meta-control designed to catch everything else — sits at 16 per cent. ICAO's own 2026 Safety Report publishes no human-factors percentage at all, saying only that "human activity remains at the core of aviation safety and performance." Boeing's statistical summary classifies by CICTT occurrence category and, as far as we can see, publishes no human-factors attribution at all.
The industry stopped counting human error as a percentage some time ago. It started counting competencies instead — and then trained against them.
| Contributing factor cited in 2025 accidents | Share |
|---|---|
| Aeroplane flight path management, manual control | 29% |
| Situation awareness and management of information | 29% |
| Adverse weather | 29% |
| Crew response and situational awareness | 25% |
| Abnormal runway contact | 25% |
| Manual handling and primary flight control errors | 24% |
| Non-compliance with standard operating procedures | 20% |
| Inadequate safety management system | 16% |
| Airport facilities | 16% |
| Bounced landing | 14% |
| Inadequate regulatory oversight | 12% |
| Vertical, lateral or speed deviation | 12% |
One caution on the headline rate. IATA counts 1.32 accidents per million sectors for 2025; ICAO's own Safety Report for the same year gives 2.23 per million departures. Different scope, different denominator, different counting rules. Both are correct and they must never be mixed in one series.
The cases, across the whole system
The flight deck dominates the public memory of aviation accidents, which distorts the picture. The behavioural failures that kill people are distributed across every part of the operation. A representative spread:
Ten cases across five parts of the system: flight deck, maintenance, weight and balance, air traffic control, and operational control under fatigue. Four are flight-deck cases, which is still an over-representation relative to where the industry's human-factor exposure actually sits — a distortion this article inherits from the public record rather than corrects.
In nine of the ten, the pattern is not incompetence. It is people behaving in entirely characteristic ways — deferring to authority, trusting a familiar number, working past a warning, not escalating — inside systems that had not been designed to expect it. The tenth, Germanwings, is a different category altogether and no generalisation about ordinary human performance covers it; it appears here because of what the BEA concluded about the barriers around the person, not about the person.
Tenerife — 27 March 1977 — KLM 4805 / Pan Am 1736
In fog at a diversion airfield, a 747 began its take-off roll while another was still back-taxiing on the same runway.
The finding: the captain "took off without clearance," "did not obey the 'stand by for take-off' from the tower," and gave an emphatic affirmative when his flight engineer questioned whether the other aircraft had cleared. The canonical authority-gradient case: the query was made, the captain overrode it, and nobody escalated.
United 173 — 28 December 1978 — Portland, Oregon
The crew held for about an hour troubleshooting a landing gear indication and ran the tanks dry.
The finding: probable cause was "the failure of the captain to monitor properly the aircraft's fuel state and to properly respond to the low fuel state and the crewmember's advisories regarding fuel state," with contributing cause "the failure of the other two flight crewmembers either to fully comprehend the criticality of the fuel state or to successfully communicate their concern to the captain." The Board recommended indoctrination in "flightdeck resource management" with "assertiveness training for other cockpit crewmembers." United initiated the first comprehensive US CRM programme in 1981 as a direct result. Every CRM course flown today descends from this accident.
British Airways 5390 — 10 June 1990 — over Didcot, Oxfordshire
A windscreen fitted the previous night departed at around 17,300 ft; the captain was blown half out of the aircraft and held by cabin crew while the co-pilot landed.
The finding: "a safety critical task, not identified as a 'Vital Point', was undertaken by one individual" with no independent check until flight. Eighty-four bolts of the wrong diameter were fitted. The engineer had thirty-three years' experience and an exemplary record, was on his first night shift after five weeks' leave, and self-certified his own work. The textbook demonstration that experience and good intentions are not barriers.
Alaska Airlines 261 — 31 January 2000 — off Point Mugu, California
The horizontal stabiliser trim jackscrew acme nut threads failed and the aircraft pitched over.
The finding: thread failure "caused by excessive wear resulting from Alaska Airlines' insufficient lubrication," with contributing cause the airline's extended lubrication and end-play check intervals "and the FAA's approval of that extension." A commercial-pressure decision chain, ratified by the regulator, that removed both the maintenance action and the inspection that would have caught its omission.
Überlingen — 1 July 2002 — southern Germany
A single controller working two positions at night issued a late descent at the moment TCAS commanded a climb. One crew followed ATC, the other followed TCAS. Both descended.
The finding: alongside the immediate causes, the BFU found that management "tolerated for years that during times of low traffic flow at night only one controller worked and the other one retired to rest." A staffing norm, not a decision, was among the causes.
MK Airlines 1602 — 14 October 2004 — Halifax, Nova Scotia
The crew used take-off performance figures generated for the previous, much lighter departure. The aircraft failed to rotate and overran.
The finding: the TSB concluded that the aircraft was heavier than the crew had calculated, that the crew did not recognise that performance was inadequate for the runway available, that reduced thrust further degraded it, and that fatigue and schedule pressure may have contributed. On the tool itself the report is verbatim: the laptop performance software "would automatically overwrite any entry in the planned weight field, without any notification to the user."
Colgan Air 3407 — 12 February 2009 — Clarence Center, New York
Airspeed decayed unnoticed on an icing approach; the stick shaker fired and the captain pulled aft into it.
The finding: probable cause was "the captain's inappropriate response to the activation of the stick shaker." The Board listed four contributing factors: the crew's failure to monitor airspeed against the rising low-speed cue, their failure to adhere to sterile cockpit procedures, the captain's failure to manage the flight effectively, and — the organisational one — "Colgan Air's inadequate procedures for airspeed selection and management during approaches in icing conditions."
The Board also documented the captain's training record: four FAA checkride disapprovals, three of them before Colgan hired him in September 2005, and three unsatisfactory or remedial proficiency events at Colgan itself. On leadership its finding was that the one-day upgrade course "did not contain significant content applicable to developing leadership skills, management oversight, and command authority." Most of the record existed before he was hired, and all of it existed before he was upgraded.
Emirates 407 — 20 March 2009 — Melbourne
A take-off weight of 262.9 t was entered instead of 362.9 t. The aircraft failed to rotate, struck its tail and hit installations beyond the runway before climbing away on TOGA.
The finding: an incorrect entry that "passed through the subsequent checks without detection," and — more broadly — "the apparent inability of flight crew to perform 'reasonableness checks' to determine when parameters were inappropriate." A single typo, defended only by cross-checks that behavioural science predicts will not catch it.
Air Midwest 5481 — 8 January 2003 — Charlotte, North Carolina
The aircraft pitched up uncontrollably on rotation and crashed.
The finding: loss of pitch control "resulted from the incorrect rigging of the elevator control system compounded by the airplane's aft center of gravity, which was substantially aft of the certified aft limit." Contributing: insufficient oversight of the contracted maintenance station, the quality assurance inspector's failure to detect the rigging error, the carrier's weight and balance programme, and "the FAA's average weight assumptions in its weight and balance program guidance at the time of the accident." Two independent human-origin defects — a maintenance shortcut and an obsolete average-passenger-weight assumption — each survivable alone, and not together. The accident drove the FAA's revision of standard average passenger weights.
Germanwings 9525 — 24 March 2015 — French Alps
The co-pilot locked the captain out of the flight deck and flew the aircraft into terrain.
The finding, as the BEA framed it: deliberately systemic rather than individual. The co-pilot continued to exercise his licence because of "the co-pilot's probable fear of losing his right to fly as a professional pilot if he had reported his decrease in medical fitness to an AME"; the absence of insurance covering loss of income from unfitness; and "the lack of clear guidelines in German regulations on when a threat to public safety outweighs the requirements of medical confidentiality." The BEA examined the reinforced door and did not recommend changing the locking system — a point frequently misreported.
2The apparatus, and its shape
What the industry built in response is genuinely impressive, and it is worth laying out in full before criticising it. Ten instruments, each with a mandating document, and — for each — the thing it does not do.
One provision does point upstream: CAT.GEN.MPA.175(b), introduced across Europe after Germanwings, requires that "flight crew has undergone a psychological assessment before commencing line flying." It is a single gate at the point of joining an operator, with no harmonised standard for what the assessment must contain or what result disqualifies. It applies to flight crew only. There is no equivalent anywhere for maintenance engineers, ramp staff, loadmasters or dispatchers. Controllers are the exception, and section 3 is about why.
| Instrument | Mandated by | What it does not cover |
|---|---|---|
| Safety management system | ICAO Annex 19; Doc 9859; ORO.GEN.200; 14 CFR Part 5 | A process standard. Says nothing about who is employed or how they are selected. Manages the risk generated by the people you already have. |
| CRM and threat & error management | ICAO Doc 9683; ORO.FC.115 and AMC1 | A training intervention delivered to whoever is on the roster. Assumes trainability. No mechanism for identifying someone temperamentally resistant to challenge. |
| CBTA and evidence-based training | ICAO Doc 9868 PANS-TRG; Doc 9995; IATA EBT Guide Ed. 2 | A recurrent training and assessment system, not a selection system. Persistently low competency produces remediation, not non-hire. |
| Line operations safety audit | ICAO Doc 9803 (recommended, not a Standard) | De-identified and voluntary by design — which is what makes it work, and what makes it structurally unable to say which individual is a problem. |
| Flight data monitoring / FOQA | ICAO Annex 6 Part I (>27 000 kg); Doc 10000; ORO.AOC.130; CAP 739 | Records the aeroplane, not the person. Captures the outcome of a decision, never the reasoning, workload or interpersonal dynamic behind it. Retrospective by definition. |
| Fatigue risk management | ICAO Annex 6 SARPs; Doc 9966; ORO.FTL; 14 CFR Part 117 | Manages exposure — rosters, rest opportunity, circadian placement. Gives you a compliant roster, not a rested person. |
| Just culture and confidential reporting | ICAO Annex 19 App. 3; Reg (EU) 376/2014 | A protection regime for information, not a competence regime for people. Deliberately makes it harder to act on an individual using safety data. |
| Maintenance human factors | ICAO Doc 9824; Part-145.A.30(e) and .A.65(b); Part-66 Module 9 | The Dirty Dozen is a taxonomy, not a control. Nothing in it speaks to selecting engineers for conscientiousness or resistance to production pressure. |
| Peer support and aeromedical mental health | Reg (EU) 2018/1042 — CAT.GEN.MPA.215; ICAO Doc 8984 | Access to help, contingent on the person choosing to use it. Not continuous surveillance — deliberately, because that would destroy the disclosure culture it depends on. |
| Pilot aptitude testing | IATA Pilot Aptitude Testing Ed. 3, 2019 — guidance material | Voluntary. No ICAO Standard, no EASA rule and no FAA regulation requires it. Contains no AI, game-based or adaptive-testing content. |
3The exception that proves it can be done
There is one safety-critical aviation role where an aptitude screen is required at the door — and where, as a result, somebody has at least published numbers on either side of a score threshold.
The FAA requires the Air Traffic Skills Assessment of Track 1 controller applicants, those without prior experience: a battery measuring, in the GAO's description, "cognitive abilities and personal abilities, including mathematical ability, decision-making, spatial comprehension, memory, planning." In its December 2025 report on the controller workforce, the GAO published the outcome by score band: applicants scoring above 85 became certified controllers or were still in training in 2022 at a rate of 29 per cent, against 9 per cent for those scoring 80 to 84.9. In the same report, roughly 43 per cent of Track 1 applicants who started at the FAA Academy were no longer controllers or in training as of 2024.
This is the fairest way to state the case. Not that the industry ignores selection, but that it has built an extraordinarily sophisticated post-hire safety apparatus on top of a pre-hire filter that remains voluntary, unstandardised and largely unvalidated against safety outcomes.
Where a screen is mandated, there is at least an argument to have about the numbers. Where none is mandated, there are no numbers to argue about.
4The load path
There is a way of thinking about reliability that aviation applies to everything it flies, and to almost nothing that walks onto the ramp.
An airframe is not designed by asking whether a member will hold. It is designed by asking what happens when it does not. Damage tolerance under CS-25.571 requires the structure to carry load with a member cracked or failed, and requires an inspection programme that will find the damage before residual strength runs out. Where a member has no alternative load path, it is named as such and given its own regime.
The same reasoning runs through the rest of the machine and the organisation around it. Reliability programmes track removals against alert levels. Systems are dual or triple. Vital Points in maintenance get an independent check for one reason: a single signature is a single load path, which is exactly the finding the AAIB made about the windscreen on BA 5390.
Then it stops. For the person there is no rating, no margin, no second path calculated in advance. What exists instead is a set of barriers — training, monitoring, rostering limits, reporting — every one of them the same size for everybody, because without a way to see the distribution there is no defensible basis for making them any other size.
A structure is designed on the assumption that one member will fail. The crew, the shift and the watch are designed on the assumption that nobody in particular will.
Where the metaphor earns its keep
The vocabulary already exists and is already understood in this industry: load, margin, redundancy, single point of failure, alternative load path. A role benchmark is the nearest thing to a load rating for the role — what the seat asks of whoever sits in it. Set a profile against it and the questions that fall out are the familiar engineering ones. Where is the margin thin? Which demands sit on one person with nothing behind them? If this quality is the one that gives, what carries the load instead?
That is precisely what the tool's output is written to produce. A shortfall of twenty to forty percentile points does not return a verdict; it returns a list of alternative load paths — a tool that carries part of the work, a colleague strong in that area, a control at the step where the quality matters, a shift of the dependent work towards a better-suited profile. That is damage-tolerant design, applied to a rota instead of a wing box.
So the honest framing: a table of ranges can be a useful input into thinking about where human load concentrates and where there is no second path behind it. It is not a reliability calculation, it produces no MTBF for a human being, and no one should present it as though it does.
5Building the reference
With that framing, the object we set out to build is narrow and specific: not a selection gate, but a shared vocabulary for what a given aviation role actually demands of a person — and a way of reading a behavioural profile against it.
The role list
The 305 roles are not a generic job catalogue filtered for aviation words. Each entry is a post, licence, rating or nominated position that exists in a published instrument, and each carries the reference that creates it. Four kinds of source were used.
The result spans 28 functional domains. Flight deck and cabin are a minority of it. Maintenance and continuing airworthiness, air traffic management and CNS engineering, aerodrome operations and rescue and fire fighting, ground handling, cargo and dangerous goods, security, the regulator itself, and the airline commercial and corporate stack all carry more roles between them than the front of the aircraft does.
- ICAO Annexes and Docs. Annex 1 for every licence and rating from PPL through the Aircraft Maintenance Licence, the Air Traffic Controller licence and the Flight Operations Officer licence; Annexes 3, 6, 11, 13, 14, 15, 17 and 19 for operational, ATS, investigation, aerodrome, information, security and safety management posts; and the Doc series — 9859, 9868, 8973, 9734, 10070, 9756, 9966, 9824, 10106, 10000 — for the competency content.
- IATA manuals and programmes. IGOM, AHM, ISAGO, the IOSA Standards Manual, the Dangerous Goods Regulations, the Cargo Agent's Handbook, WASG, SSIM and the CBTA/ITQI framework.
- EASA and FAA implementing rules where they give a post a legal name that ICAO leaves to the State: Part-ORO nominated persons, Part-FCL instructor and examiner certificates, Part-66 licence categories, Part-145, Part-CAMO and Part-21 post-holders, Regulation (EU) 2015/340 for ATCO ratings and Regulation (EU) 139/2014 for aerodromes; 14 CFR Parts 119, 121 and 65 for the US equivalents.
- Standing industry standards for the few technical roles no authority names directly — EN 4179 and NAS 410 for the three NDT levels, for example.
The eight parameters
The same eight NeuroFrame parameters, scored as percentiles against a reference base of 14 850 professionals: three cognitive — Mental Efficiency, Learning Agility, Progress Monitoring — and five personality — Openness to New, Risk Appetite, Result Focus, Agreeableness, Conscientiousness. Nothing aviation-specific was added. The aviation content lives in the ranges, not in the instrument.
How a range is produced
Each role's benchmark starts from the structure of what the work demands, expressed as a blend of two behavioural archetypes weighted for that domain — real-time safety, technical craft, engineering analysis, regulatory assurance, investigative, command leadership, operations control, commercial analytics, commercial relationship, instruction, customer frontline, security, planning and logistics, or medical and human. A seniority adjustment follows, then a life-cycle modifier for the stage of the organisation. Three parameters per role are identified as core — the ones that most distinguish this role from the average of all others — and carry a tighter band, ±7 percentile points against ±11 for the rest.
Why the life-cycle stages stop at Prime
The tool offers five stages — Prime, Stability, Aristocracy, Early Bureaucracy, Bureaucracy — and omits the four earlier ones. Aviation organisations are almost never young in the relevant sense. Airlines, air navigation service providers, airports, maintenance organisations and regulators operate under certificates, approvals and continuous oversight from their first day of operation. Infancy and Go-Go describe a company that can still improvise; a certificated operator cannot.
The direction of travel down the remaining curve is consistent: appetite for risk and openness to the new fall away, monitoring and conscientiousness rise, and the pull towards results weakens as internal process absorbs more of the day. The same job title asks for a measurably different person in a Prime airline and in a bureaucratic one.
Reading a difference
A band on its own invites the wrong behaviour — treating it as a pass mark. So the tool grades each difference on four levels, and the grading is deliberately asymmetric, because a surplus and a shortfall are not the same kind of problem.
| Level | When | What it means |
|---|---|---|
| OK | Inside the band, at any value | No signal, no note. The band is what the role asks for; meeting it is the point. |
| NOTE | Above the band up to 90 · below it by less than 20 points | Minimal attention. A small shortfall is usually made up through persistence and conscientiousness; a surplus is a matter of internal balance, not of capping the person. |
| WATCH | Above the band and above 90 · below it by 20–40 points | Structural support: tools, peer help, an added control, or shifting the part of the work that most depends on the quality. Above 90 the concern inverts — overuse, and the tension of a strength that is not being spent. |
| ACT | Below the band by more than 40 points | Systemic attention to the person, named measures on a review cadence, and — if the parameter is core to the role — delegating, distributing or redesigning the part that depends on it. |
Risk appetite is scored on its own scale
Too little of it rarely breaks a role, so a shortfall stays a note as long as it is at or above 30. Below the band and under 30, decisions under genuine uncertainty become hard and the person needs help making them. Above the band at 80 to 90 the issue is not that the person is wrong — people who take risk are doing what they are built to do — but that the balancing has to be supplied by the system rather than by their self-restraint: a second signature on irreversible calls, a written stop rule, a peer whose job is to disagree. Above 90 that structure stops being advisable and becomes necessary.
One caution about evidence
Where a large gap sits on a parameter core to the role, a strong track record is not proof that the gap has been closed. The person may have been carrying it on willpower, and that does not hold indefinitely. This is the single most common way a reassuring conversation goes wrong.
6What is wrong with it
A methodology section that lists only strengths is marketing. Here is the honest inventory.
- The ranges are model output, not measurements. Nobody has measured 305 aviation roles across five life-cycle stages. The bands are generated from the structure of each role's published competency demands and a life-cycle modifier. They have not been validated against job performance in any airline, airport, ANSP or maintenance organisation.
- No criterion validity is claimed, because none has been established. There is no published correlation between these bands and training completion, check-ride outcomes, occurrence rates or any other operational criterion. Establishing that would require an operator's own data and a study design — which is the work we would like to do next, and which is the only thing that would make the tool more than a conversation structure.
- Life-cycle stage is a judgement, not a measurement. The user chooses it. The Adizes model is one of several competing frameworks, none with consensus in organisational research, and large organisations are frequently in different stages in different functions at once.
- Role classification is imperfect. Where two instruments name the same post differently the entries were merged; where a role spans domains it was assigned to one. A reader who works in a role will sometimes disagree with where it sits, and they will usually be right.
- Aviation instruments change. The citations were accurate when compiled and some will be superseded. The tool cites them to identify roles, not as authoritative text.
- It says nothing about eligibility. A candidate can sit perfectly inside every band and still be legally unable to hold the post; the reverse is equally true. Licences, ratings, medical certificates, language proficiency and security clearance are determined by the competent authority on evidence that has nothing to do with percentiles.
7Where this fits
The interesting question is not whether behavioural data should replace any part of the existing apparatus. It should not, and could not. The question is where it attaches to what already exists.
Into competency-based training, as a targeting input
EBT already assesses the nine ICAO competencies and directs recurrent training at the ones a fleet's data shows to be weak. It does that reactively, from observed performance. A behavioural profile is available before the first line sector — which means the same remediation can be planned rather than triggered. A first officer whose profile sits well below the band on Progress Monitoring is not a problem to be screened out; they are a person whose monitoring training should be front-loaded and whose first six months should include a specific check.
Into crew pairing and rostering, as a composition input
Authority gradient is the single most persistent theme across the accident record — Tenerife, Birgenair, Air India Express 812. Rostering systems today pair on legality, currency and cost. Nothing in them knows that a very high Result Focus captain has been paired with a very low Risk Appetite first officer on a demanding night sector, which is a specific and foreseeable combination.
Into the SMS hazard register, as a population-level input
A safety manager can currently describe the hazards of their operation but not the behavioural distribution of the people running it. Knowing that a station's ramp population sits systematically low on Conscientiousness relative to the role band is a hazard with a named control — not a reason to move anyone, but a reason to add the control before the ground damage happens.
Into maintenance human factors, as a specificity upgrade
The Dirty Dozen is a taxonomy of error precursors delivered identically to everyone. It has been delivered identically to everyone for thirty years. Knowing which precursors a specific team is actually exposed to — by profile — turns a generic awareness module into a targeted one.
Into the post-hire support layer, as a design input
This is the closest fit to what the tool already produces. Every WATCH and ACT output is written as a support prescription, not a verdict: what to give the person, what control to add, what part of the work to shift. That is the same vocabulary the SMS already speaks.
The gap is not that aviation manages the human element badly. It is that it manages the person in the seat superbly and the choice of person barely at all — and the second is upstream of the first.
What we are actually claiming
It is worth being exact about this, because the temptation in this field is to overreach and the cost of overreaching in aviation is credibility that does not come back.
We are not claiming to solve aviation's human-factors problem. The apparatus described in section 2 is the product of fifty years of investigation and a great many deaths, and it works: the accident rate is 1.32 per million sectors, which is to say that a person can fly for a lifetime and never see one. Nothing in a behavioural benchmark improves on that, and any vendor who tells an airline otherwise should be shown the door.
The claim is narrower and, we think, more useful. Behavioural data does two things that the existing instruments structurally cannot.
And this matters more now than it did. The requirements themselves are moving. Recurrent training has shifted from set-piece manoeuvres to competency assessment. Automation policy is being rewritten in the direction of manual handling after two decades of the opposite. The syllabus items EASA added post-AF447 and post-Asiana — monitoring and intervention, resilience, surprise and startle — are behavioural, not technical. An extraordinary generational turnover is running through the pilot and engineer populations at the same time as new entrants with different training pathways arrive. The behaviour a role asks for in 2026 is measurably not the behaviour it asked for in 2006, and it is changing faster than the instruments that measure it.
When the target is moving, knowing the shape of your population is worth more than it was when the target stood still. That is the whole claim. It is not a safety guarantee, and it is not a hiring gate.
- It makes the distribution visible. FDM shows what happened on the line. LOSA shows what normal operations look like, de-identified. Neither can tell a head of training what the behavioural shape of their own population is — where it clusters, where it is thin, whether the intake cohort differs from the people who have been on the fleet for fifteen years. That is a trend question, and trends are what let you act before an event rather than after one.
- It gives the link between personality and behaviour a vocabulary someone can act on. Every accident report in section 1 describes behaviour: the captain who did not stop, the engineer who signed his own work, the flight engineer who asked the question and was overruled. Reports describe it after the fact, in prose, one case at a time. A parameter model describes the same thing beforehand, in a form that can be aggregated, compared across a population, and connected to a specific control.
And nobody has spare money
This is the part that usually goes unsaid, so let us say it. Every intervention in section 2 is delivered to everybody, because without a way to differentiate there is no defensible basis for doing anything else. The Dirty Dozen module is the same module for the engineer who has never missed a torque value and the one who is quietly drowning. Recurrent CRM is the same day for the captain who invites challenge and the one who suppresses it. Uniform delivery is not a training philosophy; it is what you do when you cannot see the distribution.
That is the expensive option, and it is expensive in two directions at once — money spent on people who did not need it, and money not spent on the people who did. The counterfactual costs are documented and unpleasant. Direct aircraft operating cost runs at $98.41 per block minute for US passenger carriers in 2025 (Airlines for America) — crew $37.01, fuel $29.34, maintenance $18.35, ownership $9.76, other $3.95. Four hours of block time on that basis is about $23,600. That is an accounting ceiling, not the marginal cost of a diversion: ownership and much of the crew cost is incurred whether the aircraft diverts or not, and a diversion adds the delta in block time rather than four fresh hours. The real number sits below it, and above it once passenger reaccommodation, a breached duty limit and a downline cancellation are counted. IATA projects the annual cost of ground damage could reach $10 billion by 2035. In the one role with a mandatory aptitude screen, the FAA still loses roughly 43 per cent of Track 1 Academy entrants — and can at least see which score band they came from.
Two honest caveats on the money. The often-quoted maintenance-error cost figures — engine shutdowns at $500,000, cancellations at $66,000 — trace back to 1990s Boeing MEDA material that we could not verify in the primary source, so we do not use them. And the best peer-reviewed pilot turnover figure we found, $17,405 per pilot, comes from a 2018 study of a Part 135 cargo operator and is not transferable to a mainline carrier with type-rating bonds and a larger training footprint. It is quoted here as what it is.
The argument is not that a behavioural benchmark saves money. It is that when there is no spare money, spending the same amount on everyone is the least efficient thing an operator can do — and seeing the distribution is the cheapest step towards spending it where it changes something.
8What would make this real
A benchmark nobody has tested is a hypothesis with a user interface. Three things would change that, in order of value.
Until then, what exists is a free research preview with 305 roles in it, published so that people who do this work can tell us where it is wrong. That is the honest description, and it is also the point: the first disagreement from a head of training who has actually run a fleet is worth more than another month of modelling.
- A criterion study with one operator. Profiles for an existing population, matched against training completion, check outcomes, FDM event rates at crew level where the data protection regime permits, and supervisor ratings. Even a single fleet would tell us which of the 305 role profiles are wrong, and by how much.
- Independent psychometric review of the underlying instrument. Not of this calculator — of the assessment behind it. The routes that exist are national: BPS registration in the UK against the EFPA review model, COTAN in the Netherlands, DIN 33430 certification in Germany. None of them is an aviation credential, and that is fine; the aviation credibility comes from the criterion study, not from a mark.
- Contribution to the standards conversation. At the 42nd ICAO Assembly in 2025, Kazakhstan tabled a working paper on developing standardised psychometric assessment guidelines in pilot licensing, and Argentina one on accounting for personality traits in controller selection. The gap this article describes is recognised at Assembly level. If a Standard is written, it should be written with data in front of it.
9Questions
Colophon. Role definitions and competency demands compiled from published ICAO Annexes and Docs, IATA manuals and programmes, and EASA and FAA implementing rules, August 2026. Accident findings quoted from official investigation reports as cited. Behavioural ranges are NeuroFrame's own model. Life-cycle stages are adapted from the corporate life-cycle model of Ichak Adizes — one of several competing frameworks, referenced for identification only; NeuroFrame is not affiliated with or endorsed by the Adizes Institute.
Sources.
Questions
Is this an assessment?
No. Entering values into the comparison panel does not assess, score, screen or rank anybody — it plots the numbers you typed against a modelled range. The assessment, if there is one, happened elsewhere.
Where did the behavioural ranges come from, exactly?
From the structure of each role's published competency demands, blended into two weighted behavioural archetypes for that domain, adjusted for seniority tier, modified by life-cycle stage, and then — for safety-critical domains — clamped by the safety floor. They are our own model output. The ICAO and IATA documents define the roles and their competency demands; they do not publish percentile ranges, and nothing in them endorses ours.
Why is risk appetite treated differently from the other seven parameters?
Because the asymmetry is real. A shortfall of risk appetite rarely breaks a role — it makes someone slower to commit under uncertainty, which is usually recoverable. A surplus in a safety-critical role is a different class of problem, and one that self-restraint is a poor control for. So the tool caps risk appetite for safety-critical roles at every life-cycle stage and escalates faster above the band than below it.
Does a profile outside the band mean someone should not do the job?
No, and the tool is written to resist that reading. Every flagged difference produces a support prescription — what to give the person, what control to add, what to shift — rather than a verdict. The only place the output mentions changing a role is where a gap of more than 40 percentile points sits on a parameter core to it, and even there the recommendation is to examine redesign, not to move anyone.
Has any aviation authority approved this?
No, and none has been asked. ICAO states in its own personnel licensing guidance that it does not endorse, recognise or approve training organisations or training programmes — the narrow exception being its own AVSEC and Global Aviation Training regional centres — and neither ICAO nor IATA operates a certification scheme for assessment instruments of this kind. The only ICAO scheme that evaluates a commercial test at all is AELTS, which covers aviation English language testing exclusively.
What happens to the numbers I type in?
Nothing leaves your browser. The comparison panel holds them in page memory for the duration of your visit; reloading discards them. If you enter data about another person, you are the controller of that data, not us.
Try it, and tell us where it breaks
The Aviation Role Fit Calculator is free, requires no account, and processes everything in your browser. If a band looks wrong for a role you know, we want to hear about it.
Sources
- Shappell, S. A. and Wiegmann, D. A., The Human Factors Analysis and Classification System — HFACS, DOT/FAA/AM-00/7, February 2000.
- IATA, Annual Safety Report 2025, published March 2026.
- ICAO, Safety Report, 2026 Edition (2025 data).
- EASA, Annual Safety Review 2025 (2024 data).
- Subsecretaría de Aviación Civil (Spain), report on the collision at Los Rodeos, 27 March 1977.
- NTSB, AAR-79-07, United Airlines Flight 173, Portland, Oregon, 28 December 1978.
- AAIB, AAR 1/1992, BAC One-Eleven G-BJRT, 10 June 1990.
- NTSB, AAR-02/01, Alaska Airlines Flight 261, 31 January 2000.
- BFU, AX001-1-2/02, Überlingen mid-air collision, 1 July 2002.
- Transportation Safety Board of Canada, A04H0004, MK Airlines Flight 1602, 14 October 2004.
- NTSB, AAR-04/01, Air Midwest Flight 5481, Charlotte, 8 January 2003.
- NTSB, AAR-10/01, Colgan Air Flight 3407, 12 February 2009.
- ATSB, AO-2009-012, Emirates Flight 407, Melbourne, 20 March 2009.
- BEA, final report on the accident to Airbus A320-211 D-AIPX, 24 March 2015, published March 2016.
- ICAO Annex 19 Safety Management; Doc 9859 Safety Management Manual, 4th ed.
- ICAO Doc 9683 Human Factors Training Manual; EASA AMC1 ORO.FC.115.
- ICAO Doc 9868 PANS-TRG; Doc 9995 Evidence-Based Training; IATA EBT Implementation Guide, Ed. 2, effective January 2024.
- ICAO Doc 9803 Line Operations Safety Audit (LOSA), 1st ed., 2002.
- ICAO Doc 10000 Manual on Flight Data Analysis Programmes; EASA ORO.AOC.130; UK CAA CAP 739.
- ICAO Doc 9966 Manual for the Oversight of Fatigue Management Approaches, 2nd ed., rev. 2020.
- Regulation (EU) No 376/2014 on occurrence reporting.
- ICAO Doc 9824 Human Factors Guidelines for Aircraft Maintenance, 1st ed., 2003; Dupont, G., the Dirty Dozen, Transport Canada, 1993.
- Commission Regulation (EU) 2018/1042 — CAT.GEN.MPA.175(b), CAT.GEN.MPA.215, ARO.RAMP.106.
- IATA, Pilot Aptitude Testing — Guidance Material and Best Practices, Ed. 3, April 2019.
- US Government Accountability Office, GAO-26-107320, Air Traffic Control Workforce, 17 December 2025.
- Airlines for America, U.S. Passenger Carrier Delay Costs, 2025 data.
- IATA press release, 7 May 2024, ground handling and ground damage cost projection.
- Kiernan, K. M., Calculating the Cost of Pilot Turnover, Journal of Aviation/Aerospace Education & Research, vol. 27, no. 1, 2018.
- Helmreich, R. L., Merritt, A. C. and Wilhelm, J. A., The Evolution of Crew Resource Management Training in Commercial Aviation, FAA.
- ICAO A42-WP/318 (Kazakhstan) and A42-WP/417 (Argentina), 42nd Assembly, 2025.