Publications
Explore peer-reviewed publications spanning air traffic management, AI agents, digital twins, and safety validation.
Showing 15 of 15 publications
Graph-based Complexity Forecasts in UK En Route Airspace Using Relevant Aircraft Interactions
Edward Henderson, George De Ath, Nick Pepper
Effectively managing Air Traffic Control Officer (ATCO) workload is crucial in maintaining operational safety. Group supervisors use tools that estimate upcoming traffic load to aid decision-making. However, industry-standard models can fail to capture the nuances of upcoming air traffic complexity. This study presents a probabilistic approach to forecast the complexity of an airspace sector using the number of relevant aircraft pairs, i.e., those that require monitoring or deconfliction by a controller, as a proxy measure for ATCO workload. We adapted an existing filter algorithm to make it suitable for use in London Middle Sector (LMS), a complex airspace sector with multiple flows of traffic above some of the busiest airports in Europe. Through iterative feedback with ATCOs, the algorithm was refined and extended to handle specific geometric and operational considerations. The updated algorithm outperformed the original, with an F1-score of 0.84 compared to 0.69 on a labelled set of 50 traffic scenarios. To produce forecasts of future numbers of relevant aircraft pairs in the sector, a graph representation of the LMS route network was constructed, standardising the spatial fidelity of route legs. The forecasting method accounts for uncertainty in aircraft arrival times by modelling the probability of each aircraft occupying route segments at future query times. When combined with historic distributions of relevant interactions and a live operational data stream, predictions of upcoming ATCO workload could be made up to 45 minutes in advance. The proposed method to forecast upcoming workload showed a significantly stronger correlation with actual relevant interactions than a standard traffic volume prediction. The resulting data-driven tool shows promise for use by group supervisors to inform sector configuration and ATCO rostering decisions.
Conditioning Aircraft Trajectory Prediction on Meteorological Data with a Physics-Informed Machine Learning Approach
Amy Hodgkin, Nick Pepper, Marc Thomas
Accurate aircraft trajectory prediction (TP) in air traffic management systems is confounded by a number of epistemic uncertainties, dominated by uncertain meteorological conditions and operator specific procedures. Handling this uncertainty necessitates the use of probabilistic, machine learned models for generating trajectories. However, the trustworthiness of such models is limited if generated trajectories are not physically plausible. For this reason we propose a physics-informed approach in which aircraft thrust and airspeed are learned from data and are used to condition the existing Base of Aircraft Data (BADA) model, which is physics-based and enforces energy-based constraints on generated trajectories. A set of informative features are identified and used to condition a probabilistic model of aircraft thrust and airspeed, with the proposed scheme demonstrating a 20% improvement in skilfulness across a set of six metrics, compared against a baseline probabilistic model that ignores contextual information such as meteorological conditions.
Fast Surrogate Models for Adaptive Aircraft Trajectory Prediction in En route Airspace
Nick Pepper, Marc Thomas, Zack Xuereb Conti
Trajectory prediction (TP) is crucial for ensuring safety and efficiency in modern air traffic management systems. It is, for example, a core component of conflict detection and resolution tools, arrival sequencing algorithms, capacity planning, as well as several future concepts. However, TP accuracy within operational systems is hampered by a range of epistemic uncertainties such as the mass and performance settings of aircraft and the effect of meteorological conditions on aircraft performance. It can also require considerable computational resources. This paper proposes a method for adaptive TP that has two components: first, a fast surrogate TP model based on linear state space models (LSSM)s with an execution time that was 6.7 times lower on average than an implementation of the Base of Aircraft Data (BADA) in Python. It is demonstrated that such models can effectively emulate the BADA aircraft performance model, which is based on the numerical solution of a partial differential equation (PDE), and that the LSSMs can be fitted to trajectories in a dataset of historic flight data. Secondly, the paper proposes an algorithm to assimilate radar observations using particle filtering to adaptively refine TP accuracy. Comparison with baselines using BADA and Kalman filtering demonstrate that the proposed framework improves system identification and state estimation for both climb and descent phases, with 46.3% and 64.7% better estimates for time to top of climb and bottom of descent compared to the best performing benchmark model. In particular, the particle filtering approach provides the flexibility to capture non-linear performance effects including the CAS-Mach transition.
A framework for assuring the accuracy and fidelity of an AI-enabled Digital Twin of en route UK airspace
Adam Keane, Nick Pepper, Chris Burr, Amy Hodgkin, Dewi Gould, John Korna, Marc Thomas
Digital Twins combine simulation, operational data and Artificial Intelligence (AI), and have the potential to bring significant benefits across the aviation industry. Project Bluebird, an industry-academic collaboration, has developed a probabilistic Digital Twin of en route UK airspace as an environment for training and testing AI Air Traffic Control (ATC) agents. There is a developing regulatory landscape for this kind of novel technology. Regulatory requirements are expected to be application specific, and may need to be tailored to each specific use case. We draw on emerging guidance for both Digital Twin development and the use of Artificial Intelligence/Machine Learning (AI/ML) in Air Traffic Management (ATM) to present an assurance framework. This framework defines actionable goals and the evidence required to demonstrate that a Digital Twin accurately represents its physical counterpart and also provides sufficient functionality across target use cases. It provides a structured approach for researchers to assess, understand and document the strengths and limitations of the Digital Twin, whilst also identifying areas where fidelity could be improved. Furthermore, it serves as a foundation for engagement with stakeholders and regulators, supporting discussions around the regulatory needs for future applications, and contributing to the emerging guidance through a concrete, working example of a Digital Twin. The framework leverages a methodology known as Trustworthy and Ethical Assurance (TEA) to develop an assurance case. An assurance case is a nested set of structured arguments that provides justified evidence for how a top-level goal has been realised. In this paper we provide an overview of each structured argument and a number of deep dives which elaborate in more detail upon particular arguments, including the required evidence, assumptions and justifications.
A Future Capabilities Agent for Tactical Air Traffic Control
Paul Kent, George De Ath, Martin Layton, Allen Hart, Richard Everson, Ben Carvell
Escalating air traffic demand is driving the adoption of automation to support air traffic controllers, but existing approaches face a trade-off between safety assurance and interpretability. Optimisation-based methods such as reinforcement learning offer strong performance but are difficult to verify and explain, while rules-based systems are transparent yet rarely check safety under uncertainty. This paper outlines Agent Mallard, a forward-planning, rules-based agent for tactical control in systemised airspace that embeds a stochastic digital twin directly into its conflict-resolution loop. Mallard operates on predefined GPS-guided routes, reducing continuous 4D vectoring to discrete choices over lanes and levels, and constructs hierarchical plans from an expert-informed library of deconfliction strategies. A depth-limited backtracking search uses causal attribution, topological plan splicing, and monotonic axis constraints to seek a complete safe plan for all aircraft, validating each candidate manoeuvre against uncertain execution scenarios (e.g., wind variation, pilot response, communication loss) before commitment. Preliminary walkthroughs with UK controllers and initial tests in the BluebirdDT airspace digital twin indicate that Mallard's behaviour aligns with expert reasoning and resolves conflicts in simplified scenarios. The architecture is intended to combine model-based safety assessment, interpretable decision logic, and tractable computational performance in future structured en-route environments.
Human-in-the-Loop Testing of AI Agents for Air Traffic Control with a Regulated Assessment Framework
Ben Carvell, Marc Thomas, Andrew Pace, Christopher Dorney, George De Ath, Richard Everson, Nick Pepper, Adam Keane, Samuel Tomlinson, Richard Cannon
We present a rigorous, human-in-the-loop evaluation framework for assessing the performance of AI agents on the task of Air Traffic Control, grounded in a regulator-certified simulator-based curriculum used for training and testing real-world trainee controllers. By leveraging legally regulated assessments and involving expert human instructors in the evaluation process, our framework enables a more authentic and domain-accurate measurement of AI performance. This work addresses a critical gap in the existing literature: the frequent misalignment between academic representations of Air Traffic Control and the complexities of the actual operational environment. It also lays the foundations for effective future human-machine teaming paradigms by aligning machine performance with human assessment targets.
Online Action-Stacking Improves Reinforcement Learning Performance for Air Traffic Control
Ben Carvell, George De Ath, Eseoghene Benjamin, Richard Everson
We introduce online action-stacking, an inference-time wrapper for reinforcement learning policies that produces realistic air traffic control commands while allowing training on a much smaller discrete action space. Policies are trained with simple incremental heading or level adjustments, together with an action-damping penalty that reduces instruction frequency and leads agents to issue commands in short bursts. At inference, online action-stacking compiles these bursts of primitive actions into domain-appropriate compound clearances. Using Proximal Policy Optimisation and the BluebirdDT digital twin platform, we train agents to navigate aircraft along lateral routes, manage climb and descent to target flight levels, and perform two-aircraft collision avoidance under a minimum separation constraint. In our lateral navigation experiments, action stacking greatly reduces the number of issued instructions relative to a damped baseline and achieves comparable performance to a policy trained with a 37-dimensional action space, despite operating with only five actions. These results indicate that online action-stacking helps bridge a key gap between standard reinforcement learning formulations and operational ATC requirements, and provides a simple mechanism for scaling to more complex control scenarios.
A Probabilistic Digital Twin of UK En Route Airspace for Training and Evaluating AI Agents for Air Traffic Control
Nick Pepper, Adam Keane, Amy Hodgkin, Dewi Gould, Edward Henderson, Lynge Lauritsen, Christos Vlahos, George De Ath, Richard Everson, Richard Cannon, Alvaro Sierra Castro, John Korna, Ben Carvell, Marc Thomas
This paper presents the first probabilistic Digital Twin of operational en route airspace, developed for the London Area Control Centre. The Digital Twin is intended to support the development and rigorous human-in-the-loop evaluation of AI agents for Air Traffic Control (ATC), providing a virtual representation of real-world airspace that enables safe exploration of higher levels of ATC automation. This paper makes three significant contributions: firstly, we demonstrate how historical and live operational data may be combined with a probabilistic, physics-informed machine learning model of aircraft performance to reproduce real-world traffic scenarios, while accurately reflecting the level of uncertainty inherent in ATC. Secondly, we develop a structured assurance case, following the Trustworthy and Ethical Assurance framework, to provide quantitative evidence for the Digital Twin's accuracy and fidelity. This is crucial to building trust in this novel technology within this safety-critical domain. Thirdly, we describe how the Digital Twin forms a unified environment for agent testing and evaluation. This includes fast-time execution (up to x200 real-time), a standardised Python-based ''gym'' interface that supports a range of AI agent designs, and a suite of quantitative metrics for assessing performance. Crucially, the framework facilitates competency-based assessment of AI agents by qualified Air Traffic Control Officers through a Human Machine Interface. We also outline further applications and future extensions of the Digital Twin architecture.
Towards Transparent AI Agents for Air Traffic Control
Elhassan Mohamed, Ben Carvell, Rob Procter, Eseoghene Benjamin, George De Ath, Richard Everson
Advances in the use of artificial intelligence agents for air traffic control (ATC) have the potential to reshape modern aviation operations. Despite these advances, the adoption of such agents in safety-critical ATC applications remains limited and this is, in part, due to their lack of transparency. The transparency of such agents is crucial for humans to understand why a specific action is advised or to reason about the agent's behaviour and assess their trustworthiness in real-time. The simpler the agent, the more transparent it can be, but this comes at the cost of flexibility in agent performance. We propose, illustrate, and critically examine mechanisms for providing AI agents for ATC with interpretability and explainability. We focus on agents ranging from simple rule-based systems to optimisation-based and reinforcement-learning agents, and we report initial qualitative feedback from operational ATCOs on prototype explainability mechanisms.
AirTrafficGen: Configurable Air Traffic Scenario Generation with Large Language Models
Dewi Gould, George De Ath, Ben Carvell, Nick Pepper
The manual design of scenarios for Air Traffic Control (ATC) training is a demanding and time-consuming bottleneck that limits the diversity of simulations available to controllers. To address this, we introduce a novel, end-to-end approach, AirTrafficGen, that leverages large language models (LLMs) to automate and control the generation of complex ATC scenarios. Our method uses a purpose-built, graph-based representation to encode sector topology (including airspace geometry, routes, and fixes) into a format LLMs can process. Through rigorous benchmarking, we show that state-of-the-art models like Gemini 2.5 Pro, OpenAI o3, GPT-oss-120b and GPT-5 can generate high-traffic scenarios while maintaining operational realism. Our engineered prompting enables fine-grained control over interaction presence, type, and location. Initial findings suggest these models are also capable of iterative refinement, correcting flawed scenarios based on simple textual feedback. This approach provides a scalable alternative to manual scenario design, addressing the need for a greater volume and variety of ATC training and validation simulations. More broadly, this work showcases the potential of LLMs for complex planning in safety-critical domains.
A Sector-Specific Probabilistic Approach for 4D Aircraft Trajectory Generation
Nick Pepper, George De Ath, Ben Carvell, Amy Hodgkin, Tim Dodwell, Marc Thomas, Richard Everson
Generating realistic four-dimensional trajectories is a fundamental challenge in air traffic control (ATC) that is relevant to both operational tasks and to effective simulation of airspace for the purposes of training controllers and designing new airspaces and/or procedures. Traditional trajectory generation methods are deterministic and use physics-based models with well-calibrated physical parameters and speed schedules. However, these models require knowledge of the clearances issued to aircraft in order to produce a full trajectory. Tools which can marginalise over these clearances to generate a four-dimensional trajectory are valuable in simulations as they emulate the behaviour of human controllers in background sectors, while also reflecting the level of uncertainty present in the system. Consequentially, this work proposes a probabilistic method for generating 4D aircraft trajectories that are specific to a sector of airspace, incorporating multiple routes and allowing local procedures such as co-ordinated entry and exit points to be modelled. The proposed model couples a model for generating plausible aircraft ground tracks with data-driven climb and descent models specific to an aircraft's (ICAO) wake turbulence category. A simple algorithm combines the lateral and vertical trajectories together to produce a four-dimensional (4D) trajectory. A busy sector in the United Kingdom's upper airspace was the focus of the study, which used a dataset comprising one month of aircraft surveillance data. It was found that the proposed model offered improved modelling of aircraft performance and the lateral path followed by aircraft compared to existing, deterministic methods of trajectory generation.
Public trust and blame attribution in human-AI interactions: a comparison between air traffic control and vehicle driving
Peidong Mei, Richard Cannon, Jim A. C. Everett, Peng Liu, Edmond Awad
Artificial Intelligence (AI) has potential to address the increasing demand for capacity in Air Traffic Control (ATC). However, its integration poses several challenges and requires deep understanding of public perception. Insights from the context of Autonomous Vehicles (AVs), in which more studies have been done, can inform such understanding. In this article, we investigate how the public perceives the automated future of ATC in close comparison to AVs. We conducted two studies to examine public trust and blame attribution toward human and AI operators in different Human-AI Interaction (HAI) models, covering three levels of automation (Level 0: AI tool, Level 3: AI trainee, and Level 5: AI manager). We also explored their perceptions of ATC and vehicle driving (VD) by using ten task-related measures (Familiarity, Expertise, Tech Awareness, Openness, Media Discourse, Stake, two measures of Uncertainty, Positive Safety, and Negative Safety) and five agent-related characteristics (Capability, Robustness, Predictability, Honesty and Cooperativeness). The results showed greater trust and less blame attributed to humans in both ATC and VD, except in the Level 3 AI trainee model where humans were blamed more than AI. We also found both similarities and differences in people’s perceptions of the two contexts. Our findings provide evidence-based insights into how the public attribute trust and blame to the operators in ATC and VD. These results will inform industries on the development and implementation of AI integration in aviation and advise policymakers in evaluating public opinion on AI regulation.
Air Traffic Controller Task Demand via Graph Neural Networks: An Interpretable Approach to Airspace Complexity
Edward Henderson, Dewi Gould, George De Ath, Richard Everson, Nick Pepper
Real-time assessment of near-term Air Traffic Controller (ATCO) task demand is a critical challenge in an increasingly crowded airspace, as existing complexity metrics often fail to capture nuanced operational drivers beyond simple aircraft counts. This work introduces an interpretable Graph Neural Network (GNN) framework to address this gap. Our attention-based model predicts the number of upcoming clearances, the instructions issued to aircraft by ATCOs, from interactions within static traffic scenarios. Crucially, we derive an interpretable, per-aircraft task demand score by systematically ablating aircraft and measuring the impact on the model's predictions. Our framework significantly outperforms an ATCO-inspired heuristic and is a more reliable estimator of scenario complexity than established baselines. The resulting tool can attribute task demand to specific aircraft, offering a new way to analyse and understand the drivers of complexity for applications in controller training and airspace redesign.
Learning Generative Models for Climbing Aircraft from Radar Data
Nick Pepper, Marc Thomas
Accurate trajectory prediction for climbing aircraft is hampered by the presence of epistemic uncertainties concerning aircraft operation, which can lead to significant misspecification between predicted and observed trajectories. This paper proposes a generative model for climbing aircraft in which the standard Base of Aircraft Data (BADA) model is enriched by a functional correction to the thrust that is learned from the data. The method offers three features: predictions of the arrival time with 66.3% less error when compared to BADA; generated trajectories that are realistic when compared to test data; and a means of computing confidence bounds for minimal computational cost.
A probabilistic model for aircraft in climb using monotonic functional gaussian process emulators
Nick Pepper, Marc Thomas, George De Ath, Enrico Olivier, Richard Cannon, Richard Everson, Tim Dodwell
Ensuring vertical separation is a key means of maintaining safe separation between aircraft in congested airspace. Aircraft trajectories are modelled in the presence of significant epistemic uncertainty, leading to discrepancies between observed trajectories and the predictions of deterministic models, hampering the task of planning to ensure safe separation. In this paper a probabilistic model is presented, for the purpose of emulating the trajectories of aircraft in climb and bounding the uncertainty of the predicted trajectory. A monotonic, functional representation exploits the spatio-temporal correlations in the radar observations. Through the use of Gaussian Process Emulators, features that parameterise the climb are mapped directly to functional outputs, providing a fast approximation, while ensuring that the resulting trajectory is monotonic. The model was applied as a probabilistic digital twin for aircraft in climb and baselined against BADA, a deterministic model widely used in industry. When applied to an unseen test dataset, the probabilistic model was found to provide a mean prediction that was 21% more accurate, with a 34% sharper forecast.