Validation & Statistical Rigor

Ten-Test Validation Suite
Simulation Through Flight Stack

Publication-grade methodology. Bootstrap confidence intervals. Paired effect sizes. Counterfactual fairness tests. PX4 SITL closed-loop validation. Every claim is quantified, every mechanism is isolated.

Overview — Test Architecture

Three-Phase Validation Protocol

The validation suite is organized in three phases of increasing rigor, designed to establish safety, generalization, and causal mechanism independently.

PhaseTestsPurposeKey Output
I — Controlled1-3Structure sensitivity, compute cost, worst-case safetyOperational envelope, adversarial bounds
II — Generalization4-6OOD robustness, causal mechanism, tradeoff frontier1,000-world success rate, ablation
III — Reviewer7-9Calibration, counterfactual fairness, significanceBootstrap CIs, Cohen's d, paired t
IV — Flight Stack10PX4 SITL closed-loop, MAVLink companion injection7/7 scenarios, mode transitions in real autopilot
10
Test families
1,000
OOD worlds
10,000
Bootstrap samples
5
Perturbation types
Test 4 — Out-of-Distribution Stress Test

1,000 Randomly Generated Worlds

Each world has independently randomized start/goal positions, wind structure, fuel budget, and up to five categories of perturbation applied simultaneously. No world in the test set was seen during development.

Perturbation Categories

PerturbationParametersApplied To
Non-stationary winddrift 0.3-1.0 units/sec50% of worlds
Burst sensor noisep=0.05-0.15, magnitude 3-1050% of worlds
Actuator lag1-3 timestep delay50% of worlds
Model mismatchdrag 0.5×-1.5× true value43% of worlds
Partial observabilitylimited wind sensing radius40% of worlds

Headline Results

MetricMPC BaselineWHACODelta
Mission success rate86.7%96.5%+9.8pp
Mean fuel consumed (n=860)4.4893.736-17.4%
Missions rescued—105
Missions lost—7
Rescue-to-loss ratio—15:1
WHACO win rate (all worlds)3.7%83.0%

Robustness by Perturbation Type

PerturbationN WorldsWHACO Win RateStatus
Wind drift48784.0%Stable
Burst noise48682.7%Stable
Actuator lag48788.9%Stable
Model mismatch59182.9%Stable
Partial observability43081.9%Stable
Key finding: WHACO's advantage is stable across all perturbation categories (81.9%–88.9% win rate). The benefit is not tied to exploiting a specific environmental feature — it generalizes across perturbation types.
Test 9 — Bootstrap Significance Analysis

Statistical Rigor

10,000 bootstrap resamples of the 1,000-world OOD suite establish confidence intervals for all primary metrics. All results are statistically significant with confidence intervals excluding zero.

MetricPoint Estimate95% Bootstrap CIEffect SizeStatus
Success delta+9.8pp[+7.8, +11.8]ppCI excludes 0Significant
Fuel savings+17.4%[+16.5, +18.4]%Cohen's d = 1.30Large effect
Time cost+17.6%[+17.0, +18.3]%Cohen's d = 1.65Large effect
Paired t-statistic38.0df = 859p < 0.01Significant
Rescue ratio105:7 (15:1)[87-124]:[2-13]CIs non-overlappingSignificant
Cohen's d = 1.30 indicates a large effect size for fuel savings. For reference, d > 0.8 is conventionally classified as "large." The rescue ratio confidence intervals do not overlap, confirming the 15:1 advantage is not a sampling artifact.
Test 8 — Counterfactual Fairness

Same Route, Different Throttle

The most rigorous test: all four throttle policies use the same live MPC direction oracle, called from each agent's own position with identical wind noise seeds. Only thrust magnitude varies. This isolates the throttle mechanism from any routing advantage.

ScenarioMPC 75%WHACOConstant 60%Bang-Bang
Tight fuel (5.0)44%100%100%91%
Moderate fuel (8.0)100%100%100%64%
Hostile headwind100%100%100%68%
Ultra-tight fuel (4.0)0%95%72%89%
Causal isolation: Under ultra-tight constraints, MPC throttle achieved 0% success. WHACO achieved 95% on the same directional commands. This confirms the throttle mechanism — not routing luck — is the causal driver of the performance delta.

Fuel Efficiency on Shared Routes

ScenarioWHACO vs MPC 75%WHACO vs Constant 60%
Moderate fuel-12.2% fuel+17.5% fuel (WHACO uses more)
Hostile headwind-13.8% fuel+9.0% fuel (WHACO uses more)

WHACO uses more fuel than a fixed conservative policy in easy scenarios — because it invests fuel in guaranteeing terminal arrival. This is the correct tradeoff for a feasibility controller.

Test 3 — Adversarial Safety

Worst-Case Guarantee

Six adversarial scenarios specifically designed to break WHACO — trap corridors, zero structure, extreme noise, hostile headwinds, reverse corridors, and ultra-tight fuel.

Adversarial ScenarioMPC SuccessWHACO SuccessDeltaStatus
Trap corridor (strength = 8)100%100%+0ppNo degradation
Zero structure (no corridor)100%100%+0ppNo degradation
High noise (σ = 4.0)100%100%+0ppNo degradation
Tight headwind (fuel = 5)0%86%+86ppRescue
Reverse corridor100%100%+0ppNo degradation
Ultra-tight (fuel = 4)0%100%+100ppRescue
Safety property: No degradation in mission success was observed across tested adversarial conditions. In the two hardest cases, WHACO rescued missions that were infeasible under conventional throttle policies.
Test 5 — Ablation Analysis

Causal Mechanism Identification

Systematic removal of each WHACO subsystem identifies the causal contribution of each component. The critical comparison is full WHACO vs. random mode selection under tight constraints.

VariantAvg Fuel SavingsAvg ΔSuccessTight Fuel Success
WHACO (full system)+12.9%+10.5pp100%
No look-ahead probe+15.4%+10.5pp100%
No terminal mode (PUNCH)+28.5%+10.5pp100%
Flat throttle (always 55%)+33.4%+10.5pp100%
Random mode selection+10.7%+7.5pp88%
Counter-intuitive finding: Removing components increases average fuel savings. This is because WHACO's modes (PUNCH, mode-switching) deliberately trade fuel savings for robustness. The full system invests fuel to guarantee terminal arrival. Under tight constraints, intelligent selection achieves 100% success vs. random's 88%. Mode switching minimizes tail-risk, not mean cost.
Test 6 — Regret Curves

Deployment Frontier

Composite regret analysis across three wind regimes and the full fuel budget range. Positive values indicate WHACO advantage. The composite weights: 40% success regret, 20% time regret, 40% fuel regret.

Wind RegimeWHACO-ON ZonePeak CompositePeak At
Strong corridor (S ~ 3.8)fuel 3-16+70.0%fuel = 4
Weak corridor (S ~ 0.7)fuel 3-16+66.8%fuel = 4
No corridor (S ~ 0.2)fuel 3-16+64.4%fuel = 4
Deployment rule: WHACO is net-positive across the entire tested fuel range in all three wind regimes. Peak value occurs at ultra-tight budgets (fuel 3-5), where the success rescue effect dominates. The time cost (+17.6% slower) is compensated by fuel savings and mission success improvement.

Calibration: Value vs. Constraint Tightness (Test 7)

High-resolution sweep (fuel 3.0–18.0 in 0.5 increments, 80 trials per point):

Wind RegimeSuccess Advantage ZoneFuel Savings > 10% Zone
Structured (S ~ 3.8)fuel ≤ 5.0fuel 5.0-8.0
Moderate (S ~ 1.0)fuel ≤ 6.0fuel 5.0-18.0
Hostile (S ~ 0)fuel ≤ 7.0fuel 6.5-18.0
Critical insight: In hostile conditions (no exploitable structure), WHACO's value extends further into the fuel range. When the environment is harder, WHACO helps more. This is the opposite of what a corridor-exploitation strategy would produce — confirming the constraint-aware mechanism.
Test 10 — PX4 SITL Closed-Loop Validation

Flight-Stack Integration

WHACO validated inside PX4 Autopilot v1.14.3 (SITL + jMAVSim). The companion process connects via MAVLink UDP, reads vehicle state at 50Hz, runs the full mode selector, and issues velocity commands in OFFBOARD mode — the same architecture used on physical companion computers with Pixhawk flight controllers.

7/7
Scenarios passed
50Hz
Control rate
96
Harness checks
3
Scenarios

PX4 SITL Results

Scenario / PolicyResultDistanceTimeBatteryModes
Scenario A / WHACOPASS3.6m21.0s62%PUNCH → CONSERVE → CRUISE → PUNCH
Scenario A / Fixed 75%PASS3.5m15.0s63%FIXED_75
Scenario B / WHACOPASS3.6m3.3s72%PUNCH → CRUISE
Scenario B / Fixed 55%PASS3.7m3.0s67%FIXED_55
Scenario B / Fixed 75%PASS3.7m2.7s67%FIXED_75
Phase III / WHACOPASS3.7m20.0s62%PUNCH → CONSERVE → CRUISE
Phase III / Fixed 75%PASS3.8m15.6s70%FIXED_75

Engineering Validation

ChallengeSolution
No airspeed sensor (jMAVSim)Groundspeed-based wind proxy with stall detection override to PUNCH mode
Inter-scenario state managementOFFBOARD → MANUAL → force-disarm → heartbeat confirmation → wind reset
PX4 OFFBOARD rejectionSetpoint qualification at current position before navigating to start; 5-attempt arm retry with heartbeat verification
MAVLink integrationCompanion on UDP 14540; reads LOCAL_POSITION_NED, VFR_HUD, SYS_STATUS, HEARTBEAT at configured rates
Key finding: WHACO's mode selector is fully active inside the PX4 control loop. Scenario A shows 4 mode transitions (PUNCH → CONSERVE → CRUISE → PUNCH), confirming the controller adapts to real-time wind and fuel conditions within the autopilot's attitude control envelope. The 96-check dry-run harness validates all companion logic without requiring a live PX4 instance.
Methodology

Experimental Controls

All experiments share the following controls to ensure fair comparison:

ControlImplementation
Direction oracleAll policies share the same MPC direction oracle (9 candidate angles, 4-step horizon)
Random seedsExplicit seeds via numpy.random.default_rng for full reproducibility
Physics modelShared semi-implicit Euler integration (dt = 0.5s), identical drag/mass/speed limits
Success criterionEuclidean distance ≤ 4.0 from goal before fuel depletion or time limit
Counterfactual testLive MPC oracle called from each agent's own position — not replayed trajectories
Bootstrap method10,000 resamples with replacement, BCa confidence intervals

Reproducibility

Test PhaseRuntimeDependencies
Phase I (Tests 1-3)~3.5 minPython 3, NumPy
Phase II (Tests 4-6)~3.5 minPython 3, NumPy
Phase III (Tests 7-9)~8.3 minPython 3, NumPy, Matplotlib
Full suite~15 minSingle CPU core, no GPU