Publication-grade methodology. Bootstrap confidence intervals. Paired effect sizes. Counterfactual fairness tests. PX4 SITL closed-loop validation. Every claim is quantified, every mechanism is isolated.
The validation suite is organized in three phases of increasing rigor, designed to establish safety, generalization, and causal mechanism independently.
| Phase | Tests | Purpose | Key Output |
|---|---|---|---|
| I — Controlled | 1-3 | Structure sensitivity, compute cost, worst-case safety | Operational envelope, adversarial bounds |
| II — Generalization | 4-6 | OOD robustness, causal mechanism, tradeoff frontier | 1,000-world success rate, ablation |
| III — Reviewer | 7-9 | Calibration, counterfactual fairness, significance | Bootstrap CIs, Cohen's d, paired t |
| IV — Flight Stack | 10 | PX4 SITL closed-loop, MAVLink companion injection | 7/7 scenarios, mode transitions in real autopilot |
Each world has independently randomized start/goal positions, wind structure, fuel budget, and up to five categories of perturbation applied simultaneously. No world in the test set was seen during development.
| Perturbation | Parameters | Applied To |
|---|---|---|
| Non-stationary wind | drift 0.3-1.0 units/sec | 50% of worlds |
| Burst sensor noise | p=0.05-0.15, magnitude 3-10 | 50% of worlds |
| Actuator lag | 1-3 timestep delay | 50% of worlds |
| Model mismatch | drag 0.5×-1.5× true value | 43% of worlds |
| Partial observability | limited wind sensing radius | 40% of worlds |
| Metric | MPC Baseline | WHACO | Delta |
|---|---|---|---|
| Mission success rate | 86.7% | 96.5% | +9.8pp |
| Mean fuel consumed (n=860) | 4.489 | 3.736 | -17.4% |
| Missions rescued | — | 105 | |
| Missions lost | — | 7 | |
| Rescue-to-loss ratio | — | 15:1 | |
| WHACO win rate (all worlds) | 3.7% | 83.0% |
| Perturbation | N Worlds | WHACO Win Rate | Status |
|---|---|---|---|
| Wind drift | 487 | 84.0% | Stable |
| Burst noise | 486 | 82.7% | Stable |
| Actuator lag | 487 | 88.9% | Stable |
| Model mismatch | 591 | 82.9% | Stable |
| Partial observability | 430 | 81.9% | Stable |
10,000 bootstrap resamples of the 1,000-world OOD suite establish confidence intervals for all primary metrics. All results are statistically significant with confidence intervals excluding zero.
| Metric | Point Estimate | 95% Bootstrap CI | Effect Size | Status |
|---|---|---|---|---|
| Success delta | +9.8pp | [+7.8, +11.8]pp | CI excludes 0 | Significant |
| Fuel savings | +17.4% | [+16.5, +18.4]% | Cohen's d = 1.30 | Large effect |
| Time cost | +17.6% | [+17.0, +18.3]% | Cohen's d = 1.65 | Large effect |
| Paired t-statistic | 38.0 | df = 859 | p < 0.01 | Significant |
| Rescue ratio | 105:7 (15:1) | [87-124]:[2-13] | CIs non-overlapping | Significant |
The most rigorous test: all four throttle policies use the same live MPC direction oracle, called from each agent's own position with identical wind noise seeds. Only thrust magnitude varies. This isolates the throttle mechanism from any routing advantage.
| Scenario | MPC 75% | WHACO | Constant 60% | Bang-Bang |
|---|---|---|---|---|
| Tight fuel (5.0) | 44% | 100% | 100% | 91% |
| Moderate fuel (8.0) | 100% | 100% | 100% | 64% |
| Hostile headwind | 100% | 100% | 100% | 68% |
| Ultra-tight fuel (4.0) | 0% | 95% | 72% | 89% |
| Scenario | WHACO vs MPC 75% | WHACO vs Constant 60% |
|---|---|---|
| Moderate fuel | -12.2% fuel | +17.5% fuel (WHACO uses more) |
| Hostile headwind | -13.8% fuel | +9.0% fuel (WHACO uses more) |
WHACO uses more fuel than a fixed conservative policy in easy scenarios — because it invests fuel in guaranteeing terminal arrival. This is the correct tradeoff for a feasibility controller.
Six adversarial scenarios specifically designed to break WHACO — trap corridors, zero structure, extreme noise, hostile headwinds, reverse corridors, and ultra-tight fuel.
| Adversarial Scenario | MPC Success | WHACO Success | Delta | Status |
|---|---|---|---|---|
| Trap corridor (strength = 8) | 100% | 100% | +0pp | No degradation |
| Zero structure (no corridor) | 100% | 100% | +0pp | No degradation |
| High noise (σ = 4.0) | 100% | 100% | +0pp | No degradation |
| Tight headwind (fuel = 5) | 0% | 86% | +86pp | Rescue |
| Reverse corridor | 100% | 100% | +0pp | No degradation |
| Ultra-tight (fuel = 4) | 0% | 100% | +100pp | Rescue |
Systematic removal of each WHACO subsystem identifies the causal contribution of each component. The critical comparison is full WHACO vs. random mode selection under tight constraints.
| Variant | Avg Fuel Savings | Avg ΔSuccess | Tight Fuel Success |
|---|---|---|---|
| WHACO (full system) | +12.9% | +10.5pp | 100% |
| No look-ahead probe | +15.4% | +10.5pp | 100% |
| No terminal mode (PUNCH) | +28.5% | +10.5pp | 100% |
| Flat throttle (always 55%) | +33.4% | +10.5pp | 100% |
| Random mode selection | +10.7% | +7.5pp | 88% |
Composite regret analysis across three wind regimes and the full fuel budget range. Positive values indicate WHACO advantage. The composite weights: 40% success regret, 20% time regret, 40% fuel regret.
| Wind Regime | WHACO-ON Zone | Peak Composite | Peak At |
|---|---|---|---|
| Strong corridor (S ~ 3.8) | fuel 3-16 | +70.0% | fuel = 4 |
| Weak corridor (S ~ 0.7) | fuel 3-16 | +66.8% | fuel = 4 |
| No corridor (S ~ 0.2) | fuel 3-16 | +64.4% | fuel = 4 |
High-resolution sweep (fuel 3.0–18.0 in 0.5 increments, 80 trials per point):
| Wind Regime | Success Advantage Zone | Fuel Savings > 10% Zone |
|---|---|---|
| Structured (S ~ 3.8) | fuel ≤ 5.0 | fuel 5.0-8.0 |
| Moderate (S ~ 1.0) | fuel ≤ 6.0 | fuel 5.0-18.0 |
| Hostile (S ~ 0) | fuel ≤ 7.0 | fuel 6.5-18.0 |
WHACO validated inside PX4 Autopilot v1.14.3 (SITL + jMAVSim). The companion process connects via MAVLink UDP, reads vehicle state at 50Hz, runs the full mode selector, and issues velocity commands in OFFBOARD mode — the same architecture used on physical companion computers with Pixhawk flight controllers.
| Scenario / Policy | Result | Distance | Time | Battery | Modes |
|---|---|---|---|---|---|
| Scenario A / WHACO | PASS | 3.6m | 21.0s | 62% | PUNCH → CONSERVE → CRUISE → PUNCH |
| Scenario A / Fixed 75% | PASS | 3.5m | 15.0s | 63% | FIXED_75 |
| Scenario B / WHACO | PASS | 3.6m | 3.3s | 72% | PUNCH → CRUISE |
| Scenario B / Fixed 55% | PASS | 3.7m | 3.0s | 67% | FIXED_55 |
| Scenario B / Fixed 75% | PASS | 3.7m | 2.7s | 67% | FIXED_75 |
| Phase III / WHACO | PASS | 3.7m | 20.0s | 62% | PUNCH → CONSERVE → CRUISE |
| Phase III / Fixed 75% | PASS | 3.8m | 15.6s | 70% | FIXED_75 |
| Challenge | Solution |
|---|---|
| No airspeed sensor (jMAVSim) | Groundspeed-based wind proxy with stall detection override to PUNCH mode |
| Inter-scenario state management | OFFBOARD → MANUAL → force-disarm → heartbeat confirmation → wind reset |
| PX4 OFFBOARD rejection | Setpoint qualification at current position before navigating to start; 5-attempt arm retry with heartbeat verification |
| MAVLink integration | Companion on UDP 14540; reads LOCAL_POSITION_NED, VFR_HUD, SYS_STATUS, HEARTBEAT at configured rates |
All experiments share the following controls to ensure fair comparison:
| Control | Implementation |
|---|---|
| Direction oracle | All policies share the same MPC direction oracle (9 candidate angles, 4-step horizon) |
| Random seeds | Explicit seeds via numpy.random.default_rng for full reproducibility |
| Physics model | Shared semi-implicit Euler integration (dt = 0.5s), identical drag/mass/speed limits |
| Success criterion | Euclidean distance ≤ 4.0 from goal before fuel depletion or time limit |
| Counterfactual test | Live MPC oracle called from each agent's own position — not replayed trajectories |
| Bootstrap method | 10,000 resamples with replacement, BCa confidence intervals |
| Test Phase | Runtime | Dependencies |
|---|---|---|
| Phase I (Tests 1-3) | ~3.5 min | Python 3, NumPy |
| Phase II (Tests 4-6) | ~3.5 min | Python 3, NumPy |
| Phase III (Tests 7-9) | ~8.3 min | Python 3, NumPy, Matplotlib |
| Full suite | ~15 min | Single CPU core, no GPU |