rboyd.co

Teaching a quadcopter to race itself

Anduril's AI Grand Prix is a $500,000 autonomous drone-racing competition: fly a five-inch racing quad through a gate course, from onboard FPV vision, with no human pilot.

17 / 17gates passed, full Level 2 course
32.9sflight time, every crossing clean
21stplace
2.7 → 22m/s target speed ladder
A complete Level 2 run. The policy launches from the pad and flies all seventeen gates in 32.9 seconds, passing every one cleanly — radial offsets from the gate centre range from 8 cm to 77 cm, and all seventeen clear the strict 0.75 threshold. This is the onboard camera view at the competition's exact spec: 640×360 at 30 Hz, tilted 20° up.

One honest caveat: this run flew on oracle gate observations — the policy is handed gate poses rather than estimating them from the image. So it demonstrates the control policy, not the perception stack. The camera feed is what the drone sees; it is not what this particular policy was flying on.

Building the simulator first

The official qualifier simulator was not ready when the competition opened. Rather than wait, I built an open-source practice environment so the control and perception work could start immediately — and released it so other entrants could use it too.

Putting real firmware in the loop was the decision that mattered. A policy trained against a convenient actuator model learns to exploit it, and that behaviour does not survive contact with the thing you actually fly.

An early flight in the practice simulator, with Betaflight in the loop. The physics, the flight controller, and the camera model are three separate pieces that have to agree with each other at every timestep — most of the early work was making that true.

The control problem

Policies output collective thrust and body rates — the same command interface a human pilot's sticks map to in acro mode — and are trained with PPO against full episodes run through to passage, collision, or timeout.

Some details that turned out to matter more than the hyperparameters:

The speed ladder

The gap between finishing and competing is almost entirely speed. A course leg measures roughly 242 metres between gates. My first reliable full-course policy averaged about 2.7 m/s. Tenth place on the board needed about 10.8 m/s average; first place about 19.5 m/s, with peaks near 28 m/s.

That is not a tuning gap, it is a different flight regime — so I climbed it as a bounded ladder, freezing a known-good completion policy as a baseline and training successors against progressively higher effective-speed targets and peak envelopes, rather than asking one training run to find a 20 m/s policy from scratch.

RungEffective course targetPeak-speed envelope
Baseline~2.7 m/smeasured plant
S15–6 m/s10–14 m/s
S27–8 m/s14–18 m/s
S3≥10.8 m/s18–22 m/s
S413–15 m/s22–25 m/s

The trade-off I could not break

The interesting failure is worth stating, because it is the actual state of the work. Speed and precision traded against each other along a frontier I never managed to push past:

The bar was 98% passage and 95% speed retention and 0.25 m clearance at the 10th percentile, simultaneously. No candidate met all three, so nothing was promoted. Scalar reward-weight tuning is closed as an avenue: the two objectives were being balanced by a single knob that cannot express "be fast except where geometry is tight". The honest next step is a derivative-free optimiser working directly on terminal race outcomes at target speed, rather than another attempt to find the magic weight.

Stack

Python and PyTorch for the policies, Elodin for physics, Betaflight SITL for the flight controller, containerised for reproducible runs across machines. Roughly 300 experiment write-ups accumulated over the campaign, each recording what was tried, what the evidence was, and whether the idea was promoted or closed.