Teaching a quadcopter to race itself
Anduril's AI Grand Prix is a $500,000 autonomous drone-racing competition: fly a five-inch racing quad through a gate course, from onboard FPV vision, with no human pilot.
One honest caveat: this run flew on oracle gate observations — the policy is handed gate poses rather than estimating them from the image. So it demonstrates the control policy, not the perception stack. The camera feed is what the drone sees; it is not what this particular policy was flying on.
Building the simulator first
The official qualifier simulator was not ready when the competition opened. Rather than wait, I built an open-source practice environment so the control and perception work could start immediately — and released it so other entrants could use it too.
- Physics — deterministic 6-DOF rigid-body dynamics with GPU rendering and multi-rate sensors, around a generic five-inch racing quad.
- Flight controller — a real Betaflight firmware build running in software-in-the-loop, in lockstep with the physics, speaking standard RC and PWM over UDP. The policy therefore commands the same firmware a physical drone would run, rather than an idealised actuator model.
- Camera — a forward FPV view matching the published tech spec exactly:
640×360 at 30 Hz,
fx = fy = 320, principal point at the image centre, tilted 20° up, as a real racing drone's camera is. - Course — procedurally generated gate layouts with straights, linked gates, elevation changes, tight turns, and partial visibility.
Putting real firmware in the loop was the decision that mattered. A policy trained against a convenient actuator model learns to exploit it, and that behaviour does not survive contact with the thing you actually fly.
The control problem
Policies output collective thrust and body rates — the same command interface a human pilot's sticks map to in acro mode — and are trained with PPO against full episodes run through to passage, collision, or timeout.
Some details that turned out to matter more than the hyperparameters:
- Terminal-complete returns. Truncating episodes and bootstrapping corrupted the value estimate for exactly the events that decide a race — the crash and the clean pass. Full discounted returns for the privileged critic, GAE at λ = 0.95 for the actor advantages.
- Feasible continuation starts. On hard turns, an agent starting from rest sees almost nothing but terminal failures and never observes a successful example to imitate. Seeding those segments from feasible mid-manoeuvre states gives PPO something to learn from.
- Swept-airframe scoring. Gate passage is checked against the volume the airframe sweeps between timesteps, not the position of its centre point. At 20 m/s a quad moves most of its own body length between frames, so a centre-point test will happily certify a pass that clipped the gate.
- Lexicographic objectives. Complete the course first, pass gates cleanly and with clearance second, avoid contact third, and only then optimise elapsed time. Speed is rewarded only on feasible, collision-free progress.
- Critic calibration before actor updates. Held-out calibration and first-action ranking, checked every iteration — a miscalibrated critic silently poisons every subsequent policy update, and it does so without any obvious symptom in the reward curve.
The speed ladder
The gap between finishing and competing is almost entirely speed. A course leg measures roughly 242 metres between gates. My first reliable full-course policy averaged about 2.7 m/s. Tenth place on the board needed about 10.8 m/s average; first place about 19.5 m/s, with peaks near 28 m/s.
That is not a tuning gap, it is a different flight regime — so I climbed it as a bounded ladder, freezing a known-good completion policy as a baseline and training successors against progressively higher effective-speed targets and peak envelopes, rather than asking one training run to find a 20 m/s policy from scratch.
| Rung | Effective course target | Peak-speed envelope |
|---|---|---|
| Baseline | ~2.7 m/s | measured plant |
| S1 | 5–6 m/s | 10–14 m/s |
| S2 | 7–8 m/s | 14–18 m/s |
| S3 | ≥10.8 m/s | 18–22 m/s |
| S4 | 13–15 m/s | 22–25 m/s |
The trade-off I could not break
The interesting failure is worth stating, because it is the actual state of the work. Speed and precision traded against each other along a frontier I never managed to push past:
- A continuous speed-deficit penalty reached 76.6% gate passage at 96.7% speed retention — fast, but missing too many gates.
- Stronger centering pressure reached 93.3% passage, at only 86.3% speed retention — accurate, but no longer competitive on time.
The bar was 98% passage and 95% speed retention and 0.25 m clearance at the 10th percentile, simultaneously. No candidate met all three, so nothing was promoted. Scalar reward-weight tuning is closed as an avenue: the two objectives were being balanced by a single knob that cannot express "be fast except where geometry is tight". The honest next step is a derivative-free optimiser working directly on terminal race outcomes at target speed, rather than another attempt to find the magic weight.
Stack
Python and PyTorch for the policies, Elodin for physics, Betaflight SITL for the flight controller, containerised for reproducible runs across machines. Roughly 300 experiment write-ups accumulated over the campaign, each recording what was tried, what the evidence was, and whether the idea was promoted or closed.