A two-propeller drone flies to a 2 kg brick, grabs it, and carries it to the green target. It has only left and right throttle and no attitude controller, so the network does all of the balancing. The physics and the trained policy are compiled from the project's Rust code to a 1.2 MB WebAssembly module and run here in your browser. Nothing on this page is pre-recorded.
Click anywhere in the arena to move the target, before or after the grab. You can also click inside the close-up for finer placement. The dashed box marks where targets were placed during training. Targets outside it work most of the time, but the policy never practised there.
One pixel per neuron, brighter the stronger it fires: the 16 inputs, four hidden layers of 128 tanh units, and the two outputs (collective and differential throttle). There is one output head for the whole task, so the switch from approaching to carrying at the grab has to happen inside the hidden layers. Inputs and outputs are squashed with tanh for display.
Reward is the plain task reward: +5 for the grab, +5 for first settling at the target, +0.05 per settled step, and a small alive bonus per airborne step. An episode counts as a success once the drone holds the brick within 1 m of the target, under 0.5 m/s and 8.6° of tilt, for 45 steps in a row. Moving the target restarts that count.
| This policy | Two-head | |
|---|---|---|
| Success | 92.7% | 89.0% |
| Grabbed the brick | 96.8% | 99.6% |
| Crashed | 2.4% | 4.9% |
| Mean time to settle | 7.0 s | 6.5 s |
| Seeds that learned to carry | 5 of 9 | 9 of 9 |
Both checkpoints were scored on the same 1000 deterministic episodes from random starts, and were trained with the same 800-iteration recipe. The two-head network has a separate output head for after the grab, selected by the holding input; this one has to learn that switch itself. It does so less reliably across seeds, but its best seed beats the best two-head seed. The policy on this page also acts deterministically: it applies the mean of its action distribution, without exploration noise.
In-training evaluations run every 20 iterations on 80 episodes, so single points swing by up to ±20 points. Shaded bands are where the two temporary training-aid bonuses ramped to zero; from iteration 320 on, the policy trained on the plain task reward only. The checkpoint is the last one saved, iteration 780.
single_head: one output head for the whole taskruns/single_head/seed6_iter_780