Reinforcement playground · live policy

BiProp Pickup

A two-propeller drone flies to a 2 kg brick, grabs it, and carries it to the green target. It has only left and right throttle and no attitude controller, so the network does all of the balancing. The physics and the trained policy are compiled from the project's Rust code to a 1.2 MB WebAssembly module and run here in your browser. Nothing on this page is pre-recorded.

Loading the policy…

Click anywhere in the arena to move the target, before or after the grab. You can also click inside the close-up for finer placement. The dashed box marks where targets were placed during training. Targets outside it work most of the time, but the policy never practised there.

Speed

Telemetry

Altitude–
Tilt–
Speed–
To target–
Network activationspositivenegative

One pixel per neuron, brighter the stronger it fires: the 16 inputs, four hidden layers of 128 tanh units, and the two outputs (collective and differential throttle). There is one output head for the whole task, so the switch from approaching to carrying at the grab has to happen inside the hidden layers. Inputs and outputs are squashed with tanh for display.

left throttleright throttlelast 5 s, 0–1
episode return

Reward is the plain task reward: +5 for the grab, +5 for first settling at the target, +0.05 per settled step, and a small alive bonus per airborne step. An episode counts as a success once the drone holds the brick within 1 m of the target, under 0.5 m/s and 8.6° of tilt, for 45 steps in a row. Moving the target restarts that count.

1000-episode evaluation

This policyTwo-head
Success92.7%89.0%
Grabbed the brick96.8%99.6%
Crashed2.4%4.9%
Mean time to settle7.0 s6.5 s
Seeds that learned to carry5 of 99 of 9

Both checkpoints were scored on the same 1000 deterministic episodes from random starts, and were trained with the same 800-iteration recipe. The two-head network has a separate output head for after the grab, selected by the holding input; this one has to learn that switch itself. It does so less reliably across seeds, but its best seed beats the best two-head seed. The policy on this page also acts deterministically: it applies the mean of its action distribution, without exploration noise.

How it got there

In-training evaluations run every 20 iterations on 80 episodes, so single points swing by up to ±20 points. Shaded bands are where the two temporary training-aid bonuses ramped to zero; from iteration 320 on, the policy trained on the plain task reward only. The checkpoint is the last one saved, iteration 780.

The checkpoint

Network
MLP, 16 inputs → 4 × 128 tanh → 2 outputs (collective and differential throttle), single_head: one output head for the whole task
Observations
Position, velocity, sin/cos of attitude, angular velocity, brick and target offsets (plus fine-scale tanh copies for the ~10 cm a grab needs), holding flag
Algorithm
PPO on Burn 0.21, CPU backend; 8 actors × 2048 steps = 16,384 samples per iteration; about 30 minutes for 800 iterations
Seed
6 of 9; checkpoint runs/single_head/seed6_iter_780
Relabeling
Hindsight relabeling of both sparse events (grab and arrival) with V-trace off-policy correction
Curriculum
Training only: 50% of episodes start just above the brick, 25% already holding it. Episodes here always start from the full task.
Physics
Rapier2D at 60 Hz; 1 kg drone, 10 × 4 cm, 24 N max thrust per propeller; brick 21.5 × 6.5 cm