Towards Agile Vision-Based Multi-UAV Flight: Revisiting State Estimation

Accepted to IROS 2026

Abstract

Agile multi-UAV flight requires accurate and low-latency onboard estimation of the kinematic states of neighboring UAVs for collision avoidance, motion coordination, etc. Most vision-based approaches rely on position-only measurements, inferring velocity and acceleration indirectly from displacement. We show that this introduces a fixed structural delay in the estimation of higher-order states, which limits the achievable agility. To address this, we propose to integrate tilt measurements, provided by a state-of-the-art visual detector, which inform about the thrust direction of co-planar multirotor UAVs. We benchmark four position-only and five pose-aware estimators, including a novel formulation of a linear thrust-constraining Kalman Filter, on two real-world and one high-fidelity photorealistic simulated dataset over different levels of agility (3–21 m s−2). In our setup, pose-aware estimation consistently reduces the average velocity and acceleration estimation errors by 40 % and 57 % across the three datasets with the proposed KF formulation outperforming the other estimators. Position-only filters exhibit a constant ~300 ms delay in acceleration step response independent of agility, whereas the tilt-constrained estimators operate near the physical response limit given by the camera frame-rate by observing the change in thrust direction before the displacement accumulates. In a closed-loop leader–follower simulated experiment with NMPC control, position-only estimation of the leader's state fails to facilitate stable hovering of the follower, while the proposed estimator enables tracking of lateral maneuvers exceeding 2 g of acceleration.

Video

Code & Data

📊 State estimation & CMA-ES tuning code github.com/anonymiros-bot/tilt-detection-and-estimation
📦 UAV state estimation benchmark Google Drive
🔍 YOLOv5-6D code github.com/anonymiros-bot/YOLOv5-6D-Pose
🧠 YOLOv5-6D weights Google Drive (unreal, phantom4, mavic2)
🖼️ Flight sequence images (processed) Google Drive — Unreal (ours) + Phantom 4 & Mavic 2 (MAV6D, Zheng et al.)
🔗 MAV6D dataset (Zheng et al.) Original source — unprocessed Phantom 4 & Mavic 2 images
🖼️ Detector training images (Unreal) Google Drive

Pipeline

images → YOLOv5-6D → keypoints → PnP → pnp_pos, pnp_q → estimators → filtered state
         (detection)  (9 pts)   (pose)  (measurements)  (Z-KF, etc.)  (pos, vel, acc)

Quick Start

Path A: Reproduce paper results (minimal)

Detections are pre-computed — no images or YOLOv5-6D needed.

# 1. Clone
git clone https://github.com/anonymiros-bot/tilt-detection-and-estimation
cd tilt-detection-and-estimation

# 2. Get data (choose one)
./setup_data.sh                        # downloads data + 10 independent tuning runs per filter

# 3. Evaluate (reproduces Table II)
./run_evaluation.sh unreal_train unreal_test      # Unreal
./run_evaluation.sh unreal_train phantom4_test    # Phantom 4
./run_evaluation.sh unreal_train mavic2_test      # Mavic 2

# (optional) Re-tune estimators from scratch
./run_tuning.sh unreal_train

Path B: Run detection from images (full pipeline)

Generate your own detections using YOLOv5-6D + PnP. Based on YOLOv5-6D-Pose by Viviers et al., see original repo for details.

# 1. Clone & install
git clone https://github.com/anonymiros-bot/YOLOv5-6D-Pose
cd YOLOv5-6D-Pose
pip install -r requirements.txt
cd utils && python setup.py build_ext --inplace && cd ..

# 2. Download pre-trained weights (~2.7 GB)
./setup_weights.sh

# 3. Download flight sequence images (~70 GB, includes Unreal + Phantom 4 + Mavic 2)
./setup_data.sh

# 4. Run detection + PnP
python pnp_eval_all_unreal_flying.py   # 140 flying sequences
python pnp_eval_all_unreal_step.py     # 35 step response sequences
python pnp_eval_all_real_mavic2.py     # DJI Mavic 2
python pnp_eval_all_real_phantom4.py   # DJI Phantom 4

# 5. Run state estimation on the outputs using Path A steps 2-3

Path C: Train your own detector (Unreal)

We release ~13k training images for benchmarking new UAV pose estimation methods. This training set is completely independent from the test sequences used for evaluation. See original YOLOv5-6D-Pose repo for full training documentation.

# 1. Clone & install (same as Path B step 1)

# 2. Download Unreal training images (~15 GB)
./setup_train_data.sh

# 3. Train (adjust --batch to fit your GPU memory)
python train.py \
    --data configs/unreal.yaml \
    --cfg models/yolov5x_6dpose_bifpn.yaml \
    --hyp configs/hyp.single.yaml \
    --weights yolov5x.pt \
    --batch 40 \
    --epochs 1000 \
    --optimizer Adam

# 4. Run Path B steps 3–5 using your trained weights

Setup scripts attempt automatic download via gdown — if download fails, see the script output for manual download links.

A benchmark dataset

Pre-computed detections (for estimator development)

If you are developing new state estimation methods, you can use our benchmark data directly — no images or detector needed. Each JSON contains ground truth states and YOLOv5-6D + PnP detections across multiple agility levels, ready to use as input to your estimator. See Path A above.

├── unreal_train/train.json     # 70 seq for tuning 
├── unreal_test/test.json       # 70 seq for evaluation
├── acc_step_test/test.json     # 35 seq for latency analysis (Table V)
├── phantom4_train/train.json   # from MAV6D (Zheng et al.)
├── phantom4_test/test.json     # 68 seq, static camera, 1.9–5.1 m
├── mavic2_train/train.json     # from MAV6D (Zheng et al.)
└── mavic2_test/test.json       # 61 seq, static camera, 1.9–5.1 m

Each sequence contains:

Ground truth        pos, rts_vel, rts_acc    position, velocity, acceleration
                    quat, omega_b            orientation, angular velocity
                    t                        timestamps
                    level                    agility: 3, 6, 9, 12, 15, 18, 21 m/s² (Unreal only)

Detections          pnp_pos, pnp_q           position & orientation from YOLOv5-6D + PnP

Flight sequence images (for pose detection/tracking development)

If you are developing new visual detection or tracking methods — monocular depth, feature-based tracking, alternative pose estimators, etc. — we provide photorealistic simulated (Unreal) and real-world (Phantom 4 & Mavic 2, from MAV6D) image sequences of UAV flight with varying camera distances and agility levels. A Python data loader (load_dataset.py, included in YOLOv5-6D repo) provides direct access to images, ground truth, and detections per frame:

from load_dataset import load_dataset, list_datasets

# See what's available
list_datasets()
# → unreal_flying, unreal_step, mavic2, phantom4

# Load all sequences from a dataset
data = load_dataset("unreal_flying")
seq  = data["42"]                       # one of 140 main unreal sequences

seq["t"]                    # (N,)   timestamps
seq["pos"]                  # (N,3)  GT position
seq["quat"]                 # (N,4)  GT orientation xyzw
seq["rts_vel"]              # (N,3)  GT velocity
seq["rts_acc"]              # (N,3)  GT acceleration
seq["pnp_pos"]              # (N,3)  detected position (NaN if failed)
seq["pnp_q"]                # (N,4)  detected orientation xyzw (NaN if failed)
seq["images"]               # [N]    paths to image files
seq["metadata"]["level"]    # agility level (Unreal only)


# Real-world datasets use session/sequence keys (to stay aligned with original MAV6D dataset structure)
data = load_dataset("mavic2")
seq  = data["01/0101"]
seq["pos"]                  # (N,3) GT position in VICON frame
# ... same fields as above
seq["images"]               # [N]   paths to image files

# If you don't want to study the directory structure — just iterate over keys
for key, seq in data.items():
    print(key, seq["pos"].shape, len(seq["images"]))

📝 Repository structure and documentation are being refined based on user feedback.

Questions or issues? This email address is being protected from spambots. You need JavaScript enabled to view it.