АДАПТИВНИЙ КАСКАД ЗА СКЛАДНІСТЮ СЦЕНИ ДЛЯ ВИЯВЛЕННЯ ОБ’ЄКТІВ НА АЕРОЗНІМКАХ БЕЗПІЛОТНИКІВ: КОЛИ ВИБІРКОВА ЕСКАЛАЦІЯ ОБЧИСЛЕНЬ ВИПРАВДАНА
DOI:
https://doi.org/10.28925/2663-4023.2026.34.1340Keywords:
object detection, UAV imagery, model cascade, detector ensemble, adaptive computation, domain shift, VisDrone, UAVDTAbstract
Unmanned aerial vehicles are a standard video-monitoring instrument in security tasks — from perimeter and critical-infrastructure protection to situational awareness — and detector ensembles raise the accuracy of object detection in aerial imagery at a manifold per-frame compute cost, which is critical for onboard systems. A Scene-Complexity Adaptive Cascade (SCAC) is proposed and investigated: a cheap first-tier detector processes every frame and computes a dimensionless, annotation-free complexity signal; only frames whose complexity exceeds a threshold are escalated to an expensive second-tier WBF ensemble. The threshold is not hand-tuned — it is calibrated by cross-fitting at the video-sequence level for a target average budget. The methodological contribution is an oracle feasibility gate: a greedy reference curve ordered by the realized per-frame ensemble gain is constructed from cached predictions, and realizable complexity signals are compared against it. In the training domain (VisDrone-DET) the ensemble gain is dense across frames (75–90 % of frames improve), so realizable complexity signals do not beat uniform fusion at the same budget, and the popular uncertainty signal is consistently worse than random escalation. Under domain shift (zero-shot transfer to UAVDT with official condition labels), by contrast, the ensemble gain is sparse (41 % of frames; at night the ensemble hurts), and the cross-fitted cascade with a detection-density signal reaches the quality of the full ensemble (0.618 mAP@50, scene-balanced) at 45 % of its compute, significantly outperforming the single model (+0.022 on the full sample; bootstrap mean +0.017, 95 % CI [+0.002; +0.032]) and random escalation, including its scene-stratified variant (+0.019 on the full sample; bootstrap mean +0.016, 95 % CI [+0.001; +0.029]). The realized escalation profile (9 % night / 22 % day / 45 % fog) reproduces scene adaptivity without any scene classifier, bypassing the known bottleneck of routing — the unrecognizability of adverse conditions from frame appearance. A practical criterion is formulated: selective compute escalation pays off only where the expensive model's gain is sparse across frames — typically under domain shift, not in the training domain.
Downloads
References
Solovyev, R., Wang, W., & Gabruseva, T. (2021). Weighted boxes fusion: Ensembling boxes from different object detection models. Image and Vision Computing, 107, 104117. https://doi.org/10.1016/j.imavis.2021.104117
Ruda, Kh. S., & Kromkach, V. O. (in press). Scene-dependent ensemble object detection in UAV aerial imagery: Limits of the utility of scene-based routing. Scientific Bulletin of UNFU.
Viola, P., & Jones, M. (2001). Rapid object detection using a boosted cascade of simple features. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 511-518). https://doi.org/10.1109/CVPR.2001.990517
Teerapittayanon, S., McDanel, B., & Kung, H. T. (2016). BranchyNet: Fast inference via early exiting from deep neural networks. In 23rd International Conference on Pattern Recognition (ICPR) (pp. 2464-2469). https://doi.org/10.1109/ICPR.2016.7900006
Kang, D., Emmons, J., Abuzaid, F., Bailis, P., & Zaharia, M. (2017). NoScope: Optimizing neural network queries over video at scale. Proceedings of the VLDB Endowment, 10(11), 1586-1597. https://doi.org/10.14778/3137628.3137664
Jiang, J., Ananthanarayanan, G., Bodik, P., Sen, S., & Stoica, I. (2018). Chameleon: Scalable adaptation of video analytics. In Proceedings of the ACM SIGCOMM 2018 Conference (pp. 253-266). https://doi.org/10.1145/3230543.3230574
Ghodrati, A., Bejnordi, B. E., & Habibian, A. (2021). FrameExit: Conditional early exiting for efficient video recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 15608-15618). https://doi.org/10.1109/CVPR46437.2021.01535
Matsubara, Y., Levorato, M., & Restuccia, F. (2022). Split computing and early exiting for deep learning applications: Survey and research challenges. ACM Computing Surveys, 55(5), 1-30. https://doi.org/10.1145/3527155
Han, Y., Huang, G., Song, S., Yang, L., Wang, H., & Wang, Y. (2022). Dynamic neural networks: A survey. IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(11), 7436-7456. https://doi.org/10.1109/TPAMI.2021.3117837
Rahmath P., H., Srivastava, V., Chaurasia, K., Pacheco, R. G., & Couto, R. S. (2024). Early-exit deep neural network – A comprehensive survey. ACM Computing Surveys, 57(3), 1-37. https://doi.org/10.1145/3698767
Jacobs, R. A., Jordan, M. I., Nowlan, S. J., & Hinton, G. E. (1991). Adaptive mixtures of local experts. Neural Computation, 3(1), 79-87. https://doi.org/10.1162/neco.1991.3.1.79
Shazeer, N., Mirhoseini, A., Maziarz, K., Davis, A., Le, Q., Hinton, G., & Dean, J. (2017). Outrageously large neural networks: The sparsely-gated mixture-of-experts layer. In International Conference on Learning Representations (ICLR). https://doi.org/10.48550/arXiv.1701.06538
Zhu, X., Lyu, S., Wang, X., & Zhao, Q. (2021). TPH-YOLOv5: Improved YOLOv5 based on transformer prediction head for object detection on drone-captured scenarios. In Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops (ICCVW) (pp. 2778-2788). https://doi.org/10.1109/ICCVW54120.2021.00312
Zhang, H., Liu, K., Gan, Z., & Zhu, G.-N. (2025). UAV-DETR: Efficient end-to-end object detection for unmanned aerial vehicle imagery. arXiv. https://doi.org/10.48550/arXiv.2501.01855
Nazarkevych, M. A., & Oleksiv, N. T. (2024). Object recognition system based on the YOLO model. Ukrainian Journal of Information Technology, 6(1), 120-126. https://doi.org/10.23939/ujit2024.01.120
Filimonchuk, T. V., Koltun, Yu. M., & Maslov, M. K. (2026). Adaptive model for equipment recognition with attention mechanisms based on YOLOv11. Information and Control Systems for Railway Transport, 31(2). https://doi.org/10.18664/ikszt.v31i2.362205
Zhu, P., Wen, L., Du, D., Bian, X., Fan, H., Hu, Q., & Ling, H. (2022). Detection and tracking meet drones challenge. IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(11), 7380-7399. https://doi.org/10.1109/TPAMI.2021.3119563
Du, D., Qi, Y., Yu, H., Yang, Y., Duan, K., Li, G., Zhang, W., Huang, Q., & Tian, Q. (2018). The unmanned aerial vehicle benchmark: Object detection and tracking. In Computer Vision – ECCV 2018 (pp. 370-386). https://doi.org/10.1007/978-3-030-01249-6_23
Saksena, S. K. (2025). VisDrone detection model zoo [Model collection]. Hugging Face. https://huggingface.co/collections/dronefreak/visdrone-detection-model-zoo
Wang, C.-Y., Yeh, I-H., & Liao, H.-Y. M. (2024). YOLOv9: Learning what you want to learn using programmable gradient information. In Computer Vision – ECCV 2024 (Lecture Notes in Computer Science, Vol. 15089, pp. 1–21). https://doi.org/10.1007/978-3-031-72751-1_1
Wang, A., Chen, H., Liu, L., Chen, K., Lin, Z., Han, J., & Ding, G. (2024). YOLOv10: Real-time end-to-end object detection. Advances in Neural Information Processing Systems, 37, 107984-108011. https://doi.org/10.48550/arXiv.2405.14458
Khanam, R., & Hussain, M. (2024). YOLOv11: An overview of the key architectural enhancements. arXiv. https://doi.org/10.48550/arXiv.2410.17725
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Христина Руда, Владислав Кромкач

This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.