Abstract Mainstream real-time object detectors, like the YOLO series, balance speed and accuracy but are bottlenecked by Non-Maximum Suppression (NMS) for post-processing. While end-to-end Transformer-based detectors show potential by eliminating NMS, their high computational cost impedes real-time application. Concurrently, alternatives like DECO, using a pure convolutional framework, are hampered by local receptive fields, limiting global context modeling and limiting the accuracy on lightweight models. To address this, we introduce HA-DETR, a Hybrid Architecture DETR fusing the local feature extraction of convolutions with the global context modeling of Transformers, particularly effective in resource-constrained scenarios. HA-DETR uses an efficient multi-scale encoder and an effective hybrid decoder that integrates convolutional query refinement with cross-attention to accelerate predictions. This hybrid design, however, exacerbates a training dynamic mismatch. To mitigate this, we propose the Decoupled Gamma Loss (DGL), which introduces independent modulating factors, \(\gamma _{pos}\) and \(\gamma _{neg}\), to alleviate sample imbalance from disparate convergence rates of the hybrid components. Experiments validate our method. On the COCO dataset, our lightweight HA-DETR-R18 achieves 48.4 AP at 68 FPS on a V100 GPU, surpassing RT-DETR-R18 by 1.9 AP with a 13% speedup and outperforming DECO-R18 by 7.9 AP. Our work provides a competitive architectural alternative for efficient and accurate real-time object detection. Similar content being viewed by others Data availability No datasets were generated or analysed during the current study. References Chen, X., Ma, H., Wan, J., Li, B. & Xia, T. Multi-view 3d object detection network for autonomous driving. In Proc. IEEE Conference on Computer Vision and Pattern Recognition, 1907–1915 (2017). Ess, A., Schindler, K., Leibe, B. & Van Gool, L. Object detection and tracking for autonomous navigation in dynamic environments. Int. J. Robot. Res. 29, 1707–1725 (2010). Redmon, J. You only look once: Unified, real-time object detection. In Proc. IEEE Conference on Computer Vision and