Deng, Jiajun, et al. "Transvg: End-to-end visual grounding with transformers." Proceedings of the IEEE/CVF International Conference on Computer Vision. 2021.
Kamath, Aishwarya, et al. "MDETR-modulated detection for end-to-end multi-modal understanding." Proceedings of the IEEE/CVF International Conference on Computer Vision. 2021.
Khoreva, Anna, Anna Rohrbach, and Bernt Schiele. "Video object segmentation with language referring expressions." Asian Conference on Computer Vision. Springer, Cham, 2018.
Yu, Licheng, et al. "Modeling context in referring expressions." European Conference on Computer Vision. Springer, Cham, 2016.
Zhou, Luowei, et al. "Grounded video description." Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2019.
Carion, Nicolas, et al. "End-to-end object detection with transformers." European conference on computer vision. Springer, Cham, 2020.
Links